# The Live AI Voice Screen, From Invite to Completed Conversation

*You will run a live AI voice screen start to finish: set up audio, open tight, answer adaptive follow-ups with one concrete example each, and avoid the delivery patterns that lower a competency score.*

- Canonical URL: https://www.refolk.ai/candidates/guides/live-ai-voice-screen-playbook
- Pillar: Interviewing
- Format: Playbook
- Published: 2026-08-28
- Last reviewed: 2026-08-28
- Reading time: 16 min
- Keywords: how to prepare for an AI voice interview, live AI phone screen questions, what does an AI interviewer score, AI interview follow up questions, AI voice interview tips

## Key takeaways

- The live AI voice screen runs short: documented ranges cluster from 2 to 20 minutes, so a short specific answer literally scores higher than a long vague one.
- The agent grades three things - whether the answer addresses the question, whether it gives a concrete example, and whether the reasoning is clear - not accent, speed, or vocabulary.
- Scores are won or lost on the fourth follow-up, not the first answer: memorized scripts produce clean openers but fail the probe that reconciles a claim with a constraint you mentioned earlier.
- Consistency is the integrity tell, not slowness: Fabric's analysis of 19,368 interviews found 38.5% triggered flags, and the signature pattern is a uniform 3 to 5 second delay after every question.
- A flag is reviewable, not fatal - integrity tools carry a stated 3 to 5% false-positive rate and credible vendors insist on no auto-reject, so an evidence-cited transcript gives a human reason to clear you.
- In Refolk's index, US recruiters tagging ATS skills outnumber the UK equivalent by roughly 49 to 1, so non-US candidates more often meet these agents cold.

You were invited to a live AI voice interview that replaces the recruiter phone screen, and it runs tomorrow. This is not the async one-way video interview and it is not a human call - it is a conversational agent that adapts its follow-ups in real time and grades your answers against a fixed rubric. This guide runs the format end to end, stage by stage, keyed to how these systems actually verify claims and score delivery, so you can walk in tomorrow and not tank your score.

## What a live AI voice screen is, and how it differs from the formats you already know

A live AI voice screen is a real-time conversation with a scoring agent: it asks a question, listens, and generates its next question from what you just said, then scores the whole thing against a competency rubric. It is not a video interview, where you record answers to fixed prompts with no reaction, and it is not a human phone screen, where a recruiter improvises and forgives.

Three properties define it, and each changes how you should behave:

- **It adapts.** Follow-ups are generated from the unfolding conversation, not read off a list. A vague answer earns a probe; a surprising claim earns a verification question.
- **It scores against a rubric.** Every completed screen produces a competency-by-competency score with evidence citations pulled from your transcript. The output is structured, not a recruiter's gut read.
- **It is short.** Documented screens run from roughly 2 minutes to 20 minutes. The format structurally rewards brevity, because the agent is grading relevance and one concrete example, not endurance.

The practical consequence: you cannot filibuster and you cannot charm. You have to hit the requirement with a specific example fast, then survive the probe.

**20 min - Upper end of documented AI voice screen length**

Mercor states most of its AI interviews take about 20 minutes; TalentSprout averages 6 minutes and targets under 10.

## How long it runs, and how much time to block

Block 30 minutes total: 5 to 10 for setup, and 5 to 20 for the screen itself. There is no single industry standard, but the documented ranges cluster tightly, so you can plan against them rather than guess.

The dossier's completion figures also tell you why this format exists. Voice screens complete far more often than video, so platforms keep them short and camera-optional to keep candidates from dropping. That low friction is not a signal to treat the screen casually - it is a 6-minute conversation that is fully scored and synced to the applicant tracking system.

| Format / source | Typical length | Completion rate |
| --- | --- | --- |
| Voice (HeyMilo) | not stated | ~70% |
| Voice/phone (Classet) | not stated | 80%+ |
| Video (per Classet benchmark) | not stated | 40 to 60% |
| Traditional phone screen (HeyMilo) | not stated | 30 to 50% |
| AI interview (Mercor) | ~20 min | not stated |
| AI interview (TalentSprout) | ~6 min | not stated |

The takeaway from the completion column: your edge is treating a short, low-friction screen as scorable. Most candidates read "6 minutes, no camera" as casual and answer loosely. The rubric does not.

> **Note:** The screen is short because completion is the platform's problem
>
> Voice completion (~70 to 80%) far exceeds video (40 to 60%), so platforms keep these screens brief and no-camera to stop drop-off. That design choice is why you get so little airtime to make each point.

## What the AI interviewer actually scores

The agent scores three things: whether your answer addresses the question asked, whether it gives a concrete example, and whether the reasoning is clear. It does not score your accent, your speaking speed, or your vocabulary. Reputable vendors dropped facial-expression trait scoring years ago; HireVue did so publicly in 2021.

The output that reaches a human is a structured evidence packet, not a rating out of ten. Knowing its contents tells you what to feed it.

| Scorecard element | What it holds | What you control |
| --- | --- | --- |
| Competency scores | Rating per competency, with rationale | Give one example per named requirement |
| Evidence citations | Snippets pulled from your transcript | Say specific, quotable facts |
| Executive summary | Strengths and follow-up areas | Land at least one clear strength early |
| Trust / integrity signals | Timing and consistency flags | Vary naturally; do not read scripts |

Exact rubric weights are set per role by the hiring company and are not published anywhere, so you cannot reverse-engineer them. What you can do is map the posting's named requirements to your examples. If the job lists "stakeholder management" and "SQL," walk in with one concrete stakeholder story and one concrete SQL story ready to deploy on cue.

#### What the agent grades, from the phrase you say to the packet it ships

1. **Evidence packet** - Competency scores, summary, and trust signals synced to the ATS
2. **Competency mapping** - Your example is tagged to a named role requirement
3. **Answer content** - Relevance, one concrete example, clear reasoning
4. **Spoken phrase** - The transcribed words the agent scores against

*Each answer you give is decomposed into rubric evidence, then assembled into a packet a recruiter reads.*

Because the agent grades relevance and a concrete example, a short specific answer scores higher than a long vague one. One situation, one action, one result. Rambling does not add points; it lowers your reasoning-clarity read and gives the probe more surface to poke.

> A modest claim you can defend beats an impressive claim that collapses on the fourth follow-up.

## The adaptive follow-up: where the score is actually decided

Follow-ups fire when an answer is vague, thin, or surprising, and the fourth follow-up is where scores are won or lost - not the first answer. This is the single most important thing to internalize, because it inverts how most people prepare. Candidates polish their opener and neglect the depth behind it. The agent does the opposite: it lets the opener pass, then digs.

Expect three probe types:

- **Contextual probing.** "Why did you choose that approach?" The agent tests whether your decision had a reason behind it.
- **Depth testing.** It pushes past the summary into the mechanics: who owned what, what the constraint was, what you traded off.
- **Consistency and verification checks.** When you overstate scope, a later follow-up asks you to reconcile your answer with a constraint you mentioned earlier. A flawless first answer followed by an inability to handle a basic follow-up is itself a flag.

The defense is to only claim what you can detail. If you say "I led the migration," the probe will ask who else was involved, what broke, and what you cut to hit the date. If you led it, you can answer. If you helped, say "I owned the data-validation piece" and defend that instead - a smaller, true claim survives the probe; an inflated one drops your competency score and can trip an integrity flag.

> **Rule:** Answer the probe with the four things it is fishing for
>
> When probed, add the owner (who did it), the constraint (what limited you), the tradeoff (what you gave up), and the result (what changed). Restating the claim louder is not an answer.

Prepare for the probe before you walk in. For each of your two or three core stories, write the four probe answers down so they are recall, not invention, under pressure.

**Probe-ready story card**

```
REQUIREMENT: (the skill from the posting this story proves)
SITUATION: (one sentence - where and when)
ACTION: (what you specifically did - "I", not "we")
RESULT: (one measurable or observable outcome)
OWNER: (what you owned vs. what others owned)
CONSTRAINT: (the deadline, budget, or limit you worked under)
TRADEOFF: (what you deprioritized to get the result)
RECONCILE: (if my scope sounds big, the honest limit is ___)
```

*Fill one card per core requirement in the posting. Keep answers to a sentence each.*

## Run the screen, stage by stage

Run these six stages in order. Setup and the practice run happen before the agent grades anything; the middle four are the scored conversation. Each stage has a clear "done" condition so you know when to move on.

#### The live AI voice screen, invite to completed conversation

1. **Set up the room and audio** - Open the test link on a device with a working mic, test audio, and record a clip you play back. Use a quiet room, headphones, and an external mic; confirm internet or switch to mobile data. Done when you hear yourself clearly.
2. **Take the practice run if offered** - If a free or unlimited practice interview is offered, take it and treat it as real to learn the cadence. Some platforms drop you straight into the scored screen. Done when you have heard your voice played back and know the flow.
3. **Open with a tight self-intro** - Give a role-anchored intro under a minute naming the two or three skills the job leans on. Skip the life story. Done when the intro is under a minute and specific to the posting.
4. **Answer each competency question with one concrete example** - The agent asks role-specific questions mapped to a rubric; answer each with a single concrete example, ideally STAR. Lead with situation and action, not a topic list. Done when each key requirement got one specific example.
5. **Handle adaptive follow-ups and verification probes** - Expect 'why', 'what tradeoff', and 'reconcile that with X'. Add the owner, constraint, tradeoff, and result the probe wants. Done when you supplied the detail rather than restating the claim.
6. **Close and let the output generate** - The agent signals completion and may ask for a satisfaction rating. A transcript, scorecard with evidence citations, summary, and trust score generate immediately and sync to the ATS. Done when the agent confirms the screen is finished.

A note on the practice step: sources disagree on whether it exists. Mercor offers an unlimited practice interview that does not affect your application, plus up to three retakes of the official interview. Other platforms have no practice step at all. Do not assume you get a warm-up. If one is offered, take it once, because the cadence of speaking to a machine that waits and then probes is genuinely different from a human call, and you want that surprise out of the way before the scored run.

Before you write your resume history into these stories, get the history straight. Refolk builds a candidate's resume from their own work record and tailors it to each posting, which is the same source material your voice-screen stories should come from - so the claims you make out loud match the claims on the page a recruiter reads next. Keep those consistent and the human review gate has nothing to snag on. See [Refolk](/candidates) for how that history assembles.

## How the screen goes wrong: delivery patterns that lower your score

Most of the damage in a voice screen is self-inflicted through delivery, not knowledge. The systems flag specific, documented patterns, and several of them are things a well-meaning, nervous candidate does by accident. Fabric's analysis of 19,368 interviews found 38.5% triggered cheating flags, rising to 48% in technical roles - so these patterns are common, and worth naming precisely.

| Failure mode | What it looks like | The fix |
| --- | --- | --- |
| Flatline timing | Same 3 to 5 sec pause after every question | Vary naturally; hard questions can take longer |
| Memorized script | "Three main points I'd highlight" repeated across answers | Sound conversational; allow self-correction |
| Keyword stuffing | Covers every keyword, stalls when probed | Lead with one concrete example, not a topic list |
| Overstated scope | "I led X," then fails the reconcile probe | Only claim what you can detail with owner and tradeoff |
| Skipped audio test | Agent mis-transcribes a garbled answer | Record and play back before you start |

Two of these deserve extra weight.

**Flatline timing is the overlay-tool signature.** A consistent 3 to 5 second delay after every question, regardless of difficulty, is the tell that an LLM overlay is feeding you answers. The important insight is that the flag keys on *consistency*, not slowness. Taking longer on a design question than on a "tell me about your last role" question is exactly the natural variance a real person produces. Do not try to answer instantly and do not pace yourself to a metronome. Think when you need to think.

**Second-screen reading trips gaze and timing signals on camera-enabled screens.** Off-axis eye movement plus a consistent delay reads as reading. This is also where false positives live: a nervous or neurodivergent candidate who looks away to think can get flagged for the same behavior. That false-positive risk is precisely why credible vendors insist on no auto-reject. If your screen is voice-only, gaze is not in play and integrity relies on timing and consistency instead - so the fix is the same either way: do not read a script, and vary your pace honestly.

The false positives matter for your peace of mind. The memorized-script flag can catch a genuinely polished speaker; the gaze flag can catch a thinker who looks up. You cannot eliminate the risk of a false flag, but you can make it survivable, which is the next section.

> **Watch out:** Do not use an answer-suggestion overlay
>
> Overlay tools produce the exact signature the systems hunt: uniform delay and clean first answers that stall under a specific probe. One vendor's own suggestion tool claims to start responding in ~0.8 seconds, but that speed is what gets caught. The downside of a flag outweighs any upside.

## The human gate, and why a flag is not fatal

Every completed screen is reviewed by a human before shortlisting - the AI does not make the hiring decision at credible vendors. A transcript, audio recording, competency scorecard, and integrity signals all sync to the applicant tracking system, and a recruiter reads them. This changes your strategy in one important way: your job is not only to score well, but to leave a transcript that gives a reviewer a reason to clear you if a signal fires.

Integrity tools carry a stated 3 to 5% false-positive rate, and best practice across vendors is explicit: no auto-reject. So a flag routes to a human, and that human is reading your transcript. If your answers are specific and evidence-cited - real names, real numbers, real tradeoffs - the reviewer sees a candidate who clearly did the work, and the flag looks like noise. If your answers are thin and generic, a flag has nothing to argue against.

| Segment | Flag rate |
| --- | --- |
| All interviews | 38.5% |
| Technical roles | 48% |
| Stated false-positive rate | 3 to 5% |

Read those numbers together. Nearly two in five interviews flag *something*, but the stated false-positive rate is a fraction of that - meaning most flags are attached to real overlay use, and a clean candidate who simply thought carefully is the minority the review gate exists to protect. Give the reviewer the evidence and you are on the right side of that gate.

#### From completed screen to shortlist

| Stage | Figure | Note |
| --- | --- | --- |
| Completed screens | 100% | Every finished screen generates a full evidence packet |
| Flagged for review | 38.5% | Fabric's rate across 19,368 interviews |
| True-positive share | most of the flagged | Stated false-positive rate is only 3 to 5% |
| Cleared by human gate | reviewer's call | No auto-reject; evidence-cited transcripts clear |

*The agent scores and flags, but a human reads the transcript before anyone advances.*

## Why non-US candidates meet these agents cold, and what to do about it

If you are interviewing outside the US, you are more likely to meet a voice agent with no prior exposure and less standardized guidance - because the buyer market for these tools is heavily concentrated in the US. In Refolk's index of professional profiles, US recruiters and technical recruiters tagging applicant-tracking-system skills outnumber their counterparts abroad by a wide margin.

| Country | Recruiters with ATS skills [OURS] | Ratio vs US |
| --- | --- | --- |
| United States | 6,430 | 1.0x |
| United Kingdom | 131 | 0.020x (≈49x smaller) |
| Germany | 33 | 0.005x (≈195x smaller) |

Those are Refolk's own index figures, published nowhere else. The practical read: process maturity around these agents is far higher in the US, so US candidates often arrive having done one before, while UK and German candidates more often face the format cold. If that is you, the practice run in step two is not optional polish - it is how you close the exposure gap. Take any practice interview offered, and if none is offered, run a browser-based voice mock the night before so the cadence is not new tomorrow.

You can also see who is actually deploying these screens, so you know what to expect from a given employer.

Ask me this: `Talent acquisition leaders at US companies who have rolled out AI voice or phone screening in the last year.` - [run the search](https://www.refolk.ai/start?q=Talent%20acquisition%20leaders%20at%20US%20companies%20who%20have%20rolled%20out%20AI%20voice%20or%20phone%20screening%20in%20the%20last%20year.).

*Returns the people standing up these programs, so you can gauge how established the process is at a target employer.*

## Before you start: the pre-flight checklist

Run this checklist in the ten minutes before you join. It is ordered so that the things that silently ruin a score - a bad mic, an unread posting - are caught before the agent hears a word.

#### Pre-flight, ten minutes before the screen

- [ ] I tested the link on a device with a working microphone and heard myself on a recorded playback.
- [ ] I am in a quiet room with headphones and, if I have one, an external mic; captions are on if offered.
- [ ] My internet is confirmed, with mobile data ready as a fallback.
- [ ] I took the practice interview if one was offered, and I know the question-answer cadence.
- [ ] I have two or three probe-ready story cards, each mapped to a named requirement in the posting.
- [ ] My self-intro is under a minute and names the skills the job leans on.
- [ ] I am not running any answer-suggestion overlay, and I will vary my pacing naturally.
- [ ] For each core claim, I can supply owner, constraint, tradeoff, and result on request.

## Keeping this current

These systems change, and the guide's spine is the mechanism, not any single vendor's rules. Two things are set per employer and will not appear in public docs: the exact rubric weights, and whether a practice run exists. Re-check both against the specific invite you receive - read the confirmation email and the platform's candidate help page for length, retakes, and practice availability before you assume anything from this guide applies verbatim.

Two dossier facts are worth watching over time. First, integrity detection evolves; the 3 to 5 second overlay signature and the 3 to 5% false-positive rate are current documented values, so if you are interviewing repeatedly, re-read the platform's candidate docs each time rather than trusting last quarter's behavior. Second, the completion and length ranges (2 to 20 minutes, ~70 to 80% voice completion) are the current cluster, not a law - if a screen you are invited to states a different length, believe the invite.

What does not change is the winning move. Bring real examples you can defend to the fourth follow-up, keep each answer short and specific, vary your pace like a human, and leave a transcript that gives the human reviewer a reason to advance you. Do that, and the format works in your favor: it is short, it does not judge your accent, and a modest true claim beats an impressive false one every time.

## Frequently asked questions

### How do I prepare for an AI voice interview in one evening?

Block 30 minutes tonight and do three things. First, test your link, mic, and audio playback so nothing garbles your first answer. Second, pull two or three concrete stories - one situation, one action, one result each - mapped to the skills in the posting. Third, if a practice interview is offered, take it once to learn the cadence. Mercor allows unlimited practice runs that do not affect your application, so use them.

### What does an AI interviewer actually score?

Credible agents score whether your answer addresses the question, gives a concrete example, and reasons clearly - not your accent, speed, or vocabulary. The output is a competency-by-competency scorecard with ratings, rationale, evidence citations pulled from the transcript, an executive summary of strengths and follow-up areas, and trust or integrity signals. Exact rubric weights are set per role by the hiring company and are not published, so prepare against the posting's named requirements.

### Why does the AI keep asking follow-up questions?

Follow-ups fire when an answer is vague, thin, or surprising. The agent probes for depth ('why did you choose that approach?'), tests reasoning, and runs consistency checks. A common verification probe asks you to reconcile a claim with a constraint you mentioned earlier. This is where scores are won or lost: a clean first answer that stalls under the probe reads worse than a modest claim you can detail with owner, constraint, tradeoff, and result.

### Can an AI voice interview flag me for cheating unfairly?

It can, but a flag is reviewable rather than fatal. Integrity tools carry a stated 3 to 5% false-positive rate, and credible vendors insist on a human review gate with no auto-reject. Nervous or neurodivergent candidates who look away to think can trip gaze-based signals, which is exactly why the human gate exists. Your defense is a transcript full of specific, evidence-cited answers that give a reviewer reason to clear you.

### Should I use a tool that feeds me answers during the screen?

No. Overlay tools produce a signature the systems detect: a consistent 3 to 5 second delay after every question regardless of difficulty, and clean first answers that stall under a specific probe. Fabric's analysis of 19,368 interviews found 38.5% triggered cheating flags, rising to 48% in technical roles. A modest real example you can defend on the fourth follow-up beats a manicured answer that collapses when probed.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/candidates/guides/live-ai-voice-screen-playbook*
