The Conversational AI Voice Screen, From Invite Link to a Finished Transcript
You will run a live conversational AI voice screen end to end: verify the invite, set up clean audio, and answer adaptive probes so the transcript advances you.
You were invited to an automated interview that talks back. It listens, asks a live follow-up when your answer is thin, and scores a running transcript rather than a recorded clip. This guide is for the candidate who has to get through that screen to reach a human reviewer, and it runs the job in order: verify the invite is real, set up audio that transcribes cleanly, and answer so the transcript scores high enough to advance.
The library already covers the one-way video interview, where questions are fixed and you record into a camera with no interaction. This is the other case. A conversational agent probes you in real time, and the single thing that makes it harder is that you usually cannot re-record. Everything below is built around that constraint.
What a live AI voice screen actually is
A live AI voice screen is a 10 to 20 minute phone or browser conversation with a conversational agent that chooses its next question from what you just said. It resembles a recruiter phone screen, except nobody is on the line and the output is a scored transcript, not a human's notes.
Two properties separate it from the pre-recorded one-way video interview. First, it is adaptive: the system probes deeper when an answer is vague and skips questions you have already addressed. Second, it is scored on a transcript. On most rubric-driven platforms the AI pairs each question with your response, assesses it against defined criteria, and produces an overall score that a recruiter reviews in 60 to 90 seconds instead of sitting through a 15 to 30 minute live call.
That review speed is the reason these screens exist, and it is the reason your transcript has to carry the whole case on its own. A number helps frame the stakes.
Duration and question count, by source
Plan for 10 to 20 minutes and four strong stories. Sources converge on the window but differ on how they cut it into questions, so read your own invite for the exact count and treat the table below as the range you fall inside.
| Source | Minutes | Questions |
|---|---|---|
| TestGorilla | 10-20 | not stated |
| HeyMilo | 10-20 | not stated |
| Four-Leaf (mock) | 10 / 20 / 30 | 3 / 4 / 5 |
| Koji (research) | 12-18 | up to 12-15 |
| Coril (mock) | 7-10 / 12-18 | 5 / 6-8 |
Each question is designed to take about two minutes, and one fully conversational configuration allows up to three follow-up probes per open-ended question. So a screen advertised as four questions can easily become ten exchanges once the agent starts probing. Prepare for the probes, not just the headline count.
Who you are really facing, and why the category has no name
You are facing a named tool, not a generic "AI interview." That distinction matters because the category has not settled on common vocabulary yet, which changes how you research what you will face.
In Refolk's index of professional profiles, keyword searches combining "AI interview" and "voice screening" in recruiter headlines returned zero matches in both the US and the UK. The same index shows 433 US recruiting and talent-acquisition professionals listing HireVue as a skill. The people who run these screens name them by product, not by the generic phrase a candidate would type into a search box.
The practical consequence: to learn what your screen will do, search the employer's hiring stack by tool name. Find the platform in the invite, then read that vendor's own candidate documentation for retake rules and what it scores. Refolk can trace which recruiters at a given employer list a specific interview tool as a skill, which is a faster route to the real answer than reading generic prep posts.
Where these screens cluster
Adoption is heavily US-concentrated. If you are applying from the UK, you will most often meet an AI voice screen at a US-headquartered employer.
| Market | Profiles listing HireVue | Share of US total (derived) |
|---|---|---|
| United States | 433 | 100% (baseline) |
| United Kingdom | 51 | 11.8% (derived) |
| US-to-UK multiple | - | 8.49x (derived) |
In Refolk's index, HireVue-skilled recruiters are about 8.5x more common in the US than in the UK. Among those US profiles, the top employers are EY and Deloitte with three each, with Bloomberg, Citizens and Take-Two Interactive also present. If your invite comes from a large professional-services or enterprise employer, assume a tool-based screen is in play and research accordingly.
Verify the invite before you say a word
Verify the employer independently before you speak, because a convincing website proves nothing. The one rule that survives every scam variant is the FTC's: confirm you are under consideration using a phone number you know is legitimate, not one you got from the person who approached you.
Run three checks, in this order:
- Careers page. Go to the company's official website and find the role in the Careers section. If the role is not there, treat the posting as unverified.
- Email domain. A genuine company contacts you from its own domain. A free service like Gmail, Yahoo or Hotmail is a red flag.
- Independent call. Phone the company on a number you sourced yourself and confirm both that the role is real and that the person who contacted you is affiliated with the company.
Never supply sensitive data before a signed offer. Requests for your Social Security number, passport, bank details or date of birth before a formal interview and a legitimate offer are identity harvest. Legitimate employers collect those during onboarding, after you have signed. If something is off, report it at ReportFraud.ftc.gov.
Set up audio that transcribes cleanly
On a transcript-scored screen, clean capture outranks answer polish, because the model reads text and the text comes from your audio. Background hiss alone can cut transcription accuracy by up to ten percentage points, and a garbled proper noun is a lost scoring signal.
The fixes are cheap and physical. Record in a room with soft furnishings to kill echo, position the mic about 15 cm from your mouth, and keep a steady pace. A roughly $70 USB cardioid dynamic mic, which picks up mostly what is in front of it, often beats a laptop's omnidirectional pickup in a noisy room. If anyone else is nearby, make sure only one person speaks at a time, because overlapping voices degrade the transcript.
The ceiling is high when the audio is clean. The best speech-to-text exceeds 97% accuracy on clean audio and above 92% is now common, with one current model benchmarked at a 5.6% mean word error rate. Your job is to not throw accuracy away on reverb and hiss.
What the score is built on, bottom to top
- Overall scoreRubric assessment summed across questions
- Rubric matchDid each answer fill the criterion fields
- Transcript textThe words the model actually read
- Audio captureMic, distance, room, one speaker at a time
Run a test transcription before the real call. Read a short answer that includes the proper nouns you expect to use, company names, product names, numbers, and confirm they come back correct. If they garble, move the mic, change rooms, or switch microphones until they do not.
How the adaptive follow-ups work
The follow-up is a scoring signal, not a trap. Adaptive systems probe deeper when an answer is vague, so a follow-up is the model telling you a rubric field is still empty. The single most common probe is a request for a specific example, and it means your first answer lacked detail.
Because the system listens to your answer and picks its next question from it, it will return to anything you left vague. That makes the structure of your answer, not its wording, the thing that survives probing. Use STAR with a heavy Action section: lead with a short Situation, state the Task briefly, spend most of the answer on first-person Action, and close on a Result with a number.
STAR time budget for a 90-second probed answer
| Part | Share | Seconds (derived from 90s) |
|---|---|---|
| Situation | 10-15% | 9-14 |
| Task | 10% | 9 |
| Action | 50-60% | 45-54 |
| Result | 20-25% | 18-23 |
Do not script. A scripted monologue sounds rehearsed and collapses on the second probe, because the model asks about the thing you did not plan for. Practice the structure and pre-stage branches instead: for each story, be ready to say why you chose that approach, what alternative you rejected, who disagreed, and how you confirmed the result. Those are the three questions the agent asks when it probes.
The probe loop
- You answerDeliver a complete STAR unit, capped near 90 seconds
- Model scores silentlyChecks which criteria your answer actually filled
- Probe fires"Can you give a specific example?" means a field is thin
- You add a specificOne named example or number, not an apology
- Model advancesFills the field and moves on, or skips the redundant question
A follow-up is not a verdict on your answer, it is the model handing you the rubric field you forgot to fill.
The procedure, invite link to finished transcript
Run the screen in nine stages. Earlier stages are preparation you do once; the launch and answer stages are the live call itself.
Running the screen end to end
- Verify the invite independentlyCheck the role on the official careers page, confirm the sender's email domain, and call the company on a number you found yourself. Done when the employer independently confirms you are under consideration.
- Decode the format from the inviteRead the question count, timing, retake rules, deadline, and device requirements. If you get no answer, assume no retakes and that a human may also review.
- Set up audio and environmentUse a wired or USB cardioid mic about 15 cm away in a quiet, soft-furnished room, then run a test transcription. Done when your proper nouns transcribe correctly.
- Prepare three or four real stories in STARBuild Action-heavy stories and pre-stage follow-up branches: the alternative rejected, who disagreed, how you verified. Done when you can be probed on any detail without a script.
- Launch and confirm it is the AIWatch for instant re-prompting, no social acknowledgment, verbatim follow-ups, and clean consistent audio. Done when you know the format you are in.
- Answer in 60 to 90 second unitsDeliver each answer as a complete Situation, Task, Action, Action, Result, capped near 90 seconds. The interviewer moves on when you ramble.
- Handle follow-ups by adding specificsTreat a request for a specific example as a cue to add one concrete detail, not as a failure. Done when the probe resolves with a named example.
- Recover from a stumble without a retakeFinish the answer, then explicitly summarize the key action you took so the transcript ends clean. Done when the last words are a clear summary.
- Close and request a redo only on genuine failureClick Finish, or email the recruiter only if a real technical failure cut the call off. A dropped call is valid, disliking your answers is not.
Confirming it is the AI, not a person
At launch, four signals tell you the format: instant re-prompting the moment you stop talking, no social acknowledgment of what you said, exact verbatim follow-ups, and consistent audio with no ambient noise. Once you have confirmed it is the agent, stop performing for a human and start filling rubric fields.
Why the no-retake rule changes your tactics
The retake asymmetry is the core risk versus one-way video. The pre-recorded video interview often lets you re-record per question, so a weak first take is recoverable. The live voice screen captures one continuous conversation with no mid-session retake, so there is no take to throw away and redo.
| Screen type | Retake behavior | Recovery tactic |
|---|---|---|
| One-way video | Sometimes one retake per question | Re-record the weak answer |
| Live phone/voice | Usually none | Summarize-to-close inside the same answer |
| Mercor (exception) | Up to 3 attempts, latest evaluated | Still answer as if it is the only take |
Because you cannot re-record, the recovery move lives inside the answer. If you fumble, do not apologize and restart. Complete the thought, then explicitly summarize: state the key action you took in one clean line. The transcript then ends on clarity, which is what gets read.
To summarize, the key action I took was rebuilding the reconciliation process, and the result was cutting month-end close from nine days to four.
Say this after a fumble, before the agent advances, so the transcript's last words are clean. Swap in your real action.
How this goes wrong
Most failures on an AI voice screen are not weak content. They are predictable mistakes about how the format works, and each one has a tell you can check for.
- Treating a probe as rejection. You hear a follow-up and backpedal or apologize. The probe usually means a detail was missing, so it is a gift, not a verdict. Add a specific example instead of defending the first answer.
- Scripting your answers. A fluent memorized monologue that collapses the moment the agent probes a detail you did not plan for. Practice the STAR structure and the branch questions, never the exact words.
- Assuming retakes exist. You deliberately give a weak first answer expecting to redo it. If the retake policy is unconfirmed, assume there are none and treat every answer as final.
- Website-only verification. A professional-looking site convinces you the employer is real. Scammers build a phony online presence, so the site proves nothing. Call a number you sourced yourself.
- Laptop mic in a hard-surfaced room. It sounds fine to you and garbles proper nouns in the transcript, where echo and reverb blur consonants and word boundaries. Run a test transcription first.
- Rambling past the cap. You think more detail scores higher, but the interviewer moves on when you ramble and can truncate before your Result lands. Cap answers near 90 seconds and make sure the Result is in.
- Over-trusting that delivery is or is not scored. You optimize vocal tone for a rubric-only platform, or ignore delivery on one that scores it. Sources conflict, so confirm the specific platform named in your invite.
That last one deserves its own judgement call, because the two assumptions pull your preparation in opposite directions.
Should you optimize delivery, or only content?
Greenhouse Voice AI, which sits inside an ISO/IEC 42001 audit scope, explicitly excludes accent, tone, fluency, pace, stutters, pauses and confidence from scoring. Some candidate-prep tools instead sell vocal scorecards for clarity and persuasiveness. You cannot know which world you are in without naming the platform, which is why that step sits at the front of the procedure.
Research the platform and the people behind it
Before the call, learn what your specific platform does and who runs it at the employer. The fastest route is to find the tool in your invite, then confirm its retake and scoring rules from the vendor's own candidate documentation rather than averaging across generic advice.
Because recruiters name these screens by tool, searching for the people who deploy a named platform tells you more than searching the generic category. Looking up who at a target employer lists a specific interview tool as a skill shows you the real stack you will face.
Verify before you click Finish
Run this check before you end the session. Each item is something you can confirm, not a topic to think about.
Pre-finish check
- I confirmed the role exists on the official careers page and the sender used the company domain.
- I called a number I sourced myself and the employer confirmed I am under consideration.
- I ran a test transcription and my proper nouns came back correct.
- I know whether this platform allows retakes, and I assumed none where it was unconfirmed.
- I identified whether my named platform scores delivery or only transcript content.
- Each answer landed a Result with a number before the agent advanced.
- Every probe for a specific example was resolved with a concrete, named detail.
- Any fumble ended with a one-line summary of the key action.
- I shared no SSN, passport, bank, or date-of-birth details.
Keeping this current
This category changes by tool, not by rule, so the maintenance work is re-checking the named platform each time. Two facts are stable enough to rely on: most phone-based voice screens give no mid-session retake, and the most common probe is a request for a specific example. Build your habits on those.
Everything time-sensitive lives in the vendor's own documentation. Before each new screen, re-confirm three things from the invite and the platform's candidate help pages: the retake policy, whether delivery is scored, and the question count. When the invite names a tool you have not seen, search the employer's recruiting staff by that tool name to learn what they actually use, then read the vendor docs for that product. The procedure above stays the same. Only the platform-specific answers move, and you re-verify them one screen at a time.
Questions job seekers ask
How long is an AI voice interview and how many questions will it ask?
Most run 10 to 20 minutes, with each question designed to take about two minutes. Vendor and research sources map length to roughly 3 questions at 10 minutes, 4 at 20, and 5 at 30, with methodology guidance capping sessions at 12 to 15 questions. Read your invite for the exact count, and if it is not stated, plan for four strong stories that each survive a couple of follow-up probes.
Can I retake an AI phone screen if I stumble?
Usually not. Most phone-based live voice screens capture one continuous conversation with no mid-session retake, unlike one-way video where some platforms allow one retake per question. A few systems are exceptions, for example Mercor documents up to three attempts across related applications with only the most recent evaluated, but assume no retakes unless the invite says otherwise. If a genuine technical failure interrupts the call, email the recruiter for a fresh attempt.
Does the AI score how I sound, or only what I say?
It depends on the platform, and the sources genuinely disagree. Enterprise systems like Greenhouse Voice AI explicitly exclude accent, tone, fluency, pace, stutters and confidence, scoring only transcript content against a rubric. Some candidate-prep tools instead sell vocal scorecards for clarity and persuasiveness. Confirm the specific platform named in your invite rather than optimizing for both, because the two assumptions pull your preparation in opposite directions.
How do I know if an AI voice interview invite is a scam?
Verify the employer independently before you speak. Check the role on the company's official careers page, confirm the sender uses the company's own email domain rather than a free service, and call a phone number you sourced yourself, not one the recruiter gave you. Never share your SSN, passport, bank details or date of birth before a signed offer, since that is identity harvest. Report suspected scams at ReportFraud.ftc.gov.
What is the best answer structure for an AI interviewer's follow-up questions?
Use STAR with a heavy Action section: Situation 10 to 15 percent, Task 10 percent, Action 50 to 60 percent in the first person, and Result 20 to 25 percent with a number. Practice the structure, not a word-for-word script, because scripted answers collapse on the second probe. Pre-stage branches for the common follow-ups: the alternative you rejected, who disagreed, and how you verified the result.
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.