# The Conversational AI Voice Screen, From Invite Link to a Finished Transcript

*You will run a live conversational AI voice screen end to end: verify the invite, set up clean audio, and answer adaptive probes so the transcript advances you.*

- Canonical URL: https://www.refolk.ai/candidates/guides/conversational-ai-voice-screen
- Pillar: Interviewing
- Format: Playbook
- Published: 2026-10-02
- Last reviewed: 2026-10-02
- Reading time: 17 min

You were invited to an automated interview that talks back. It listens, asks a live follow-up when your answer is thin, and scores a running transcript rather than a recorded clip. This guide is for the candidate who has to get through that screen to reach a human reviewer, and it runs the job in order: verify the invite is real, set up audio that transcribes cleanly, and answer so the transcript scores high enough to advance.

The library already covers the one-way video interview, where questions are fixed and you record into a camera with no interaction. This is the other case. A conversational agent probes you in real time, and the single thing that makes it harder is that you usually cannot re-record. Everything below is built around that constraint.

## What a live AI voice screen actually is

A live AI voice screen is a 10 to 20 minute phone or browser conversation with a conversational agent that chooses its next question from what you just said. It resembles a recruiter phone screen, except nobody is on the line and the output is a scored transcript, not a human's notes.

Two properties separate it from the pre-recorded one-way video interview. First, it is adaptive: the system probes deeper when an answer is vague and skips questions you have already addressed. Second, it is scored on a transcript. On most rubric-driven platforms the AI pairs each question with your response, assesses it against defined criteria, and produces an overall score that a recruiter reviews in 60 to 90 seconds instead of sitting through a 15 to 30 minute live call.

That review speed is the reason these screens exist, and it is the reason your transcript has to carry the whole case on its own. A number helps frame the stakes.

**60-90s - How long a recruiter spends on your scored transcript**

The AI compresses a 15 to 30 minute live screen into a transcript a human skims in under two minutes, so the text must stand alone.

### Duration and question count, by source

Plan for 10 to 20 minutes and four strong stories. Sources converge on the window but differ on how they cut it into questions, so read your own invite for the exact count and treat the table below as the range you fall inside.

| Source | Minutes | Questions |
| --- | --- | --- |
| TestGorilla | 10-20 | not stated |
| HeyMilo | 10-20 | not stated |
| Four-Leaf (mock) | 10 / 20 / 30 | 3 / 4 / 5 |
| Koji (research) | 12-18 | up to 12-15 |
| Coril (mock) | 7-10 / 12-18 | 5 / 6-8 |

Each question is designed to take about two minutes, and one fully conversational configuration allows up to three follow-up probes per open-ended question. So a screen advertised as four questions can easily become ten exchanges once the agent starts probing. Prepare for the probes, not just the headline count.

## Who you are really facing, and why the category has no name

You are facing a named tool, not a generic "AI interview." That distinction matters because the category has not settled on common vocabulary yet, which changes how you research what you will face.

In Refolk's index of professional profiles, keyword searches combining "AI interview" and "voice screening" in recruiter headlines returned zero matches in both the US and the UK. The same index shows 433 US recruiting and talent-acquisition professionals listing HireVue as a skill. The people who run these screens name them by product, not by the generic phrase a candidate would type into a search box.

The practical consequence: to learn what your screen will do, search the employer's hiring stack by tool name. Find the platform in the invite, then read that vendor's own candidate documentation for retake rules and what it scores. [Refolk](/candidates) can trace which recruiters at a given employer list a specific interview tool as a skill, which is a faster route to the real answer than reading generic prep posts.

### Where these screens cluster

Adoption is heavily US-concentrated. If you are applying from the UK, you will most often meet an AI voice screen at a US-headquartered employer.

| Market | Profiles listing HireVue | Share of US total (derived) |
| --- | --- | --- |
| United States | 433 | 100% (baseline) |
| United Kingdom | 51 | 11.8% (derived) |
| US-to-UK multiple | - | 8.49x (derived) |

In Refolk's index, HireVue-skilled recruiters are about 8.5x more common in the US than in the UK. Among those US profiles, the top employers are EY and Deloitte with three each, with Bloomberg, Citizens and Take-Two Interactive also present. If your invite comes from a large professional-services or enterprise employer, assume a tool-based screen is in play and research accordingly.

> **Note:** Name the tool before you prepare
>
> Because the category is named by product, not by phrase, find the platform in your invite first. Retake rules, timing, and what gets scored all differ by vendor, and generic advice averages over those differences.

## Verify the invite before you say a word

Verify the employer independently before you speak, because a convincing website proves nothing. The one rule that survives every scam variant is the FTC's: confirm you are under consideration using a phone number you know is legitimate, not one you got from the person who approached you.

Run three checks, in this order:

1. **Careers page.** Go to the company's official website and find the role in the Careers section. If the role is not there, treat the posting as unverified.
2. **Email domain.** A genuine company contacts you from its own domain. A free service like Gmail, Yahoo or Hotmail is a red flag.
3. **Independent call.** Phone the company on a number you sourced yourself and confirm both that the role is real and that the person who contacted you is affiliated with the company.

Never supply sensitive data before a signed offer. Requests for your Social Security number, passport, bank details or date of birth before a formal interview and a legitimate offer are identity harvest. Legitimate employers collect those during onboarding, after you have signed. If something is off, report it at ReportFraud.ftc.gov.

> **Rule:** Call a number you found yourself
>
> Do not rely on the existence of a website or a number inside the invite email. Scammers build a phony online presence and supply their own callback line. Independent verification means a number you located on the official site.

## Set up audio that transcribes cleanly

On a transcript-scored screen, clean capture outranks answer polish, because the model reads text and the text comes from your audio. Background hiss alone can cut transcription accuracy by up to ten percentage points, and a garbled proper noun is a lost scoring signal.

The fixes are cheap and physical. Record in a room with soft furnishings to kill echo, position the mic about 15 cm from your mouth, and keep a steady pace. A roughly $70 USB cardioid dynamic mic, which picks up mostly what is in front of it, often beats a laptop's omnidirectional pickup in a noisy room. If anyone else is nearby, make sure only one person speaks at a time, because overlapping voices degrade the transcript.

The ceiling is high when the audio is clean. The best speech-to-text exceeds 97% accuracy on clean audio and above 92% is now common, with one current model benchmarked at a 5.6% mean word error rate. Your job is to not throw accuracy away on reverb and hiss.

#### What the score is built on, bottom to top

1. **Overall score** - Rubric assessment summed across questions
2. **Rubric match** - Did each answer fill the criterion fields
3. **Transcript text** - The words the model actually read
4. **Audio capture** - Mic, distance, room, one speaker at a time

*Each layer depends on the one beneath it, which is why audio setup protects more marginal score than another rehearsal hour.*

Run a test transcription before the real call. Read a short answer that includes the proper nouns you expect to use, company names, product names, numbers, and confirm they come back correct. If they garble, move the mic, change rooms, or switch microphones until they do not.

> **Watch out:** A laptop mic in a hard room sounds fine to you
>
> It will sound clean in your own ears and still blur consonants and word boundaries in the transcript, mangling exactly the proper nouns that carry your specifics. The only reliable check is a test transcription, not how it sounds live.

## How the adaptive follow-ups work

The follow-up is a scoring signal, not a trap. Adaptive systems probe deeper when an answer is vague, so a follow-up is the model telling you a rubric field is still empty. The single most common probe is a request for a specific example, and it means your first answer lacked detail.

Because the system listens to your answer and picks its next question from it, it will return to anything you left vague. That makes the structure of your answer, not its wording, the thing that survives probing. Use STAR with a heavy Action section: lead with a short Situation, state the Task briefly, spend most of the answer on first-person Action, and close on a Result with a number.

### STAR time budget for a 90-second probed answer

| Part | Share | Seconds (derived from 90s) |
| --- | --- | --- |
| Situation | 10-15% | 9-14 |
| Task | 10% | 9 |
| Action | 50-60% | 45-54 |
| Result | 20-25% | 18-23 |

Do not script. A scripted monologue sounds rehearsed and collapses on the second probe, because the model asks about the thing you did not plan for. Practice the structure and pre-stage branches instead: for each story, be ready to say why you chose that approach, what alternative you rejected, who disagreed, and how you confirmed the result. Those are the three questions the agent asks when it probes.

#### The probe loop

1. **You answer** - Deliver a complete STAR unit, capped near 90 seconds
2. **Model scores silently** - Checks which criteria your answer actually filled
3. **Probe fires** - "Can you give a specific example?" means a field is thin
4. **You add a specific** - One named example or number, not an apology
5. **Model advances** - Fills the field and moves on, or skips the redundant question

*Every probe is the model refilling an empty rubric field, so treat each one as a prompt to add one concrete detail.*

> A follow-up is not a verdict on your answer, it is the model handing you the rubric field you forgot to fill.

## The procedure, invite link to finished transcript

Run the screen in nine stages. Earlier stages are preparation you do once; the launch and answer stages are the live call itself.

#### Running the screen end to end

1. **Verify the invite independently** - Check the role on the official careers page, confirm the sender's email domain, and call the company on a number you found yourself. Done when the employer independently confirms you are under consideration.
2. **Decode the format from the invite** - Read the question count, timing, retake rules, deadline, and device requirements. If you get no answer, assume no retakes and that a human may also review.
3. **Set up audio and environment** - Use a wired or USB cardioid mic about 15 cm away in a quiet, soft-furnished room, then run a test transcription. Done when your proper nouns transcribe correctly.
4. **Prepare three or four real stories in STAR** - Build Action-heavy stories and pre-stage follow-up branches: the alternative rejected, who disagreed, how you verified. Done when you can be probed on any detail without a script.
5. **Launch and confirm it is the AI** - Watch for instant re-prompting, no social acknowledgment, verbatim follow-ups, and clean consistent audio. Done when you know the format you are in.
6. **Answer in 60 to 90 second units** - Deliver each answer as a complete Situation, Task, Action, Action, Result, capped near 90 seconds. The interviewer moves on when you ramble.
7. **Handle follow-ups by adding specifics** - Treat a request for a specific example as a cue to add one concrete detail, not as a failure. Done when the probe resolves with a named example.
8. **Recover from a stumble without a retake** - Finish the answer, then explicitly summarize the key action you took so the transcript ends clean. Done when the last words are a clear summary.
9. **Close and request a redo only on genuine failure** - Click Finish, or email the recruiter only if a real technical failure cut the call off. A dropped call is valid, disliking your answers is not.

### Confirming it is the AI, not a person

At launch, four signals tell you the format: instant re-prompting the moment you stop talking, no social acknowledgment of what you said, exact verbatim follow-ups, and consistent audio with no ambient noise. Once you have confirmed it is the agent, stop performing for a human and start filling rubric fields.

### Why the no-retake rule changes your tactics

The retake asymmetry is the core risk versus one-way video. The pre-recorded video interview often lets you re-record per question, so a weak first take is recoverable. The live voice screen captures one continuous conversation with no mid-session retake, so there is no take to throw away and redo.

| Screen type | Retake behavior | Recovery tactic |
| --- | --- | --- |
| One-way video | Sometimes one retake per question | Re-record the weak answer |
| Live phone/voice | Usually none | Summarize-to-close inside the same answer |
| Mercor (exception) | Up to 3 attempts, latest evaluated | Still answer as if it is the only take |

Because you cannot re-record, the recovery move lives inside the answer. If you fumble, do not apologize and restart. Complete the thought, then explicitly summarize: state the key action you took in one clean line. The transcript then ends on clarity, which is what gets read.

**Summarize-to-close recovery line**

```
To summarize, the key action I took was rebuilding the reconciliation process, and the result was cutting month-end close from nine days to four.
```

*Say this after a fumble, before the agent advances, so the transcript's last words are clean. Swap in your real action.*

## How this goes wrong

Most failures on an AI voice screen are not weak content. They are predictable mistakes about how the format works, and each one has a tell you can check for.

- **Treating a probe as rejection.** You hear a follow-up and backpedal or apologize. The probe usually means a detail was missing, so it is a gift, not a verdict. Add a specific example instead of defending the first answer.
- **Scripting your answers.** A fluent memorized monologue that collapses the moment the agent probes a detail you did not plan for. Practice the STAR structure and the branch questions, never the exact words.
- **Assuming retakes exist.** You deliberately give a weak first answer expecting to redo it. If the retake policy is unconfirmed, assume there are none and treat every answer as final.
- **Website-only verification.** A professional-looking site convinces you the employer is real. Scammers build a phony online presence, so the site proves nothing. Call a number you sourced yourself.
- **Laptop mic in a hard-surfaced room.** It sounds fine to you and garbles proper nouns in the transcript, where echo and reverb blur consonants and word boundaries. Run a test transcription first.
- **Rambling past the cap.** You think more detail scores higher, but the interviewer moves on when you ramble and can truncate before your Result lands. Cap answers near 90 seconds and make sure the Result is in.
- **Over-trusting that delivery is or is not scored.** You optimize vocal tone for a rubric-only platform, or ignore delivery on one that scores it. Sources conflict, so confirm the specific platform named in your invite.

That last one deserves its own judgement call, because the two assumptions pull your preparation in opposite directions.

#### Should you optimize delivery, or only content?

Horizontal axis runs from Platform excludes delivery to Platform scores delivery. Vertical axis runs from You practiced content only to You practiced content plus vocal tone.

| Quadrant | What it means |
| --- | --- |
| Content prep, delivery-blind platform | Ideal fit, spend all effort on rubric fields and audio |
| Content prep, delivery-scored platform | Underprepared, add pace and clarity practice |
| Content plus vocal prep, delivery-blind platform | Wasted effort on tone, redirect to specifics |
| Content plus vocal prep, delivery-scored platform | Ideal fit, both dimensions covered |

*The right preparation depends on whether your named platform scores how you sound, and the sources genuinely disagree.*

Greenhouse Voice AI, which sits inside an ISO/IEC 42001 audit scope, explicitly excludes accent, tone, fluency, pace, stutters, pauses and confidence from scoring. Some candidate-prep tools instead sell vocal scorecards for clarity and persuasiveness. You cannot know which world you are in without naming the platform, which is why that step sits at the front of the procedure.

## Research the platform and the people behind it

Before the call, learn what your specific platform does and who runs it at the employer. The fastest route is to find the tool in your invite, then confirm its retake and scoring rules from the vendor's own candidate documentation rather than averaging across generic advice.

Because recruiters name these screens by tool, searching for the people who deploy a named platform tells you more than searching the generic category. Looking up who at a target employer lists a specific interview tool as a skill shows you the real stack you will face.

Ask me this: `Talent acquisition leaders at US tech companies who list HireVue or conversational AI interviewing as a skill.` - [run the search](https://www.refolk.ai/start?q=Talent%20acquisition%20leaders%20at%20US%20tech%20companies%20who%20list%20HireVue%20or%20conversational%20AI%20interviewing%20as%20a%20skill.).

*Returns the recruiters who actually run these screens, so you can confirm the tool and its rules rather than guessing from generic prep posts.*

## Verify before you click Finish

Run this check before you end the session. Each item is something you can confirm, not a topic to think about.

#### Pre-finish check

- [ ] I confirmed the role exists on the official careers page and the sender used the company domain.
- [ ] I called a number I sourced myself and the employer confirmed I am under consideration.
- [ ] I ran a test transcription and my proper nouns came back correct.
- [ ] I know whether this platform allows retakes, and I assumed none where it was unconfirmed.
- [ ] I identified whether my named platform scores delivery or only transcript content.
- [ ] Each answer landed a Result with a number before the agent advanced.
- [ ] Every probe for a specific example was resolved with a concrete, named detail.
- [ ] Any fumble ended with a one-line summary of the key action.
- [ ] I shared no SSN, passport, bank, or date-of-birth details.

## Keeping this current

This category changes by tool, not by rule, so the maintenance work is re-checking the named platform each time. Two facts are stable enough to rely on: most phone-based voice screens give no mid-session retake, and the most common probe is a request for a specific example. Build your habits on those.

Everything time-sensitive lives in the vendor's own documentation. Before each new screen, re-confirm three things from the invite and the platform's candidate help pages: the retake policy, whether delivery is scored, and the question count. When the invite names a tool you have not seen, search the employer's recruiting staff by that tool name to learn what they actually use, then read the vendor docs for that product. The procedure above stays the same. Only the platform-specific answers move, and you re-verify them one screen at a time.

## Frequently asked questions

### How long is an AI voice interview and how many questions will it ask?

Most run 10 to 20 minutes, with each question designed to take about two minutes. Vendor and research sources map length to roughly 3 questions at 10 minutes, 4 at 20, and 5 at 30, with methodology guidance capping sessions at 12 to 15 questions. Read your invite for the exact count, and if it is not stated, plan for four strong stories that each survive a couple of follow-up probes.

### Can I retake an AI phone screen if I stumble?

Usually not. Most phone-based live voice screens capture one continuous conversation with no mid-session retake, unlike one-way video where some platforms allow one retake per question. A few systems are exceptions, for example Mercor documents up to three attempts across related applications with only the most recent evaluated, but assume no retakes unless the invite says otherwise. If a genuine technical failure interrupts the call, email the recruiter for a fresh attempt.

### Does the AI score how I sound, or only what I say?

It depends on the platform, and the sources genuinely disagree. Enterprise systems like Greenhouse Voice AI explicitly exclude accent, tone, fluency, pace, stutters and confidence, scoring only transcript content against a rubric. Some candidate-prep tools instead sell vocal scorecards for clarity and persuasiveness. Confirm the specific platform named in your invite rather than optimizing for both, because the two assumptions pull your preparation in opposite directions.

### How do I know if an AI voice interview invite is a scam?

Verify the employer independently before you speak. Check the role on the company's official careers page, confirm the sender uses the company's own email domain rather than a free service, and call a phone number you sourced yourself, not one the recruiter gave you. Never share your SSN, passport, bank details or date of birth before a signed offer, since that is identity harvest. Report suspected scams at ReportFraud.ftc.gov.

### What is the best answer structure for an AI interviewer's follow-up questions?

Use STAR with a heavy Action section: Situation 10 to 15 percent, Task 10 percent, Action 50 to 60 percent in the first person, and Result 20 to 25 percent with a number. Practice the structure, not a word-for-word script, because scripted answers collapse on the second probe. Pre-stage branches for the common follow-ups: the alternative you rejected, who disagreed, and how you verified the result.

---

*From the Refolk guide library. I revise these guides rather than replacing them, so the current version is always at https://www.refolk.ai/candidates/guides/conversational-ai-voice-screen*
