RefolkCandidates
9 min read

Sarah the AI Interviewer Went Berserk. Here's What She Actually Scores.

The viral gibberish AI interview clip exposed how HireVue style bots really score candidates: keyword mirroring, STAR structure, and the AI threshold.

On October 6, 2026, a TikTok creator sat down with an AI avatar named "Sarah," answered her questions in pure gibberish, and watched the bot escalate from stock acknowledgements to accusing him of a mental breakdown before hanging up. If you have an AI video interview on the calendar this quarter, that clip is the clearest free lesson you will get on what these systems actually measure, and what they do not.

What the gibberish video actually proved

The bot did not detect gibberish. It detected the absence of expected keywords, collapsed its own scoring signals to zero, and layered LLM-generated frustration on top. That is the entire mechanism in one sentence.

Here is how the viral clip, posted by TikTok's The Dreadful Studio, played out. The candidate interviewed for an assistant manager role with an AI avatar called Sarah. He answered every question with nonsense syllables. Sarah first offered stock acknowledgements, then escalated to "I do not know if you're having a breakdown or if you simply lack the capacity to speak English," and finally, per Futurism's October 6 coverage, "I am ending this screening now. You are completely unqualified. Goodbye."

Related trolling-the-AI videos on TikTok have stacked 278,600 likes and nearly 5,000 comments in the same week. The engagement is a tell: candidates are getting enough AI interviews to make pranking them a genre.

What the bot was doing under the hood is boring and important:

  • Transcribing audio to text (HireVue uses Rev.ai for this step).
  • Scoring the transcript against a competency rubric that expects specific phrases.
  • Checking narrative structure (does the answer have a situation, an action, a result).
  • Triggering scripted follow-ups when a signal is missing or weak.

Gibberish produces an empty transcript. Empty transcript equals zero keyword overlap equals zero completeness score. The hostile dialogue was generative theater on top of a dead scoring pipeline.

0.25%
Lift HireVue got from facial expression analysis before killing it in 2021

The nonverbal cues candidates still coach for were worth a rounding error, which is why the market leader removed them.

How AI interviewers score candidates in 2026

AI interviewers score the transcript of what you said against a fixed rubric of competency keywords and narrative-structure signals, not your face, your tone, or your pauses. That is a direct quote from HireVue's 2025 AI Explainability Statement: "our AI relies only on what is said by the candidate and does not use any video analysis or other audio characteristics."

Three things follow from that, and most interview prep content has not caught up to any of them.

The model is static and deterministic

HireVue's scoring model is trained in a controlled environment by Industrial-Organizational Psychologists before deployment, then shipped. The same keywords always trigger the same signals. There is no live learning from your specific answer. If "cross-functional collaboration" is in the rubric for a product manager role, you either say it or you do not, and the score moves accordingly.

The AI threshold is the real gate

HireVue is used by 700+ employers worldwide including JPMorgan, Goldman Sachs, Unilever, and Bain. The AI scores every answer, ranks you, and surfaces a shortlist for human recruiters. If your score does not clear the AI threshold, no human watches your recording. The "human in the loop" line is technically true and operationally hollow for everyone below the cutoff.

Facial expression hacks are dead weight

HireVue stopped scoring facial expressions in 2021. Its own data (Fortune, 2021) said the nonverbal cues added about 0.25 percent to the model in most roles. If a coach is still telling you to practice "warm eye contact" for a HireVue screen, they are optimizing for a signal that no longer exists at the market leader.

Why US candidates feel this first

The AI interviewer is not replacing a recruiter shortage. It is rationing throughput, and US employers hit the throughput wall first because the US recruiting function is roughly 12 times the size of the UK's.

That ratio comes from Refolk's index of professional profiles, filtered on titles like Recruiter, Talent Acquisition, and Talent Partner:

SegmentCountNote
US recruiters / TA professionals98,093Refolk's index, title filter
UK recruiters / TA professionals7,964Refolk's index, same filter
US-to-UK recruiter ratio~12.3xDerived from the two rows above
Share of UK recruiter sample in London48% (12 of 25)Refolk's index, top regions
Top US employer concentration (HMBL Talent) in sample8% (2 of 25)Refolk's index, top companies

The US recruiting function is bigger in absolute terms but still gets drowned by application volume, especially since auto-apply tools pump hundreds of applications per candidate per week. Wonderin.ai users have reported receiving so many AI interviews they started pranking them for sport. That is the Dreadful Studio clip's actual context: an infrastructure response to a volume crisis, with candidates on the receiving end.

The keyword mirroring tactic that beats polish

The highest-leverage prep move for an AI video interview is to repeat the job posting's exact competency phrases back inside your answers, verbatim, not paraphrased. This inverts normal interview advice and it is the single tactic that moves the score most.

Because the model is deterministic and trained on exact competency phrases, saying "cross-functional collaboration" scores higher than the smoother "working with different departments." The rubric is looking for a token match, not for prose quality.

How to actually do this before the interview:

  1. Pull the top 10 to 15 noun phrases from the posting. Verbs matter less; the rubric is noun-heavy.
  2. Group them into three or four competency buckets (ownership, collaboration, measurement, technical scope).
  3. Draft two STAR stories per bucket where the exact phrase lands inside the Action or Result sentence.
  4. Rehearse out loud until the phrase comes out whole, not reworded.

The failure mode is improvisation. Under pressure, trained speakers smooth language into their own idiom and lose the exact token the model is scanning for. This is also where the mirror between your resume and the posting matters: if your resume already uses the posting's language, your spoken answers tend to as well because you have said the words a dozen times editing bullets. Pasting a posting into Refolk and getting your own resume back rewritten with the posting's phrasing is a shortcut, because the rewrite surfaces the exact tokens you then rehearse into your HireVue answers.

The bot cannot tell a brilliant candidate from a bad one. It can only tell a matching transcript from a non-matching one.

STAR is not advice, it is structural compliance

The STAR method (Situation, Task, Action, Result) is not a nice-to-have for AI video interviews. It maps directly to the narrative completeness signals the scoring algorithm checks for, which is why answers that skip any of the four pieces lose points even when the content is strong.

Think of STAR as a checklist the model runs on your transcript:

  • Situation: establishes context tokens the rubric expects (industry, team size, scope).
  • Task: creates the goal statement that makes "Result" scorable.
  • Action: the slot where competency keywords live.
  • Result: the quantified outcome that trips the measurement signal.

Drop Result and your answer looks incomplete to the model, even if the Action was impressive. Drop Situation and the keywords in Action float with no context, which the rubric reads as weak evidence. The structure is doing scoring work, not just helping you sound organized.

Why you will get grilled on fewer topics, harder

Expect AI interviews to produce roughly twice as many follow-up questions as a human screen but cover fewer distinct topics. Prep two or three deep stories, not eight shallow ones.

A 2026 arXiv study of AI-conducted interviews (arxiv.org/pdf/2608.01640) found AI-assisted interviews averaged 54.5 interviewer turns versus 32.1 in human-led training sessions, and 33.7 follow-up questions versus 16.3. Topic breadth went the other way: AI interviews covered about 34 percent fewer distinct topics.

The practical read: the bot will pick two or three themes from your opening answers and keep drilling. If your best story is a 14-month platform migration where you led a cross-functional team and shipped a measurable revenue outcome, that story needs five or six layers of detail you can produce on demand. Breadth loses. Depth on the right themes wins.

700+
Employers running HireVue AI screens, including JPMorgan, Goldman Sachs, Unilever, Bain

If your score does not clear the AI threshold, no human recruiter watches your recording.

The Goldman Sachs double standard, and what to do about it

Employers are deploying AI interviewers at scale while asking candidates not to use AI tools in return. Goldman Sachs, per Futurism, wants applicants to stop relying on AI during interviews even as it runs AI-scored screens on the front end. That tension is real, and it is why Cluely, a viral AI "cheating" startup from Roy Lee, raised $15 million led by Andreessen Horowitz specifically to help candidates beat AI interviews.

You do not need Cluely to be competitive. You need two things most candidates skip:

  • A resume whose bullets already use the posting's exact competency language, so your spoken answers inherit it. This is the exact work Refolk does when you paste a posting: it rewrites your resume from your own history into the words the AI rubric is scanning for.
  • A fit read before you record, so you know which of your stories hit the posting's rubric hardest and which two or three themes to go deep on before the AI starts drilling.

The ethics argument will outlast this news cycle. The scoring mechanics will not change meaningfully in the next six months. Prep for the mechanics.

Prepare for an AI video interview in 2026: a short checklist

If you have a HireVue, Sapia, or avatar-based screen on the calendar, work this list in order the night before.

  1. Extract 10 to 15 competency phrases from the posting, verbatim.
  2. Draft three STAR stories where those phrases live inside the Action sentence.
  3. Quantify every Result. One number per story, minimum.
  4. Rehearse out loud so the phrases come out whole under pressure.
  5. Plan for 50-plus turns and heavy follow-ups on your top two themes.
  6. Ignore facial expression coaching for HireVue; it has not been scored since 2021.
  7. Speak clearly for the transcriber; Rev.ai is good but not perfect with accents and crosstalk.
  8. If a question lands on a theme you did not prep, bridge back to one of your three stories within two sentences.

The gibberish video is a gift. It made the scoring machinery visible for free. The candidates who treat it as entertainment will keep losing to the ones who treat it as a spec sheet.

FAQ

Does HireVue analyze my face or tone of voice in 2026?

No. HireVue's 2025 AI Explainability Statement is explicit that the model "relies only on what is said by the candidate and does not use any video analysis or other audio characteristics." The company stopped scoring facial expressions in 2021 after its own analysis showed the nonverbal cues added only about 0.25 percent to model performance in most roles. If a coach is still selling you facial expression drills for HireVue, they are selling you prep for a 2019 system.

What are the highest-leverage AI avatar interview answers?

Answers that repeat the posting's exact competency phrases inside a complete STAR structure with a quantified result. The scoring model is static and deterministic, trained on specific competency keywords by Industrial-Organizational Psychologists before deployment, so verbatim phrase matches beat polished paraphrases. One number in your Result sentence trips the measurement signal that many candidates miss entirely.

How does AI interview keyword matching actually work?

Your spoken answer is transcribed to text (HireVue uses Rev.ai), then the transcript is scored against a competency rubric that looks for specific noun phrases and narrative structure. The model checks whether the expected phrases appear, whether the answer has situation and result components, and how dense the keyword overlap is. It does not understand your answer; it matches tokens. That is why the gibberish video failed catastrophically: no tokens matched, so every signal collapsed.

Will a human recruiter watch my recording either way?

Usually not if you score below the AI threshold. HireVue is used by 700-plus employers who rely on the AI ranking to produce a shortlist, and recruiters typically only watch recordings above the cutoff. The "human in the loop" marketing claim is technically accurate but operationally meaningless for candidates in the bottom half of the AI score distribution, which is why clearing the AI threshold is the first job.

Put this to work

Paste your career in once. Every application after that is written for you.

Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.

  1. 01Drop your resume

    A PDF or a LinkedIn URL. About a minute, once.

  2. 02I rank the openings

    Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.

  3. 03Each one is written up

    Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.

  • New matches ranked and written before you are up.
  • Every bullet stays inside what your history supports. Nothing invented.
  • Queued, submitted, interviewing, offer: one screen, not a spreadsheet.

500 free credits on sign-up. No card. Nothing is sent until you say so.

Keep reading