The Proctored Interview Flag Reference, and What Trips an Honest Candidate
You will be able to name every signal an AI-monitored interview collects, spot which of your honest habits trips each one, and set up your room, disclosures, and answers so a clean interview reads clean.
Key takeaways
- An AI-monitored interview never observes cheating; it observes correlates of cheating, then a threshold turns a glance or a pause into an accusation.
- The load-bearing false-positive risk is textual, not visual: text detectors flag non-native English writers 61.22% of the time versus about 5.1% for native writers, roughly a 12x gap.
- The behavioral flag rate more than doubled in six months, from 15% in June 2025 to 35% in December 2025, and hit 48% on technical roles, which raises every honest candidate's baseline suspicion.
- Voice biometrics and gaze answer different questions: voice auth only proves who is on camera and says nothing about coaching, so passing identity does not clear you on gaze.
- Only about 11% of FAANG interviewers report using detection software, and vendors call the output probabilistic, which makes demanding human review your strongest recourse.
- Accommodations become violations if not pre-approved: one documented test was terminated at the flag limit for using a screen magnifier the candidate never cleared in advance.
You have an AI-monitored or auto-proctored interview scheduled, and you want to know what the system watches so your normal, honest behavior does not get you flagged as a cheater. This is a lookup document for the person being watched, not the person watching. Jump to the signal you are worried about, read what it claims to prove and what it looks like when it lies, and set your room, disclosures, and answering habits before you log in.
Every page that ranks for AI interview detection is written for the employer buying the tool. This one inverts it: each monitored signal maps to the honest behavior that trips it, the concrete countermeasure, and the recourse no vendor page publishes.
What signals does an AI-monitored interview actually collect?
An AI-monitored interview layers several distinct signals, and each claims to prove a different thing. Knowing which is which is the whole game, because passing one does not clear you on another.
The common signals are gaze and eye-line tracking, response-latency and answer-unfolding analysis, voice biometrics or speaker recognition, second-screen and tab or audio detection, speech-uniformity or coherence analysis, and liveness or identity checks. Vendors are explicit that these observe correlates, not the act itself.
Two clarifications matter most. Gaze is a coarse orientation signal, not identity: the camera is used for gaze direction analysis only, a coarse read on whether you are looking at your screen, and it is distinct from facial recognition, which characterises individuals from biometric features. Voice biometrics targets substitution, not coaching: voice authentication verifies that the person who registered is the person taking the interview, and it says nothing about whether the person on the call is being coached.
The signal stack, outermost first
- Identity and livenessProves who is on camera; a selfie or an active prompt to perform a simple action.
- Voice biometricsProves the registered person is speaking; silent on coaching.
- Gaze and framingCoarse read on whether your eyes stay screen-ward.
- Environment signalsTab, second-screen, and audio detection for outside help.
- Text and speech analysisLatency and coherence scored for AI-like uniformity.
The signal you are most likely to be falsely accused by is not on the video layer at all. It is the text-detector layer that scores your written answers, and it is where an honest candidate carries the most exposure.
Which signal is which, and what trips each one?
The table below is the core of this reference. Read the row for the signal you are worried about: what it claims to prove, the honest behavior that trips it, and the fix.
| Signal | What it claims to prove | Honest behavior that trips it | Countermeasure |
|---|---|---|---|
| Gaze tracking | You are reading off-screen | Looking up to recall or think | Keep eyes screen-ward, narrate aloud |
| Note-glancing detection | You are reading from a phone | Reading your own scratch notes | Declare notes on camera, or skip them |
| Response latency | You are fetching an answer | Pausing to compose a careful reply | Think aloud so the pause is filled |
| Text-detector score | An AI wrote your answer | Precise, low-perplexity or formal English | Keep drafts and version history |
| Voice biometrics | Someone else is substituting for you | None if you registered yourself | Complete registration cleanly, once |
| Second-screen or tab | You have outside help open | Nothing open if you closed everything | Single screen, phone in another room |
Every one of these signals turns a correlate into an accusation at a threshold. None of them observe cheating directly. They observe correlates of cheating, and then a threshold turns a correlate into an accusation. That single fact is what your setup and your recourse both hang on.
Where do false positives actually come from?
The load-bearing false-positive risk is textual, not visual. The reproducible number is a detector gap for non-native English writers, and it dwarfs anything documented on the video side.
A Stanford study published in the Cell Press journal Patterns found a 61.22% false-positive rate on essays from non-native English speakers versus about 5.1% for native speakers. The mechanism is measurable and fixable in a way that indicts the tool: enriching the vocabulary of those essays cut the average false-positive rate to 11.77%, and afterward only 1 of 91 essays was unanimously flagged as AI-written. That proves the detector measures vocabulary range, not authorship.
| Writer group | Avg false-positive rate | Source |
|---|---|---|
| Non-native English | 61.22% | Stanford / Patterns |
| After vocabulary enrichment | 11.77% | same study |
| Native English | ~5.1% | dev.to reporting |
Non-native writers are flagged roughly 12x more often than native writers (61.22 divided by 5.1), a figure derived from the two rows above. This is why take-homes and written assessments are where an honest candidate is most exposed. For context, OpenAI retired its own AI-text classifier at 26% accuracy, and seven detectors together flagged non-native writing as AI 61% of the time.
On the behavioral side, the documented false-positive triggers are all innocent: a nervous candidate who looks away to think, a non-native speaker who pauses, a neurodivergent applicant who does not hold eye contact. Note-taking is a specific trap. A candidate looking down to read a physical notebook they took for themselves, in a country where taking notes during an interview is normal, will trigger the same flag as a candidate reading from a phone.
The detector measures vocabulary range, not authorship, so a careful answer in a second language reads as a machine.
Why is your baseline suspicion higher than it used to be?
Your baseline suspicion is higher because real cheating rose and thresholds tightened around it, sweeping in innocent correlates. The behavioral flag rate more than doubled in six months.
Fabric's tracking of more than 50,000 candidates shows the share flagged for AI-assisted cheating rose from 15% in June 2025 to 35% in December 2025. On technical roles the rate hit 48% across 19,368 interviews. As the genuine fraud signal climbs, platforms lower their thresholds, and a pause or a glance that would have passed before now clears the bar for a flag.
| Point | Flag rate | Source |
|---|---|---|
| June 2025 | 15% | Fabric tracking |
| December 2025 | 35% | same |
| Technical roles | 48% | Fabric via qbsglobal |
The flag share more than doubled, 35 divided by 15 is 2.33x, a figure derived from the rows above. This is why disclosing your setup up front is now defensive rather than optional. At the same time, adoption of detection tooling is low precisely because accuracy is low: only about 11% of FAANG interviewers report their company uses cheating-detection software, and HireVue found that fewer than 1% of candidates produced highly similar shared responses. Many flags come from platforms whose own vendors call the output probabilistic.
Context on scale: 63% of active job seekers have now been interviewed by an AI, up 13 percentage points in six months, from a survey of 2,950 people. Only 21% believe most employers use AI responsibly, and 38% have walked away from a process because it included an AI interview.
How do I set up my room so honest behavior reads as clean?
Set up a front-facing camera with your full face in frame, light in front of you, a single screen, your phone in another room, the door closed, and the desk cleared. Framing and camera angle cause more false flags than any single behavior.
Framing matters because the proctor reads what it can see. If you look down and the view is primarily of your forehead, or if you turn far enough that only one eye is visible, the system reads it as looking at unauthorized resources. Camera angle compounds this: a front-facing view allows accurate face and gaze tracking, while a side-facing camera may misread normal movements as suspicious.
Camera angle and answering style
Real test-takers report concrete tactics that reduce flags: laying out their materials, keeping their appearance presentable, staring intently at the screen, sitting up straight, staying in range of the camera, and using breathing techniques to stay still. A dropped feed is its own trap: a missing video or audio feed is a major flag, so if your connection stutters, contact the proctor immediately rather than continuing silently.
If your written history is what a text detector will score, the defensible move is to have your resume and answers assembled from your own record, with your own drafts on file. Refolk writes your resume from your own history and tailors it per posting, which leaves you a clear authorship trail rather than a polished document you cannot account for.
The setup-to-recourse procedure
Run these steps in order, from reading the policy through the appeal path. Each step has a clear done-state so you can tell when it is complete before the next one starts.
From policy to appeal
- Read the AI policy and get it in writingFind the stated rule on permitted AI assistance; if none exists, email to ask, since whether AI use disqualifies is disclosure-dependent. Done when you can state exactly what is allowed.
- Request accommodations in advanceSecure written approval for any assistive tool or extra time days ahead, because accommodations cannot be added once an exam has started. Done when approval is on file.
- Complete the identity and liveness checkFollow the selfie or active liveness prompt exactly on the first try. Done when identity is confirmed.
- Set up the room and cameraFront-facing camera, light in front, single screen, phone in another room, door closed, desk cleared. Done when your full face is in frame with no second device reachable.
- Interview under continuous monitoringThe system collects gaze, latency, audio, tab activity, and coherence throughout; keep eyes screen-ward and narrate. Done when the session ends with a timestamped flag log.
- Let the suspicion score assembleThe platform resolves collected correlates into a probability or trust level, not a verdict. Done when scoring resolves to a number.
- Route flagged sessions through human reviewBest practice is a human examining flagged moments before any outcome, with the option to contact you. Done when a flag is confirmed, dismissed, or queried.
- Use the appeal or redo pathRequest the flag log and the human-adjudication workflow. Done when the outcome arrives with a stated path to contest it.
Sources disagree on whether human review comes before or after the outcome. Integrity Advocate places review before the decision issues; other workflows treat it as a post-decision appeal. Request review either way, because you cannot tell in advance which order your employer uses.
How this goes wrong: failure modes and false positives
The failure modes below are the most valuable part of this reference, because each is an honest behavior that the system cannot distinguish from cheating. Read the check column as your defense.
- Gaze flag on honest thinking. Looking up or away to recall triggers the same signal as reading a screen; there is no built-in distinction. Check: keep your eyes screen-ward and narrate your reasoning aloud so the thinking is audible.
- Note-glancing reads as phone-reading. Physical notes you wrote yourself flag identically to a phone. The false positive is your own scratch paper. Check: declare notes on camera at the start, or do not use them.
- Text detector penalizing careful English. A precise, low-perplexity answer reads as AI. The false positive is an ESL candidate or a formal, PhD-style writer. Check: keep your drafts, and hold that employers should never reject on a detector score alone.
- Side camera misreads motion. A profile view turns normal head movement into looking away. Check: front-facing camera with your face fully in frame.
- Accommodation behavior flagged as cheating. Screen-magnifier clicks or leaning in get flagged; one documented test was terminated at the flag limit. Check: approve accommodations in advance so approved behaviors are not treated as violations.
- Dropped audio or video read as evasion. A missing feed is a major flag. Check: contact the proctor immediately, and do not continue silently.
- Treating a probability as a verdict. A suspicion score is a correlate, not proof. Check: demand human review and the flag log rather than accepting the outcome.
The through-line across all seven is that voice biometrics and gaze answer different questions, and conflating them mislabels honesty. You can pass the identity check cleanly and still be flagged on gaze, and the reverse is equally true. They are not interchangeable safety nets.
What recourse do you have when you are flagged?
Your strongest lever is demanding the human review that best practice already prescribes, plus the timestamped flag log. Because a flag is probabilistic and cannot prove intent or authorship, a human reviewer gives context, examines edge cases, and may contact you for an explanation.
Documented redress includes advance human review and a formal appeal workflow: platforms are advised to tune thresholds to reduce false positives, enable human review, and provide an appeal workflow with human adjudication so flagged candidates can explain potential false positives. When you appeal, ask for the specific flagged moments rather than raw footage. Recruiters receive flagged moments, not hours of raw footage, and timestamped logs let a reviewer jump straight to the problem spots, so those same timestamps let you explain each one.
The stakes are real. Terminations happen at flag limits, where a test is ended midway once a candidate hits the maximum number of allowed flags. That is exactly why front-loading your disclosures and setup matters more than any argument after the fact.
Subject: Request for flag log and human review, [role] interview Hi [name], I understand my recent [assessment or interview] was flagged by the monitoring system. I want to resolve this and give you the context an automated score cannot. Could you share the timestamped flag log so I can address the specific moments? I can already account for the likely triggers: - I looked up several times to recall details; I was thinking, not reading. - I used notes I wrote for myself, which I disclosed on camera at the start. - My written answers are my own; I can share drafts and version history on request. Since a suspicion score is a correlate rather than proof, I would appreciate a human review before any decision, and a redo if the log is inconclusive. I am happy to walk through any moment live. Thank you, [your name]
Send to the recruiter or coordinator; adapt the flagged behaviors to your own.
Who sits on the other side of these decisions is a real and searchable population. In Refolk's index of professional profiles, there are about 4,479 US Trust and Safety, fraud, and identity-verification professionals, against about 3,559 in the same roles in India, a US pool roughly 1.26x the size (4,479 divided by 3,559, derived). Top employers in that set included Walmart in the US and Bank of America and Barclays in India.
| Market | Trust, fraud, and identity roles | Source |
|---|---|---|
| United States | ~4,479 | Refolk's index |
| India | ~3,559 | Refolk's index |
If you want to understand how these systems are built and where they misfire, you can find the people who build them. Refolk lets you search that population directly.
Pre-interview checklist
Run this before you log in. Each item is a check you either pass or fix, not a topic to think about.
Clear before you connect
- I have the employer's AI-use policy in writing, or an email asking for it.
- Any accommodation I need is approved and on file, not requested mid-session.
- My camera is front-facing and shows my full face with light in front of me.
- My phone is in another room and only one screen is active.
- I have decided whether to use notes, and if so I will declare them on camera.
- I know I will narrate my reasoning aloud instead of pausing silently.
- I have my own drafts and version history saved in case a text detector flags my writing.
- I know how to contact the proctor immediately if my audio or video drops.
How to keep this reference current
The signals in this document are stable, but the thresholds are not, so re-check the numbers rather than the mechanism. The flag rate that doubled from 15% to 35% in six months tells you the moving part is the threshold, not the list of signals, and thresholds tighten as real fraud rises.
Two things are worth re-checking before each interview. First, the disclosure policy, because permission is disclosure-dependent and each employer sets its own rule; get it in writing every time. Second, whether the specific platform routes flags through human review before or after the decision, since sources genuinely disagree on that order. When in doubt, ask for the flag log and human adjudication regardless. The one claim you can rely on across every platform is that these tools observe correlates, not cheating, so your job is to make your honest behavior legible and your authorship provable before the session starts.
Questions job seekers ask
Will looking away to think get me flagged for cheating?
It can. Gaze tracking is a coarse orientation signal, and looking up or away to recall triggers the same flag as reading off a screen, because the system has no built-in way to tell honest thinking from reading. The countermeasure is to keep your eyes screen-ward and narrate your reasoning aloud so your process is audible even when your eyes drift. This turns a silent glance into visible thinking a human reviewer can read as honest.
Can I use my own handwritten notes during an AI-monitored interview?
Only if you declare them. Physical notes you wrote yourself flag identically to reading from a phone, because the camera sees you looking down and cannot read the page. Show the notes on camera at the start and say you took them for yourself, or do not use them at all. In some countries note-taking is normal, but the system does not know your context, so disclose it explicitly.
I am a non-native English speaker. Am I more likely to be falsely flagged?
On text detectors, yes, and by a large margin. A Stanford study found a 61.22% false-positive rate on essays from non-native English writers versus about 5.1% for native writers, because detectors measure vocabulary range rather than authorship. Enriching vocabulary cut that rate to 11.77%. Take-homes are where you are most exposed, so keep your own drafts and version history to prove authorship if challenged.
How do I appeal being flagged for cheating in an interview?
Ask for the timestamped flag log and the human-adjudication workflow. A flag is probabilistic and cannot prove intent or authorship on its own, and best practice routes it through a human who can examine edge cases and contact you for context. Request the specific moments that triggered the score, explain each honest behavior, and point out that a suspicion score is a correlate, not proof.
Does passing the voice or identity check mean I am cleared of cheating?
No. Voice biometrics and identity checks only verify that the person who registered is the person taking the interview, and they say nothing about whether that person is being coached or reading from a second screen. You can pass identity cleanly and still be flagged on gaze or tab activity. The signals answer different questions and are not interchangeable safety nets.
What should I do if my audio or video drops mid-interview?
Contact the proctor immediately and do not continue silently. A missing video or audio feed is a major flag, because the system reads the gap as possible evasion rather than a connection fault. Flag the problem the moment it happens so there is a record that the drop was technical. Continuing without your feed lets the silence accumulate against you.
Put this to work
Paste your career in once. Every application after that is written for you.
Drop a resume or a LinkedIn URL. I rank the live openings against it, rewrite the resume and write a cover letter for the best of them, and fill in the employer's form when you press the button. You read, you decide what goes out.
01Drop your resume
A PDF or a LinkedIn URL. About a minute, once.
02I rank the openings
Every weekday morning, the live catalog scored against your history. Up to 20 worth your time, not two hundred links.
03Each one is written up
Resume rewritten for the posting, a cover letter, a fit score. Press send, or let me fill in the form.
- New matches ranked and written before you are up.
- Every bullet stays inside what your history supports. Nothing invented.
- Queued, submitted, interviewing, offer: one screen, not a spreadsheet.
500 free credits on sign-up. No card. Nothing is sent until you say so.