48% of Tech Interviews Are Flagged for Cheating. Source the Signal Instead.
Cluely and Interview Coder broke live coding rounds. Here is why sourcing verified GitHub history is now the only technical signal you can trust.
The live technical interview stopped being a measurement instrument sometime in Q3 2025. Fabric's analysis of 19,368 AI-powered interviews between July 2025 and January 2026 found 48% of software engineering candidates flagged for AI-cheating behavior, and the tools driving that number are architecturally invisible to Zoom, Meet, and Teams. If your hiring loop still weights the coding round like it did in 2023, you are hiring on noise.
The single number that should end the debate
In Fabric's 19,368-interview dataset, 48% of technical candidates were flagged for AI-cheating behavior, and 61% of those flagged would have advanced through a standard hiring process without detection tooling. That is roughly one in four technical hires clearing your loop on a fraudulent signal.
The rate did not creep. It stepped:
- July 2025: 9% of interviews flagged
- September 2025: 45% flagged
- January 2026: sustained near the September peak
- Sales roles, same window: 12%
- Sunday interviews, all roles: 47.1%
The 5x ramp in 90 days is not a bell curve of new bad actors. It is a tooling adoption curve. Cluely, Interview Coder, and Final Round AI now have more than one million combined users, and a $20 to $50 monthly subscription against a $150,000 engineering salary produces a risk/reward ratio that makes the decision trivial for anyone without a strong honor reflex. 83% of candidates in one survey said they would use live AI assistance if they believed they would not get caught.
Why the "live coding round" is now the least reliable round
Live coding is the round most exposed to fraud, not the most protected, because modern cheating overlays render below the screen-share capture layer. The interviewer literally cannot see the assistance on the shared screen.
Interview Coder and its successor Cluely exploit a graphics-stack quirk:
- On Windows, Cluely uses DirectX overlays.
- On macOS, it renders through Metal framework layers.
- Both sit at a compositor depth below what Zoom, Google Meet, and Microsoft Teams capture when a candidate shares their screen.
- Final Round AI runs in a hidden browser tab; Interview Hammer runs on a second device with the camera pointed at the primary screen.
Roy Lee and Neel Shanmugam, the two Columbia students who built Interview Coder, were suspended, dropped out, rebranded as Cluely, and raised $5.3M in seed and then $15M Series A from Andreessen Horowitz. Lee has said the Leetcode-style questions the tool defeats are "useless." Whether or not you agree, an a16z-backed company with a million users is now productizing the exact failure mode your interview loop assumes cannot happen.
Karat, whose whole business is running technical interviews, estimates 80% of candidates use LLMs during code tests despite explicit prohibition. Dedicated overlay tools like Cluely and Interview Coder account for 45% of flagged cases in the Fabric dataset. Everything else is browser tabs, phones, and second monitors.
The detection arms race is a trap
Truely, a counter-tool built by two other Columbia students specifically to detect Cluely, is already shipping. Cluely will patch around it. Truely will patch back. Any hiring process whose integrity depends on winning that cycle in real time, on a Tuesday afternoon, against a candidate whose only job for the next hour is to defeat you, is not a hiring process. It is a bet.
Detection is a losing arms race. Provenance is not. Git history from March 2024 cannot be generated the night before the interview.
What is actually verifiable in 2026
The only technical signals that survive the Cluely era are ones a candidate cannot generate retroactively: dated commit history, merged PRs in named repos, co-authors, issue threads, and shipped artifacts with public provenance. Everything else, including live coding, take-homes, and resume claims, is now assumed contaminated.
The asymmetry matters. A cheater can pass a 45-minute Leetcode round with a $20 subscription. They cannot go back to March 2024 and author 40 commits to a Kubernetes operator in a repo with 800 stars, with review threads from maintainers, tagged releases, and issue triage across nine months. The cost curve of faking prior work at scale is not $20 a month. It is "hire a ghost committer for a year," which nobody is doing for a $150K job.
Here is the comparable dataset from Fabric, laid out plainly:
| Segment | Flag rate | Notes |
|---|---|---|
| All roles (avg, 19,368 interviews) | 38.5% | July 2025 to January 2026 |
| Software engineering | 48% | Where checkable answers exist |
| Sales roles | 12% | 4.0x lower than engineering |
| July 2025 (window start) | 9% | Pre-Cluely mass adoption |
| September 2025 (peak) | 45% | 5.0x ramp in 90 days |
| Sunday interviews | 47.1% | Highest weekday cohort |
The 4x engineering-to-sales gap is the tell. Sales interviews reward improvisation and rapport, which LLMs help with only marginally. Coding rounds have clean, checkable answers, which is exactly what an overlay LLM is optimized to produce. The rounds you designed to be "objective" are the rounds that got hollowed out first.
Big Tech's fix does not work for remote-first companies
Google, Amazon, Jane Street, and Hudson River Trading are quietly re-shoring key interview rounds in person because AI-assisted cheating has broken the remote signal. If you do not have a Manhattan office and a candidate travel budget, that fix is unavailable to you.
Amazon has explicitly banned AI tools in interviews. Google has openly discussed pulling final rounds back on-site. Jane Street and HRT keep final rounds in person. This works if you are a firm whose median offer clears $400K and whose loop is 15 candidates a quarter for one desk. It does not work if you are a 40-person Series B running 300 top-of-funnel a month for six roles, with candidates in Lisbon, Bangalore, and Warsaw.
For everyone else, the substitute is pre-interview verification: source on shipped work first, then use the interview to confirm depth, not to generate the primary competence signal. That inversion is the point. This is the exact gap Refolk closes for engineering hiring: you describe the person in plain English (say, "backend engineers who have shipped production Postgres extensions in the last two years, active on GitHub, based in EU time zones") and get a ranked shortlist built from GitHub, LinkedIn, and the open web, with the provenance signals attached.
The four sourcing signals that hold up under adversarial conditions
Weight these signals ahead of any live coding score, in this order. They are the ones an overlay cannot fake in a 45-minute window.
- Dated commit history in named repos. Look for consistent activity over 12+ months, not a March 2026 sprint of green squares. Co-authored-by trailers and PR review threads add human corroboration.
- Merged PRs in projects the candidate does not own. Owning your own repo proves nothing. Getting code accepted into a project maintained by strangers proves the candidate can read a codebase, follow contribution guidelines, and respond to review.
- Public artifacts with provenance. Talks with recordings, RFCs with author attribution, released packages with download counts, blog posts indexed before the job search started. The date stamps matter as much as the content.
- Weak-tie corroboration. Who has starred or forked their work, who has co-authored commits, who reviewed their PRs. This is the hardest signal to fabricate because it requires other people's calendars and reputations.
60% of engineers applying for roles lack the skills they claim on their profiles, and that number predates Cluely. Resume signal degradation and interview signal degradation are now stacked. What is left is the artifact trail, and GitHub's roughly 100 million developers make the artifact trail unusually rich for engineering hires compared to almost any other field.
The operational problem is that filtering 100 million profiles on plain-English criteria was, until recently, a Boolean-string exercise that most recruiters could not run and most engineering managers did not have time for. Refolk turns that into a natural-language query: "senior Rust engineers with published crates over 10K downloads, currently at Series A or B companies, open to remote," returns a ranked list with the underlying signals visible. That is a screen you can actually trust in 2026.
What the new loop looks like
Move screening weight up-funnel into verified sourcing, and use interviews for judgment and collaboration signals that overlays cannot help with. Concretely:
- Cut Leetcode-style rounds to one, or drop them entirely. They now measure Cluely subscription status, not skill.
- Add a "walk me through your PR" round. Pick a specific merged PR from the candidate's public history. Ask them to explain the tradeoffs, the review feedback, and what they would do differently. Overlays cannot help with the candidate's own history.
- Weight collaboration signals. Design review, code review of your team's real code, incident postmortem discussion. Judgment and communication are still measurable in a live setting.
- Verify before you interview. If the candidate's public work does not clear the bar for the role, the interview will not save you. If it does, the interview is a confirmation round, not a discovery round.
- Assume assistance, design accordingly. The 83% "would if they could" number kills any honor-code strategy. Design rounds where AI assistance either does not help or is explicitly allowed and observed.
Engineering managers I have talked to are quietly running this loop already. The recruiters who work with them are spending less time scheduling and more time on candidate verification engineering, which is the actual job now. Sourcing verified engineers, not filling calendars, is where the leverage sits.
FAQ
Are AI cheating detection tools reliable enough to keep live coding rounds?
Not durably. Detection works for the current generation of overlays, but Truely versus Cluely is a live arms race with new releases on both sides monthly, and the incentive gradient favors the attackers because the payoff per successful cheat is a $150K salary and the cost is a $20 subscription. Detection also runs after the fact on behavioral signals, so even when it works you have already spent the interviewer time. The durable defense is shifting evaluation weight to signals that cannot be generated in the interview window at all, which means verifiable prior work.
How do I verify GitHub activity is actually the candidate's own?
Cross-reference commit email addresses, co-author trailers, and PR review threads. Look for consistent activity over 12+ months rather than a recent burst, check that the candidate can talk fluently about specific merged PRs in an interview, and look for weak-tie corroboration: who reviewed the code, who co-authored commits, who has starred related repos. A candidate who cannot walk you through the tradeoffs in a PR they allegedly wrote 14 months ago is telling you something. Tools that aggregate this across GitHub, LinkedIn, and the open web, including Refolk, make the cross-reference tractable at pipeline scale.
What about candidates without a public GitHub presence?
They exist, and they are often excellent, especially engineers coming out of finance, defense, and older enterprise shops. For those candidates, lean on artifact provenance in other forms: internal talks that got recorded and posted, patents, published papers, package registries other than GitHub, and reference calls with named former colleagues you can independently verify are real people at real companies. The principle is the same: something dated, attributable, and hard to fabricate under time pressure.
Does re-shoring interviews in person actually solve the problem?
Partially, and only for companies that can afford it. Google, Amazon, Jane Street, and HRT are pulling key rounds in person, but that assumes you can fly candidates to an office, that candidates will accept the travel, and that you are competing on total comp levels that justify the friction. For a Series B remote-first team hiring in five countries, in-person final rounds are not economically viable at pipeline volume. Verified sourcing is the remote-first substitute, and for most companies now it is not optional.