61% of AI Cheaters Still Pass. Source From GitHub Before the Screen.
Live-coding interviews just crossed a 38% AI-assist flag rate, and 61% of flagged candidates still pass. Here is how to flip the funnel to public artifacts.
The live-coding screen is finished as a hiring signal. Fabric HQ's analysis of 19,368 technical interviews between July 2025 and January 2026 put the AI-cheating flag rate at 38.5%, climbing from 9% in July to 45% by September. Anthropic quietly rewrote its own interview questions in January because too many candidates were solving them with Claude in real time. If the screen is broken, the fix isn't better proctoring. It's sourcing.
The 38% headline hides a worse number: 61%
The real crisis isn't detection, it's that detection doesn't matter. Fabric HQ found that 61% of candidates flagged for AI assistance still cleared the interview score threshold. The format is producing positive signals from fraudulent inputs, which makes it a false-positive machine, not merely a leaky one.
The pass-through happens because scoring rubrics reward the artifact (working code, correct answer, articulate walkthrough) and the AI overlay produces exactly that. A rubric that grades output cannot distinguish between a candidate who reasoned to the output and one who transcribed it from a Cluely overlay in the 1 to 2 seconds it took the LLM to answer. The interviewer sees the same screen either way.
That is why "add better proctoring" is the wrong first move. You would be tightening a filter that lets more than half of caught cheaters through on merit as scored.
Why the proctoring arms race is already lost
You cannot screen-share your way out of this. Modern cheating tools operate at the OS graphics layer, beneath the surface that screen-share capture can see, so they are architecturally invisible to CodeSignal, HackerRank, and every browser-tab proctor on the market.
Here is the stack candidates are actually running against you in 2026:
- Cluely. Chungin "Roy" Lee's Columbia dropout project. $5.3M seed from Abstract Ventures and Susa Ventures, then a $15M Series A from Andreessen Horowitz two months later. 70,000 signups in the first week at $20 per month. Tagline: "cheat on everything."
- Interview Coder. Same founder, earlier product. Lee has said publicly he used it to land an Amazon internship; Amazon declined to comment on his case but told TechCrunch that candidates must acknowledge they won't use unauthorized tools.
- LeetcodeWizard, Parakeet AI, FinalRound AI. Same category, invisible overlays, sub-2-second suggested responses driven by interviewer audio and on-screen text.
These tools display AI-generated answers directly over the coding environment. The overlay is visible only to the candidate. When a candidate runs Cluely alongside a CodeSignal browser tab, the AI assistant runs as a desktop application at a layer the assessment vendor cannot access. That is a permanent architectural gap, not a bug to be patched.
CodeSignal itself now publicly acknowledges its own proctored channel doubled in cheating rate, from 16% in 2024 to 35% in 2025, with APAC hitting 48% against 27% in North America. Any detection vendor citing 90%+ catch rates is measuring last season's tools.
The interview is banning the exact skill the job requires
Testing engineers without AI in 2026 is like testing them without Google in 2010. According to Dice, 71% of US tech job postings now require some form of AI fluency, a 181% year-over-year increase. Live-coding-without-AI screens are measuring a skill the role no longer needs.
The mechanism is quietly absurd: hiring is optimizing against the tool the new hire will use every day after they start. A candidate who cannot use Claude, Cursor, or Copilot productively on day one is a worse hire in 2026 than one who cannot solve a medium Leetcode by hand. But the screen inverts that ranking.
The escape is not "allow AI in the interview," which only accelerates the 61% pass-through problem. It is to stop letting synchronous interviews carry the weight of the decision at all.
Flip the funnel: verify public artifacts before the first call
The durable fix is a sourcing-first funnel that anchors on artifacts a human did not perform in front of you. Commits, PRs, issue comments, package ownership, and contribution graphs are dated, signed to an account, and impossible to fake in the 1 to 2 seconds an overlay can respond.
Concretely, the pre-interview evidence stack for a senior backend hire should include:
- Commit history depth and cadence on repos the candidate claims. Look for multi-year presence, not a September spike.
- PR authorship in repos the candidate did not create. External contributions signal collaboration and review acceptance, which no overlay can fabricate retroactively.
- Issue comments and design discussion. Prose reasoning under a stable handle, dated years back, is the highest-fidelity signal of how someone thinks.
- Package ownership on npm, PyPI, crates.io, or Maven. Download counts and maintenance cadence are audit-trail evidence.
- Cross-referenced identity. The same handle, email, or commit signature appearing on GitHub, a personal site, and a conference talk from 2022 is not something Cluely can produce mid-call.
None of this replaces judgment. It replaces the fraudulent input the current funnel is training on.
The scale problem: 358,700 engineers, 4.8% with visible GitHub
Sourcing from artifacts sounds obvious. The reason nobody does it at scale is that most engineers do not advertise the artifacts. In Refolk's index of US professional profiles, roughly 358,700 people carry Software Engineer, Backend Engineer, or Full-Stack Engineer titles. Only about 17,000 of them, or 4.8%, list GitHub as a surfaced skill or artifact on their profile.
That 4.8% is a self-selected minority you want to prioritize, but the other 95.2% are not artifact-free, they are artifact-hidden. Their commits exist. Their PRs exist. The link between profile and handle is the missing piece, and it is the single most valuable thing to resolve before you schedule a call.
This is the exact gap Refolk closes. You describe the person in plain English ("senior Python engineer in Austin who has shipped to a large open-source data tool in the last 24 months") and get a ranked shortlist with the GitHub, LinkedIn, and open-web signals stitched together, so the artifact check happens before the calendar invite, not after.
The math of why this has to be automated
A human recruiter cannot manually audit 15 GitHub histories per hire in a market with 40% more interviews per hire than 2021. In Refolk's index, the US has roughly 23,970 technical recruiters, sourcers, and TA professionals against 358,700 engineers in the three core SWE titles alone. That is a ratio of about 15:1, concentrated in the SF Bay, Austin, and Seattle markets where the hiring bar is already highest.
Layer the interview inflation on top and the operational picture is stark:
| Metric | Value | Source |
|---|---|---|
| Tech candidates flagged for AI cheating, 2025 to 2026 | 38.5% | Fabric HQ, 19,368 interviews |
| Same metric, purely technical roles | 48% | Fabric via metaintro.com |
| Proctored assessment cheating, 2024 to 2025 | 16% to 35% | CodeSignal, Feb 2026 |
| Entry-level cheating rate, 2024 to 2025 | 15% to 40% | CodeSignal |
| APAC vs North America cheating rate | 48% vs 27% | CodeSignal |
| Detected cheaters who still passed | 61% | Fabric HQ |
| US SWE profiles with GitHub listed as a skill | ~17,057 of ~358,701 (4.8%) | Refolk index, derived |
| US ratio of engineers to technical recruiters | ~15:1 | Refolk index, derived |
If you accept that 38% of interviews are corrupted and 61% of corruptions pass, you have to spend recruiter hours earlier in the funnel, not later. The interview cannot be salvaged by adding a second interview. It can only be de-risked by making the first synchronous touch a conversation about verified prior work.
The interview cannot be salvaged by adding a second interview. It can only be de-risked by moving the burden of proof to before the call.
Entry-level is where this hurts most, and where GitHub helps least
Juniors are cheating the most and have the least verifiable history. CodeSignal's entry-level cheating rate nearly tripled year over year, from 15% to 40%, roughly double the shift seen at senior levels. That is the segment where "just check the commits" advice collapses, because a 22-year-old will not have five years of merged PRs.
For juniors, the artifact stack looks different:
- Contribution graphs going back to school years. A dated pattern of green squares from 2022 to 2025 is harder to fake than any live-code answer.
- Hackathon repos with commit history from the event window. Timestamps that match a public event are strong evidence.
- Issue comments in libraries they used in class projects. Real questions, engaged answers, actual reading of source.
- Small utility packages published to npm or PyPI with real, if modest, install counts.
- Fork-and-modify patterns on repos related to their stated interests, showing exploration rather than portfolio theater.
Trained sourcers can read these differently than they read staff-level PRs. The screening question flips from "did this person write great code" to "does this person have a plausible multi-year trail of curiosity in this stack." Refolk's plain-English search is built for exactly that pivot, since you can ask for "recent grads with 18+ months of consistent GitHub activity in Rust" rather than filtering on the wrong signals from a resume.
What to do this quarter
Rip the interview off the critical path for the top-of-funnel decision. Concretely:
- Kill the standalone coding screen for anyone with a verifiable public artifact history. Replace it with a 45-minute conversation about a specific PR they authored. They cannot fake a discussion about code they wrote a year ago.
- For candidates without public artifacts, require a take-home with a Loom of the process, and grade the reasoning in the Loom, not the code. The overlay tools can write the code. They cannot yet fake 20 minutes of coherent narration about tradeoffs.
- Move recruiter effort from scheduling to verification. The 15:1 ratio makes this only viable with automated artifact stitching, so human hours land on the 10 candidates whose artifacts actually check out, not the 100 whose resumes look identical.
- Stop measuring recruiters on interviews scheduled. Start measuring them on verified-artifact shortlists delivered. The metric change is the intervention.
- Assume every unattested claim in a synchronous interview is now compromised. The Checkr data showing 23% of employers lost more than $50K per fraudulent hire is the CFO version of this argument.
The screen was a convenient fiction that a synchronous hour with a stranger predicted job performance. Cluely, Interview Coder, and their invisible-overlay peers have made that fiction untenable. The response is not a better overlay-detector. It is treating the candidate's public code history as the primary evidence and the interview as the confirmation.
FAQ
Is banning AI in interviews still a reasonable policy?
No, not as the primary defense. Ban policies are unenforceable against overlay tools that operate below the screen-share layer, and they contradict the 71% of US tech postings that now require AI fluency. A written expectation is fine, but the real defense is moving the decision-carrying evidence to artifacts that predate the interview, so the ban question stops mattering.
What if the candidate has no public GitHub at all?
Then use a take-home with narrated reasoning (Loom or Tuple recording) and grade the narration, not the code. Many strong engineers work exclusively in private repos at previous employers, which is legitimate, so absence of GitHub is not disqualifying. What is disqualifying is inability to talk coherently about their own decisions for 20 minutes without prompts.
How is sourcing from GitHub different from what LinkedIn Recruiter does?
LinkedIn Recruiter surfaces self-reported claims. Sourcing from GitHub surfaces dated, third-party-verifiable artifacts. The two are complementary, but only the second one is resistant to interview fraud, because the evidence exists in a system the candidate does not control. Refolk stitches both together so you do not have to choose.
Does this mean CodeSignal and HackerRank are dead?
Not dead, but demoted. They remain useful for structured comparison and for candidates who genuinely have no public footprint, especially early-career hires. What has ended is their role as the primary go/no-go gate. In 2026 they are a supporting instrument, and the sourcing-first funnel is the lead.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.