Three AI Screeners Agreed on 14%. The Other 55% Are Your Pipeline.
Three frontier LLMs shortlisted the same candidates only 14% of the time and buried 55% entirely. Here is why outbound sourcing is the only fix.
If you run inbound at any scale in 2026, a study Greg Savage recirculated on September 14 should stop your quarter cold. Three frontier LLMs, given the same 109 resumes for the same role, agreed on 14% of their top-10 shortlists and never surfaced 55% of the candidates on any day. Your ATS is not filtering. It is flipping coins.
What the Eunomia study actually measured
Eunomia HR ran 300 head-to-head resume screens in May 2025 using ChatGPT-4o, Gemini 2.0 Flash, and Grok 3 against 109 anonymized HR Business Partner resumes, sourced for a global Meta role and evaluated with a typical recruiter prompt. The setup was deliberately generous to the models: real candidates, real seniority, real job spec, no adversarial tricks.
The findings, now cited by Recruiterflow's 2026 trends report and driving fresh discussion on Savage's LinkedIn feed:
- 14% shortlist overlap across the three models on the same resumes, roughly half the human inter-rater reliability band (Cohen's k around 0.49), with twice the volatility.
- 55% of candidates were never shortlisted by any model on any day. Not "ranked low". Invisible.
- Rankings reshuffled by around ±2.5 places day to day on the same model with the same input.
- Rationale bullets recycled the same three phrases 96% of the time, evidence the models were pattern-matching surface features rather than reading the resume.
An independent arXiv paper ("Fairness Is Not Enough", 2507.11548) found Gemini's slow model rated the same underqualified resume nearly 30 points lower on average than its fast sibling. The instability is not a single-vendor problem. It is the class.
Why "AI screening reliability" is a worse story than the ATS era
AI resume screening reliability is genuinely lower than the keyword ATS it replaced, because legacy ATS failed deterministically and LLMs fail stochastically. A boolean filter that rejects your candidate at 9 a.m. rejects them at 3 p.m. An LLM that rejects them at 9 a.m. might shortlist them at 3 p.m., and vice versa.
That has three consequences most hiring leaders have not internalized:
- AI shortlist consistency degrades exactly where inbound volume is highest. Recruiterflow reports 82% of firms will depend on AI to sift resumes by 2026, applied to a candidate corpus that is itself increasingly LLM-generated. Homogenous inputs plus stochastic scoring compounds the ranking-roulette effect.
- Two recruiters running the same tool interview different people. The ±2.5 rank drift means Monday's top 10 is not Tuesday's top 10. Your calibration meetings are arguing about noise.
- The 55% ghosts are not weak candidates. The Eunomia pool was pre-qualified HRBPs applying for a Meta role. The mechanism is that LLMs anchor on a small set of surface phrases (the recycled 96%) and anyone whose resume does not trigger those anchors gets zero visibility across every model.
The algorithmic blind spot in hiring is not the tail of the distribution. It is the middle.
Two recruiters running the same screener at 9 and 3 will interview different people. Your shortlist is a coin flip with a UI.
The invisible 55%, sized against a real talent pool
Applied to any real market, the 55% figure describes a majority of your qualified pipeline that no inbound funnel will ever surface. Refolk's index makes the scale concrete for the exact role Eunomia tested.
| Role | Geo | Pool (Refolk index) | Invisible layer at 55% | Note |
|---|---|---|---|---|
| HR Business Partner | US | 9,601 | ~5,281 | Never shortlisted by any model on any day |
| HR Business Partner | UK | 9,348 | ~5,141 | Near-identical pool; volatility not scarcity is the issue |
| Technical / Talent Sourcer | US | 1,251 | n/a | Capacity to do outbound |
| HRBP ÷ US sourcing pros | US | 7.7x | n/a | Demand vs outbound-capacity gap |
| HRBP US vs UK delta | cross | +2.7% | n/a | Pools essentially match |
In Refolk's index of professional profiles, there are 9,601 current HR Business Partners in the US and 9,348 in the UK. If the Eunomia result generalizes, a US employer running the same LLM stack has roughly 5,281 qualified HRBPs it will never see, regardless of how long the job stays open or how much it spends on LinkedIn promotion. The UK number is 5,141. The top employers of these people (HCA Healthcare, American Credit Acceptance, Maury Regional Health in the US; ManpowerGroup, dunnhumby, BAE Systems Digital Intelligence, Macquarie Group in the UK) are not obscure. They are simply not on the anchor list your screener recognizes.
The demand-capacity math is worse. The US pool of dedicated sourcing specialists (Technical Sourcer, Talent Sourcer, Sourcing Recruiter) is 1,251. That is 7.7x smaller than the HRBP pool the screeners are miscounting. The profession trained specifically to reach the invisible layer is the smaller number. Most firms do not have the internal headcount to close the gap manually, which is the outbound-is-the-only-way argument rendered in profile counts.
Outbound sourcing vs AI screening, as a channel
Outbound sourcing beats AI screening not because sourcers read faster but because outbound bypasses the anchor phrases the LLM is fixated on. The screener looks at 109 resumes and sees three clichés recycled 96% of the time. A sourcer starts from the population and works backward to the person.
That inversion is the point. Inbound plus AI is a filter over a self-selected, LLM-tuned resume corpus. Outbound is a query over the whole professional graph, including people who never applied, never wrote a "keyword-optimized" resume, and never triggered your ATS's anchor list. In practice this is where Refolk fits: you describe the person you want in plain English ("senior HRBP with M&A integration experience in US healthcare, currently at a system with 40k+ employees") and get a ranked shortlist drawn from the full professional graph. No boolean, no anchor phrases, no ranking drift between Monday and Tuesday.
The legal exposure sits with you
The EEOC's 2023 technical assistance on AI confirmed employers are liable for adverse impact caused by third-party AI tools. "The vendor's tool did it" is not a Title VII defense. Two live cases sharpen this:
- EEOC v. iTutorGroup (August 2023) was the first AI hiring discrimination settlement: $365,000 paid to more than 200 candidates auto-rejected by age-filtering software.
- Mobley v. Workday was conditionally certified in May 2025 by Judge Rita Lin (NDCA) as an ADEA collective on behalf of all applicants 40+ rejected by Workday's AI screening since 2020. The largest AI-hiring action in US history.
Now overlay ±2.5 rank drift on that legal posture. The same candidate can be rejected Monday and shortlisted Tuesday, which is exactly the fact pattern a disparate-impact plaintiff needs. Eunomia flags this directly: invisible disqualifiers raise EU AI Act and GDPR risk, and volatile, unexplainable decisions are hard to defend.
Why AI screening false negatives compound in 2026
AI screening false negatives compound because the resume corpus itself is now machine-generated, and the screener anchors on the same phrases the generator produces. In 2024 an LLM screener at least saw variance in resume prose. In 2026, with 82% of firms sifting with AI and, per Recruiterflow, roughly 40% of tech candidates believed to be inflating resumes with LLM assistance, the input and the filter are trained on overlapping distributions.
The downstream numbers are already visible:
- 65% of business leaders say AI rejects candidates before humans have seen the application (HRExecutive, August 2026).
- 14% say the tech rejects more than half of all applicants before a human sees them.
- 53% of job seekers reported being ghosted in the last year, up from 48% in 2025 and 38% in 2024 (Criteria, reported in Fortune).
Ghosting used to be a courtesy failure. It is now a measurement of how much of the funnel is disappearing into false negatives. And ghosted candidates do not reapply. Every quarter your screener runs at 14% agreement, the addressable pool for your next req shrinks by the people you silently rejected last time.
Outbound is the only channel where inputs are not machine-optimized to match the filter. That is not an aesthetic preference. It is the only place the invisible 55% still lives.
What to change on Monday
Do not rip out the LLM. Move it to the part of the process it is genuinely good at, and route hiring intent through outbound. Eunomia's own defense case is real: LLMs screen 109 CVs in under a minute, enforce template discipline, and draft acceptable first bullets. That is a triage tool, not a decision tool.
A workable split:
- Use LLMs for speed blitzes and template discipline, not for shortlist ranking. Cap their authority at "surface" versus "hide", not "hire" versus "reject".
- Require a second, deterministic pass (structured rubric, human calibration) on any inbound the LLM elevated. This is what keeps the ±2.5 drift out of your interview loop.
- Fund outbound to cover the 55%. For an HRBP req in the US that means budgeting sourcing capacity against a 5,281-person invisible layer, not against the applicant count in your ATS.
- Log the prompt, the model, the version, and the rationale bullets for every AI-touched decision. If the rationale recycles across 96% of candidates, your audit trail is already telling you the tool is not reading.
- Route senior and hard-to-fill roles through named-list outbound first. For roles like HRBP where the pool is under 10,000 in a country, the cost of missing the median candidate is higher than the cost of a sourced conversation.
The companies that have already institutionalized this pattern are visible in Refolk's index: Pinterest, Anthropic, Rippling, Verkada, and Gartner employ the largest US sourcing benches. They are not avoiding AI. They are refusing to let it be the gate.
FAQ
Does the 14% overlap number apply to my ATS's AI screener?
The Eunomia study tested general-purpose frontier LLMs (ChatGPT-4o, Gemini 2.0 Flash, Grok 3) with a recruiter-style prompt, which is the same class of model most ATS "AI screening" features wrap. Purpose-built vendors may perform better on specific dimensions, but the arXiv "Fairness Is Not Enough" paper shows ranking volatility across model families, not just consumer chatbots. Until your vendor publishes head-to-head reliability numbers on your resumes, assume the 14% ceiling is a reasonable prior.
Isn't outbound sourcing just as biased as AI screening?
Outbound has different failure modes, but it does not silently delete 55% of your pool the way stochastic LLM ranking does. A sourcer's misses are visible in the shortlist and reviewable against a rubric. An LLM's misses never appear anywhere, which is the exact problem the EEOC and the plaintiffs in Mobley v. Workday are attacking. Outbound also draws from the full professional graph, not the self-selected inbound sample, so the base rate of qualified candidates is higher before any judgment is applied.
How big does a role's pool have to be before outbound is worth it?
Any role where the invisible 55% is larger than the number of candidates you can reasonably interview. For an HRBP req in the US that is a ~5,281-person layer against a ~9,601-person pool, and outbound pays back on the first hire. For very small pools of a few hundred qualified profiles country-wide, outbound is not a strategy choice, it is the only channel that works because inbound will not produce a viable candidate at all.
What should I tell my hiring managers who love the AI screener's speed?
Keep the speed, move the authority. Frame the LLM as a triage layer that flags obviously off-target applications and drafts template bullets, and make the shortlist decision on a structured human rubric with sourced candidates included. That preserves the 109-CVs-per-minute throughput Eunomia validated, removes the ±2.5 ranking drift from your interview loop, and gives your legal team an audit trail that will survive a disparate-impact claim.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.