14% Overlap, 55% Invisible: The Screener Math That Kills Screening
The same AI resume screener run twice returns only 14% overlap. Here is why sourcing a named shortlist is now the only defensible workflow.
If you run the same AI resume screener twice on the same 109 candidates, you get two shortlists that agree on 14% of names. That is the finding Greg Savage anchored his February 2026 podcast around, and it is what Recruiterflow's July 2026 industry roundup put at the top of its AI risk section. The number is not a warning about tuning. It is a verdict on the workflow itself.
The 14% number is worse than it looks
Fourteen percent overlap on a top-10 shortlist drawn from 109 résumés is barely above the roughly 9% you would expect from random chance. The screener is doing coin-flip work while charging for judgment.
The primary source is a May 2025 study by hair.ventures: 300 head-to-head résumé screens using ChatGPT-4o, Gemini 2.0 Flash, and Grok 3 on 109 anonymized HR Business Partner résumés for a global Meta role. The study measured overlap across models and across repeated runs of the same model on the same inputs. The mechanism is not mysterious. Any temperature above zero, plus prompt sensitivity in the ranking step, means "identical inputs" is a fiction once the tokens hit the model. Two runs are two experiments.
For contrast, human recruiter inter-rater reliability sits around κ ≈ 0.49. Humans disagree plenty. But the 14% LLM overlap is roughly half the human κ band with twice the volatility, and unlike humans, the model cannot explain why it changed its mind between 9:14 and 9:17.
The hidden 55% is the real scandal
55% of the résumés in the talent pool were never shortlisted by any model, on any day. More than half the population vanished into an algorithmic blind spot no recruiter ever saw. The reproducibility number gets the headlines. The exclusion number is what should end the category.
This inverts the usual framing of "AI resume screening accuracy". The dominant failure mode is not that the wrong ten candidates get ranked in the wrong order. It is that a majority of the candidate pool is silently deleted from consideration before any human is asked a question. If you buy a screener to save time on ranking, you are paying to have the exclusion step done in a way you cannot audit.
Sourcing inverts this by construction. When you start from a named list of the people you want, the invisible 55% problem disappears, because you are not filtering a pile, you are describing a person and going to find them. That is the exact gap Refolk closes: you describe the candidate in plain English and get a ranked shortlist across GitHub, LinkedIn, and the open web, with the names attached so a human can see who was and was not considered.
Model rationales are recycled slop
The screeners are not just inconsistent, they are shallow. In the hair.ventures study, rationale bullets recycled the same three phrases 96% of the time.
That is the deeper tell. If the model's own explanations for why it picked a candidate are copy-pasted across 96 out of 100 decisions, the "reasoning" is decorative. It exists to satisfy an ATS field, not to justify a hiring choice. Add the intra-model instability - individual models reshuffled identical résumés by an average of ±2.5 rank places between runs - and you have a system that is fluent, fast, and produces text that looks like a defensible decision without being one.
The rationales are decorative. They exist to satisfy an ATS field, not to justify a hiring choice.
Florian Fisch, ex-BMW recruiter, put the polite version in his Savage Truth guest post: agencies globally are paying premium prices for "AI-powered" tools that are, in practice, regular expressions from 1987 with a chat interface glued on top. His term for it, "AI-washing", is the right frame for what most of the current screening stack ships, from HireVue and Paradox to the AI features shipped inside Workday, Greenhouse, and iCIMS.
Adoption is scaling faster than reliability
AI resume screening adoption roughly doubled in a year, from 26% to 43% of HR teams between 2024 and 2025 per SHRM data, while the underlying reliability numbers did not move. That gap is the story.
Zoom out and it gets sharper. 69% of companies now use AI in talent acquisition in some form, but only 18% deploy it broadly across hiring workflows. Most buyers are still in pilot. The screening budget has not calcified yet, which means the argument for a different architecture - precision sourcing over stochastic screening - is still winnable inside most companies this quarter.
| Metric | Figure | Source |
|---|---|---|
| Shortlist overlap, same tool, identical inputs | 14% | hair.ventures, May 2025 |
| Résumés never shortlisted by any model | 55% | hair.ventures, May 2025 |
| Rationale phrases recycled across decisions | 96% | hair.ventures, May 2025 |
| Intra-model rank drift, same résumé, different day | ±2.5 places | hair.ventures, May 2025 |
| AI screening adoption, 2024 to 2025 | 26% to 43% | SHRM |
| Companies using AI in TA, any form | 69% | 2026 HR guide |
| Companies with broad AI deployment in hiring | 18% | 2026 HR guide |
| Recruiters wanting final authority over AI picks | 85% | Aptitude Research |
Legal exposure is asymmetric
A 14% reproducibility rate means that in any EEOC, EU AI Act, or GDPR challenge, the employer literally cannot reproduce their own hiring decision. That is not a compliance risk. That is the absence of a defense.
The 2024-25 FAIRE academic study found measurable disparate impact across protected characteristics in multiple commercial screening systems, consistent enough to shift shortlist composition at scale. One GDPR complaint at a European staffing firm surfaced a system that had been rejecting applicants with no human review trail at all. When the regulator asked for the reasoning, the answer was a rationale bullet that appeared on 96% of decisions.
Named-shortlist sourcing produces a very different artifact. The record says "a human evaluated these 25 named candidates on these criteria on this date." The record does not say "the model preferred them, twice, and disagreed with itself in between." One of those is auditable. The other is a lawsuit exhibit.
Where AI screening fails hardest is exactly your hiring plan
The screeners underperform most on senior roles, niche specialisms, international candidates, and anyone who has made a non-linear career move. That is not a corner case. That is the population every serious hiring plan is built around.
Precision sourcing is the term for starting from a description of the person you want and returning a named shortlist, rather than starting from an applicant pile and filtering down. It is the workflow retained and executive search have always used, and it is the workflow Savage's Feb 2026 piece flagged as structurally positioned to grow. His argument is that the 14% finding takes the senior-end pathology (screening never worked for VPs) and generalizes it to the whole market.
The AI sourcing vs screening choice is not two flavors of the same thing. They point in opposite directions:
- Screening starts with a pool of self-selected applicants and asks a model to rank them. The failure mode is silent exclusion.
- Sourcing starts with a description of the ideal candidate and asks a system to find them by name. The failure mode is a shortlist you can argue with.
- Screening produces a rationale bullet. Sourcing produces a person, an employer, and a link.
- Screening compounds candidate screening bias inside the model. Sourcing exposes the criteria to the recruiter before any names are ranked.
The market already voted with headcount
The clearest evidence that sourcing is where humans still add value is where companies put humans. In Refolk's index of professional profiles, there are roughly 12 dedicated Sourcers in the US for every 1 Recruiting Operations person.
Here is the shape of that market, from Refolk's index.
| Segment | Count | Note |
|---|---|---|
| Sourcers (Technical / Talent Sourcer), US | 1,745 | Current-title match |
| Sourcers, UK + DE + CA combined | 269 | US market is roughly 6.5x larger |
| Recruiting Operations / TA Ops, US | 143 | Sourcing headcount is roughly 12x ops headcount |
| Sourcers concentrated in SF Bay Area | ~20% of top-25 sample | Densest single hub |
The named employers in the US top-10 of Refolk's sourcer index include Rippling, MongoDB, Anthropic, Verkada, Gartner, Zipline, EvolutionIQ, Fetch, Church's Texas Chicken, and Mission Resourcing. The UK/DE/CA top-10 includes Meta, Deloitte, Siemens, Cohere, Apple, eBay, Planet, NHS, AMS, and Crowe Watson. Notice who is on both lists. Anthropic and Cohere sell frontier models. Meta ships one of the most cited open-weight families in the industry. They still pay humans to name candidates. That is the market's revealed preference.
What to do on Monday
Stop buying reproducibility from a system that cannot deliver it, and start buying names. Concretely:
- Freeze new screener spend for one quarter. You are in the 82% of companies still piloting. Use the pause to run the comparison honestly.
- Run the hair.ventures test on your own stack. Feed the same 100 résumés through your screener twice and measure the overlap. If it is under 30%, you have a stochastic decision system in production.
- Move the budget to sourcing. 85% of recruiters already want to keep final authority over AI recommendations. Give them tools that produce named shortlists instead of ranked piles.
- Write down the criteria before you see the names. This is what makes a sourcing shortlist auditable and a screener rationale not.
- Keep the human in the "who" step, not the "how many" step. Screening automates the wrong step. Sourcing automates the search, then hands humans a list they can defend.
FAQ
Is the 14% overlap number specific to one tool or general?
It came from a head-to-head of three frontier models (ChatGPT-4o, Gemini 2.0 Flash, Grok 3) on 109 résumés, so it is not a single-vendor artifact. The underlying cause - non-zero temperature, prompt sensitivity, and shallow rationale generation - is common to every commercial LLM-based screener currently shipping. If your vendor claims materially better reproducibility, ask them to publish the same test design.
How is precision sourcing different from an AI screener with better prompts?
Precision sourcing starts from a description of the person you want and returns named candidates from the open web, GitHub, and LinkedIn. Screening starts from a pile of applicants and ranks them. The difference matters because the 55% "invisible" failure mode in screening happens before ranking, so no prompt tweak fixes it. Sourcing avoids the exclusion problem by construction.
Does this mean AI has no role in hiring?
AI has a large role in the search step, where the job is "find candidates matching this description across the open web" and the output is a list a human then argues with. It has a bad role in the exclusion step, where the job is "silently decide which applicants a recruiter never sees" and the output is a rationale bullet that appears on 96% of decisions. The distinction is not AI vs no AI. It is sourcing vs screening.
What does a defensible audit trail actually look like?
Written criteria dated before the search, a named shortlist returned against those criteria, a human sign-off on which names advance, and links to the public evidence (GitHub, publications, LinkedIn) that justified each pick. Screener output rarely produces any of these artifacts in a form a regulator will accept. A sourcing workflow produces all four as a byproduct of doing the work.
Try it on the search you came here for
Stop building boolean strings. Just describe the person.
Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.
01Describe them
One plain sentence. Role, city, stack, stage, whatever matters to you.
02I read the web live
GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.
03You read the shortlist
Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
- No boolean, no filters, no seat to buy. One box.
- Read at search time, so a profile updated yesterday counts today.
- Every step visible as it runs, every name with its reason.
500 free credits on sign-up. No card, no demo call. See real searches.