The 14% Overlap Test: Why AI Shortlists Fail Reproducibility
Same AI, same candidates, 14% shortlist overlap. What the reproducibility crisis in AI screening means for TA procurement in 2026, and how to test it.
Greg Savage dropped a number in February 2026 that should have ended a lot of vendor demos: run the same AI screening tool twice on the same candidates, and the shortlists overlap by 14%. Recruiterflow's 2026 trend report picked it up. Deloitte's 2026 Human Capital Trends survey of 9,000+ leaders across 89 countries piled on with fake resumes, deepfake interviews, and "workslop." The reproducibility gap is now a procurement question, not a philosophy one.
What the 14% overlap number actually means
Fourteen percent overlap means the same AI screening tool, given the same job description and the same candidate pool twice in a row, agreed with itself on roughly one in seven finalists. That is not a rounding error. It is a tool that fails the test-retest reliability bar any occupational psychologist would demand of a pre-employment assessment before it goes near a hiring decision.
The figure surfaced on a Savage podcast episode logged at timestamp "06:34 AI shortlisting chaos and the 14% overlap test" and was folded into Recruiterflow's 2026 recruitment trends writeup. It is the first hard, cite-able number attached to a problem TA teams have been muttering about for two years: AI candidate shortlists that look different every time you press the button.
The mechanism is not mysterious. LLM-based scorers use temperature sampling and stochastic embeddings. Running the same JD against the same 500 resumes twice invokes different token paths, different nearest-neighbor tiebreaks, and different chain-of-thought branches. Determinism was traded for fluency somewhere in the stack, and nobody told the buyer.
Why this is a reliability problem, not an accuracy problem
AI screening tool accuracy is the wrong frame. The right frame is reliability engineering: what is the seed-to-seed variance on the top 25? If a vendor cannot answer that question, they do not know either.
Recruiters intuitively grasp this. Aptitude Research found that 85% of recruiters want final decision authority over AI recommendations. That is not Luddism. It is a rational response to a scoring system that will not pass a test-retest check. The industry has spent a decade telling recruiters to trust the model. The model has spent 2026 telling recruiters it does not trust itself.
Here is the practical translation, using only numbers from published research:
- 69% of companies now use AI somewhere in talent acquisition (Aptitude Research / iCIMS).
- Same-tool shortlists overlap 14% run-to-run (Savage / Recruiterflow).
- That implies roughly 59% of AI-screened funnels (0.69 x 0.86) are producing candidate lists that would change materially on a re-run.
Six in ten AI-mediated hiring funnels in 2026 are, in effect, a coin flip with better UX.
The comparable numbers, in one table
The reproducibility crisis is not one number. It is a stack of them that only makes sense together. Here is the dataset a TA leader should carry into the next vendor call.
| Metric | Value | Source |
|---|---|---|
| Shortlist overlap, same AI tool, identical data, run twice | 14% | Savage / Recruiterflow 2026 |
| Executives worried about candidate data accuracy | 95% | Deloitte 2026 HC Trends |
| Organizations making significant progress fixing that data | 5% | Deloitte 2026 HC Trends |
| Concern-to-action gap (derived) | 19x | 95 / 5 |
| Companies using AI in talent acquisition | 69% | Aptitude Research / iCIMS |
| Candidates who trust AI to evaluate fairly | 26% | Gartner |
| Tech candidates believed to have meaningfully inflated resumes | ~40% | Recruiterflow 2026 |
| Talent leaders planning to use AI in recruiting in 2026 | 84% | Korn Ferry, n=1,670+ |
| Recruiters wanting final decision authority over AI | 85% | Aptitude Research |
Two rows do most of the work. Ninety-five percent of executives are worried about candidate data accuracy. Five percent of organizations are making meaningful progress on it. That is a 19x concern-to-action ratio, and it is the market opening for any vendor that leads with reproducibility guarantees instead of another accuracy claim.
Sourcing drift is almost certainly worse than screening drift
If bounded-pool screening drifts 86%, open-web sourcing drifts more. Screening runs against a fixed set of applicants who already applied. Sourcing runs against an index of the open web that shifts daily, on top of the same stochastic scoring layer. Nobody has published the number yet. It will not be flattering.
This is the part of the AI sourcing reliability story that vendors quietly do not want tested. Three drift sources compound:
- Index drift. New profiles appear, old ones update, dead links get pruned. The candidate universe on Tuesday is not the candidate universe on Wednesday.
- Query drift. Natural-language queries get re-interpreted by the LLM each run. "Senior backend engineer with payments experience" resolves to slightly different embeddings each time.
- Ranking drift. The final scoring pass adds its own temperature-sampled noise.
Any one of these is survivable. Stacked, they turn "find me 25 candidates" into a lottery whose ticket you re-buy every session. The honest fix is to expose the drift, log the seed, and let the operator lock a run. That is the discipline behind how Refolk treats a plain-English query: the same prompt returns a stable, ranked shortlist you can re-open, not a fresh dice roll.
Six in ten AI-mediated hiring funnels in 2026 are a coin flip with better UX.
The AI-washing trap: deterministic does not mean intelligent
The scariest second-order effect: if you buy on "consistency," you may accidentally select for the dumbest tool in the room. Regex from 1987 is perfectly reproducible. It is also useless.
Florian Fisch, writing on Savage's blog in March 2026, put it plainly: agencies globally are paying premium prices for "AI-powered" tools using technology from the 1990s. If a system cannot match "Led development teams" to "Team Leader," it is not AI. It is a badge.
The uncomfortable implication for anyone running a bake-off:
- A real LLM-based tool will fail the 14% overlap test hard, because it is doing real inference under temperature.
- A regex-with-lipstick tool will pass it easily, because deterministic string matching is trivially reproducible.
- Buyers who rank on "consistency" alone will pick the regex.
The fix is a two-part test. Measure reproducibility (does the top-25 stabilize?) and measure recall against a known-good set of candidates you already hired or interviewed. Any tool that passes one but flunks the other is telling you what it actually is.
The candidate side got worse at the same time
Deloitte's 2026 report warns that generative AI now produces fake resumes, deepfake interviews, and "workslop" at scale. Roughly 40% of tech candidates are believed to have meaningfully inflated their resumes (Recruiterflow 2026). AI writing tools have collapsed the cost of applying to near zero, and candidates responded rationally: they apply to everything, with machine-tailored documents for every posting.
So the funnel now looks like this:
- Inbound volume: inflated by generative AI on the candidate side.
- Signal quality per resume: degraded by the same tools.
- Screening layer: non-reproducible by 86%.
- Human reviewer: told to trust the machine.
This is why so many teams are quietly moving budget from inbound screening to outbound sourcing. When the applicant pool is polluted, the reliable move is to go pick the people you want, from primary evidence rather than self-reported PDFs. That is the exact gap Refolk closes: describe the person in plain English, get a ranked shortlist built from GitHub, LinkedIn, and the open web, with the underlying evidence attached to each hit so you can audit why someone ranked where they did.
What to actually test before you sign
Run three tests, in this order, before any AI screening or sourcing tool gets past procurement. This is the minimum bar a serious TA org should adopt in 2026.
1. The 14% overlap test, but on your data
Pick a real requisition. Feed the tool 300 to 500 candidates. Run it twice, an hour apart, same JD, same pool. Compute Jaccard overlap on the top 25. Anything under 70% is a red flag. Anything under 30% means the vendor is selling you noise with a UI.
2. The known-good recall test
Take 10 people you actually hired for similar roles in the last 18 months. Drop them into a pool of 500 plausible-but-wrong candidates. See how many the tool surfaces in the top 25. If it misses more than three of your known-good hires, it does not understand your bar.
3. The audit trail test
Ask the vendor to show you, for a single candidate, exactly which signals drove the ranking. If the answer is "the model decided," you cannot defend the tool under NYC Local Law 144's bias audit requirement, and you will not survive the EU AI Act's August 2026 transparency rules either (high-risk employment rules land December 2026). Reproducibility failures may now trigger audit findings, not just bad hires.
Refolk's index shows only 482 senior TA leaders in the US at Director, VP, or CXO level with Head of Talent, TA, or Recruiting Ops titles. Five in San Francisco, four in NYC, two in Boston. That is the entire population being asked to green-light AI screening across the American tech economy. A group that small can, and should, standardize a procurement test.
Where the market goes from here
The vendors that survive 2027 will publish their reproducibility numbers the way SaaS companies publish uptime. Everyone else will get quietly cut at renewal.
Three shifts are already visible:
- Reproducibility SLAs. Expect "same query, same shortlist" guarantees, with logged seeds and re-openable runs. This is the 19x concern-to-action gap collapsing.
- Evidence-linked ranking. Scores nobody can explain will lose to scores with GitHub commits, verified employers, and open-web citations attached.
- Regulatory forcing functions. NYC Local Law 144 is live. The EU AI Act's transparency rules land August 2026, high-risk employment rules December 2026. Non-reproducible tools will fail audits before they fail hiring managers.
Korn Ferry, using AI on its own workflow in 2024, reported a 50% increase in sourcing and a 66% decline in time-to-interview. The search firm Andiamo reported 4x revenue and a 76% reduction in time-to-fill after rebuilding workflows around Recruiterflow and AIRA. AI works. Non-reproducible AI does not. The distinction is now measurable, and 2026 is the year buyers started measuring it.
FAQ
Is the 14% overlap number a benchmark or a one-off?
It was surfaced by Greg Savage in February 2026 and reflected in Recruiterflow's 2026 trend report, based on running the same AI screening tool twice on identical candidate data. It is not yet an industry-wide replicated benchmark, and no vendor has published a competing number. That absence is itself informative: if any vendor had a better figure, they would have posted it by now.
Does non-determinism mean AI screening is useless?
No. It means AI screening is a probabilistic layer that needs the same reliability discipline as any other statistical tool. Run it multiple times, look at the top-25 intersection, and combine model outputs with recruiter judgment. The 85% of recruiters who want final decision authority are not wrong; they are running an implicit ensemble over a noisy scorer.
How is sourcing different from screening on reproducibility?
Sourcing queries the open web or a professional index, so the underlying candidate universe changes between runs, on top of any model-level nondeterminism. That should make sourcing worse on paper. In practice, a well-built sourcing tool logs the index snapshot and the query, so a re-run returns the same ranked list. Ask any sourcing vendor to demo a re-opened run. If they cannot, you have your answer.
What is the single fastest test to run on a vendor this quarter?
The Jaccard overlap test on top-25 shortlists, using one real requisition and 300 to 500 candidates, twice, one hour apart. It takes an afternoon, produces a defensible number, and separates real ML tools with reproducibility engineering from AI-washed regex and unmanaged LLM wrappers. If the vendor will not let you run it, that is the result.
Try it on your own search
Stop building boolean strings. Just describe the person.
Type one sentence and I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web live, then hand back a ranked shortlist with the reasoning behind every name. No filters to learn, no export to clean up, no sales call to sit through.
- One sentence in, a ranked shortlist out. No boolean, no filters, no seat to buy.
- Read live at search time, not from a database that went stale last quarter.
- Watch every step as it runs, and see why each name made the list.
- Staff backend engineers in NYC who shipped Rust in production
- Series A fintechs in SF under 50 people, growing headcount this year
- Maintainers of fast-growing Rust web frameworks on GitHub
500 free credits on sign-up. No card, no demo call. See real searches.