Refolk
September 12, 2026·9 min read

Three AI Screeners, 109 Resumes, 14% Agreement: The Outbound Case

A head-to-head test of GPT-4o, Gemini 2.0, and Grok 3 left 55% of resumes invisible to every model. Why outbound sourcing is the only fix.

ai resume screening accuracyai screening tool consistencyresume screener blind spotsoutbound sourcing vs inboundai recruiting bias
Three AI Screeners, 109 Resumes, 14% Agreement: The Outbound Case

If you run an AI resume screener today, three different models will hand you three different shortlists, and none of them will contain the majority of your qualified candidates. That is what happens when you point ChatGPT-4o, Gemini 2.0 Flash, and Grok 3 at the same 109 resumes and ask each of them, every day, for a top ten.

The Hair Ventures test from May 2025 has been re-circulating hard through 2026, most loudly via recruitment analyst Greg Savage. It deserves the airtime. It quietly rewrites the business case for outbound sourcing.

What the Meta HRBP screen actually found

Three frontier LLMs, given the same 109 anonymized HR Business Partner resumes for a global Meta role and a typical recruiter prompt, agreed on only 14% of their daily top-ten shortlists and never surfaced 55% of the pool at all.

Hair Ventures ran 300 head-to-head screens across ChatGPT-4o, Gemini 2.0 Flash, and Grok 3. The specific findings that matter:

  • 14% shortlist overlap. Two AI recruiters looking at the same CVs differ four times out of five.
  • ±2.5 rank places of drift. The same model, given the same resumes on a different day, reshuffled candidates by an average of 2.5 places.
  • 96% recycled rationales. Rejection explanations reused the same three phrases 96 out of every 100 times.
  • 55% never shortlisted, ever. Roughly 60 of the 109 candidates were invisible to every model on every day of the test.
55%
of resumes never shortlisted by any AI model on any day
From the Hair Ventures May 2025 head-to-head across GPT-4o, Gemini 2.0 Flash, and Grok 3.

A 2026 follow-up from i10X Research made the mechanism explicit. It ran 400 resumes past GPT-5.4, Claude Sonnet 4.6, Gemini 3 Pro, and Grok 4.3, and found Claude approved 42% of GPT-written resumes but 84% of Claude-style ones. Same candidates. Same qualifications. Different score, driven by writing style the model recognized as its own.

Why "ai resume screening accuracy" is the wrong frame

Accuracy assumes there is a right answer the model is approximating. There is not. AI screening is a lottery draw that reshuffles the ticket order every morning, so the right question is not "how accurate?" but "how much of the addressable market am I discarding?"

Recruiters read "AI screener" and picture a filter: noisy inbound goes in, qualified candidates come out. The Hair Ventures numbers say something different. If Model A's top ten and Model B's top ten overlap by only 14%, then at most 1.4 candidates are the "consensus" pick. The other 8.6 are essentially random draws from the pile. Stack a third model for a "second opinion" and the intersection shrinks further. This is the counter-intuitive part: the more AI opinions you stack, the fewer candidates survive, and the survivors correlate with resume style, not talent.

That is ai screening tool consistency in one paragraph. It doesn't exist. It cannot exist, because the models are learning stylistic fingerprints, not competence signals. The i10X 42%-vs-84% split is the smoking gun.

The 55% is a supply problem in disguise

The invisible 55% isn't a filter miscalibration; it is proof that the whole funnel is looking at the wrong pool. When 9,514 US HR Business Partners exist and inbound plus AI surfaces roughly ten of them, you are not selecting talent. You are discarding it.

The Meta HRBP screen tested 109 resumes. That is a rounding error against the real US market for the role.

SliceCountSource
HR Business Partner, all seniorities, US9,514Refolk's index
HR Business Partner, all seniorities, UK9,259Refolk's index
Director HR / VP People / VP HR, US5,916Refolk's index
Resumes in the Hair Ventures study109Hair Ventures
Candidates never shortlisted by any model on any day~60 (55%)Hair Ventures
Ratio of real US HRBP pool to resumes tested~87xDerived
Ratio of invisible candidates to daily consensus picks~39xDerived

The applicant pool the study drew from was already 87 times smaller than the addressable US market for the role. Then the models threw away 55% of what was left. Then the "consensus" filter compressed it to roughly one and a half candidates a day. That is the funnel that companies with 82% AI screening adoption (SHRM 2025) are running.

The mechanical fix is outbound: query the whole market directly instead of ranking the sliver that self-applied. Describing an HRBP you actually want in plain English and getting back a ranked shortlist across LinkedIn, GitHub, and the open web is the exact gap Refolk closes. You are no longer choosing from 109 resumes filtered by three coin-flips. You are choosing from 9,514.

What "resume screener blind spots" really look like in 2026

The blind spot is not a category of candidate the model dislikes. It is a random 55% of the pool that never enters the model's field of view on any given day, refreshed daily, with no audit trail explaining why.

Three overlapping forces are widening this hole through 2026:

  1. Application volume is exploding into the same broken filter. LinkedIn saw a 45% jump in applications year over year, averaging ~11,000 submissions per minute. 64% of recruiters report more look-alike applications.
  2. Applications themselves are increasingly AI-written. 78% of applications now contain AI-generated content, which the i10X study proved is scored by stylistic self-recognition, not merit.
  3. Screener adoption is now default. SHRM 2025 puts AI resume-screening adoption at 82% of companies for resume review and 51% specifically for recruiting. AI recruiting adoption doubled from 26% to 43% of HR teams between 2024 and 2025.

The compounding effect is grim. AI-written applications flow into AI screeners that reward writing that looks like the screener's own output. Signal-to-noise inside the inbound channel is decaying, not improving. Outbound sourcing vs inbound is not a stylistic preference anymore. It is the difference between reaching the market and sampling one algorithm's fingerprint of it.

The more AI opinions you stack, the fewer candidates survive, and the survivors correlate with resume style, not talent.

The compliance time bomb nobody is pricing in

The 96%-recycled-rationale finding is not just embarrassing. Under GDPR Article 22 and the EU AI Act's August 2026 deadline, it is arguably illegal. And it puts the entire inbound-plus-AI workflow on borrowed time.

Two rules matter here:

  • GDPR Article 22 already forbids solely automated decisions with significant effects on a person. A screener that auto-rejects 55% of a pool without human review is on shaky ground today.
  • The EU AI Act classifies recruitment screening as high-risk (Annex III), with full compliance required from August 2026 and penalties up to €30M or 6% of global turnover.

If your rejection rationale is one of three phrases pasted onto 96% of candidates, that is not a defensible human-in-the-loop record. It is an audit finding waiting to happen. Outbound sourcing sidesteps this entirely because there is no automated rejection to justify. You reach out to people you chose. Nobody was rejected by a machine.

Deloitte's 2026 Global Human Capital Trends and Harvard Business Review's June 2026 analysis of 120 TA leaders and 6,300 recorded screening sessions are both catching up to what recruiters already suspected. The compliance risk of ai recruiting bias is now a boardroom problem, not a hiring problem.

What outbound sourcing vs inbound actually looks like at the top of the funnel

Outbound sourcing means building a candidate list from the whole market and reaching out, instead of ranking whoever applied. In the Meta HRBP case, that is the difference between choosing from 109 and choosing from 9,514.

The practical playbook for a role like the one the study used:

  • Start from the market, not the ATS. For HRBPs, the US pool is ~9,514 and the UK pool is ~9,259. Your inbound is a sample of maybe 0.5% of that.
  • Segment by employer type, not just title. HRBP at HCA Healthcare (regulated, high-volume ops) is not the same job as HRBP at Intel (tech, matrixed) or dunnhumby (analytics-forward). BAE Systems Digital Intelligence, Samsonite, Macquarie Group, and ManpowerGroup all employ people with the same title doing very different work.
  • Pull the leadership slice separately. ~5,916 US professionals hold Director HR, VP People, or VP HR titles. These people almost never appear in an inbound funnel because they are employed and not applying.
  • Skip the model-consensus trap. Do not run three screeners and take the intersection. The intersection is 14%. Instead, use one model to enrich a sourced list, then have a human read the top 30.
87x
real US HRBP market vs. the resumes the LLMs were tested on
9,514 identifiable HRBPs in Refolk's index vs. the 109 anonymized resumes in the Hair Ventures screen.

The HRBP number is instructive because the role is not especially rare. For scarcer roles, the blind-spot math gets worse. In a US pool of 5,916 HR Directors and VPs, if only ten apply per opening and three models agree on 14% of them, you are effectively hiring from a shortlist of one or two candidates for a job with thousands of qualified holders. That is the number that should worry any engineering leader making a senior people hire this quarter.

Where to start this week

Start by counting: how many resumes did your last hire draw from, and how large is the actual market for that role? If the ratio is worse than 1:100, your funnel is not a filter, it is a lottery.

Then pick one open req and run it two ways:

  • Inbound-only. Whatever came through the ATS, ranked however you rank it.
  • Outbound-first. Describe the person you actually want in plain English, pull 40 to 60 names from the full market, and let a human read the top 20.

Compare the two lists in a week. The overlap will be small. That is the point. The names on the outbound list are the ones your AI screener was never going to see, because they never applied. That is the invisible 55%, plus the much larger slice who do not show up in your ATS at all.

FAQ

Isn't a 14% overlap fine if the models are all pulling from qualified candidates?

No. Overlap that low means the models are not converging on quality signals, they are drifting on stylistic ones. The i10X study proved it: Claude approved 42% of GPT-written resumes and 84% of Claude-style ones. Same candidates, different score. If the "agreed" 14% is really the intersection of three style biases, you are hiring from a filter that rewards resume writing, not competence.

Does using more models fix the consistency problem?

It makes it worse. Stacking screeners for a "second opinion" shrinks your candidate pool to the intersection. With 14% pairwise overlap, adding a third model roughly halves the survivors again. The candidates who survive multi-model consensus are the ones whose resumes happen to match every model's stylistic priors, which correlates weakly (or negatively) with the job.

Is outbound sourcing actually faster than inbound plus AI screening for senior roles?

For senior and specialist roles, the market math forces the answer. The best candidates for a senior HR role are almost never in your applicant pool of ten to twenty; they are among the 5,916 US HR Directors and VPs who are employed and not applying. Reaching them requires querying the market directly rather than ranking whoever self-selected in.

What should I do about the EU AI Act deadline in August 2026?

Two things. First, document human review on every rejection, especially if you use AI to rank or filter. Recycled rationale phrases across 96% of rejections will not survive an audit. Second, shift senior and specialist hiring toward outbound wherever you can. There is no automated rejection to justify when you chose to reach out to a specific person, which removes the highest-risk part of the workflow entirely.

Try it on the search you came here for

Stop building boolean strings. Just describe the person.

Type one sentence. I plan the search, read GitHub, public LinkedIn and Crunchbase records, and the open web as it is right now, and hand back a ranked list with the reason next to every name.

  1. 01Describe them

    One plain sentence. Role, city, stack, stage, whatever matters to you.

  2. 02I read the web live

    GitHub, public LinkedIn and Crunchbase records, the open web. Not a database that went stale last quarter.

  3. 03You read the shortlist

    Ranked, with the reasoning under every name. Open a profile, ask a follow-up, narrow it down.

  • No boolean, no filters, no seat to buy. One box.
  • Read at search time, so a profile updated yesterday counts today.
  • Every step visible as it runs, every name with its reason.

500 free credits on sign-up. No card, no demo call. See real searches.

Read next