If your resume feels like it disappears into a void, the Hair Ventures head-to-head test now has receipts. Across 300 screens of the same 109 HR Business Partner CVs, three frontier models never shortlisted 55% of the pool, on any day. That is not a fairness debate anymore. That is a slot machine you need to learn how to play.
The finding: 55% of resumes were invisible to every model
In May 2025, Hair Ventures ran 109 anonymized HRBP resumes for a global Meta role through ChatGPT-4o, Gemini 2.0 Flash, and Grok 3 with a normal recruiter prompt, across 300 total screens. The headline number: 55% of resumes in the pool were never shortlisted by any model on any day. They vanished, cleanly, from every top-ten list the models produced.
The other numbers are just as ugly:
- The three models agreed on only 14% of daily top-ten shortlists.
- Two AI screeners looking at the same CVs disagreed four times out of five.
- Individual models reshuffled identical resumes by an average of ±2.5 rank places day-to-day.
- One resume went from #10 to #1 in 48 hours on Gemini with no edits.
- Rationale bullets recycled the same three phrases 96% of the time.
- Human recruiter inter-rater reliability sits around κ ≈ 0.49. The LLMs delivered half that agreement with roughly twice the volatility.
Greg Savage brought the study back into wide circulation on Feb 25, 2026 in "Your recruitment business model is dying," as the EU AI Act's high-risk hiring rules took effect on Aug 2, 2026. The audience shifted from recruiters to candidates, and the practical question changed with it: if your resume never shortlists, is the resume broken, or is the reader?
From the Hair Ventures 300-screen test of ChatGPT-4o, Gemini 2.0 Flash, and Grok 3 on 109 HRBP resumes.
Why identical resumes get inconsistent AI screen results
LLM resume screeners are not reading, they are phrase-matching against a small habitual output vocabulary. The 96% rationale-recycling number is the tell. When a model can only produce three variants of "strong stakeholder management, cross-functional experience, aligned to JD," it is not evaluating you against the job, it is checking whether your words look like its three sentences.
That produces three distinct failure modes:
- Cross-model disagreement (80% of pairs). Different tokenizers, different training data, different system prompts. The recruiter's choice of tab (ChatGPT vs Gemini vs Grok) decides your fate more than your work does.
- Same-model drift (±2.5 rank places). Ranks move day-to-day with no CV change. A candidate ranked 8 to 12 is functionally at the mercy of a re-run.
- Structural invisibility (the 55%). If your phrasing sits outside the model's habitual output vocabulary, no roll of the dice puts you on the list. You are not competing for rank. You are not in the eligible set.
The right frame for "beat the AI resume screener" advice is not "trick one model." It is: write for the intersection of what all three models tokenize the same way, and mirror the job description's grammar rather than your own voice. That is exactly what Refolk does when you paste a posting. Refolk rewrites your resume against that specific JD, using the phrasings the screener is priming on, and scores how well you actually fit before you submit.
What the 55% vanish looks like at scale
At the scale of a real job market, "55% never shortlisted" is not a lab curiosity. It is thousands of people who never get read. In Refolk's index of professional profiles, the HRBP market the Hair Ventures study mimicked is large enough to make the number concrete.
| Segment | Total pool | Note |
|---|---|---|
| HR Business Partner, United States | 8,848 | Refolk index |
| HR Business Partner, Germany | 7,878 | Refolk index |
| Senior HR Business Partner, United States | 874 | Refolk index |
| US HRBP-to-Senior HRBP ratio | 10.1 : 1 | ~10 base-title HRBPs per senior HRBP, all competing on the same keyword |
| Hair Ventures test as share of US HRBP pool | 109 / 8,848 ≈ 1.2% | A realistic funnel slice |
| Implied "vanished" if 55% applied to US HRBPs | ~4,866 | 8,848 × 0.55 |
That last row is the one to sit with. If the Hair Ventures blind-spot rate held at national scale for a single common title, roughly 4,866 US HRBPs would be structurally unreadable to the current generation of LLM screeners. Not rejected. Not ranked low. Never seen.
The pool also tells you who is doing the screening. Top US employers of HRBPs in Refolk's index include HCA Healthcare, Intel, American Credit Acceptance, and Samsonite. Top German employers include Siemens, Infineon, Festo, and Deutsche Bahn. These are the ATS instances a Hair Ventures-style prompt would run against, and the German ones are now subject to the EU AI Act's high-risk hiring rules.
The ATS blind spot is a phrasing problem, not a keyword one
The blind spot is not caused by missing keywords. It is caused by phrasing outside the model's habitual output vocabulary, which means your bullets need to look like the JD's bullets, not like your voice.
Here is the working checklist that survives all three models:
- Job title strings, exact. If the posting says "HR Business Partner II," use that string. Do not translate to "Senior People Partner."
- Dates in "Month YYYY" or ISO. Not "Summer 2023." Not "2 yrs."
- Verb grammar mirrored from the JD. If the posting says "partner with leaders to drive," your bullet says "partnered with leaders to drive." Same verb family, same object shape.
- No infographics, no icons, no two-column layouts. Tokenizers mangle them and rationale bullets never quote them.
- One resume per posting. The 25% top-ten swing from ±2.5 rank volatility means a generic resume rolls the dice; a JD-mirrored one shortens the distribution.
The tedious part is doing this for every application. That is the exact work Refolk takes off you: paste the posting, get your own resume back rewritten for it, with the cover letter drafted in the same phrasing family and a fit score that tells you whether the application is worth sending at all.
The 55% vanish cohort is not unqualified. It is phrased outside the model's habitual output vocabulary.
Rank volatility means you should apply on more than one day
If a model reshuffles identical resumes by ±2.5 rank places day-to-day, applying on a single day is statistical malpractice. Where the ATS allows updates, resubmit.
The math is blunt. For a top-10 shortlist:
- ±2.5 places = a 25% swing. A candidate ranked 8 to 12 today is a coin flip for tomorrow's top ten.
- Hair Ventures logged one resume that jumped from #10 to #1 in 48 hours on Gemini.
- Two re-runs against a model with ±2.5 drift roughly doubles the odds of catching a favorable draw.
None of this helps if you are in the 55% vanish cohort. Volatility is a top-of-funnel game. Structural invisibility is a "wrong resume" game. You have to fix the second before the first is worth playing.
What the EU AI Act changed on Aug 2, 2026
From Aug 2, 2026, EU-based applicants rejected by AI-assisted screening have a right to explanation for that decision under Article 86 of the AI Act. In a world where 96% of rationales are three recycled phrases, that right is unusually cheap leverage.
The specifics that matter for a job seeker:
- Annex III classifies recruitment, candidate selection, performance evaluation, task allocation, monitoring, and promotion or termination decisions as high-risk AI use. Resume screeners are squarely in scope.
- Article 86 grants a right to explanation for individual decisions by high-risk AI that produce legal or similarly significant effects. A hiring rejection qualifies.
- The transparency obligations took effect Aug 2, 2026 and are now subject to enforcement.
If you are in the EU and you get a form rejection you suspect came from an LLM screen, asking for the rationale in writing costs you nothing and raises the compliance cost of a lazy no. It also, occasionally, restarts the conversation with a human.
The candidate-side asymmetry nobody talks about
Candidates using ChatGPT to write their resumes fare worse against ATS than candidates using dedicated tools, even as 83% of companies say they will use AI to screen resumes in 2025 (ResumeBuilder.com's October 2024 survey of 948 business leaders). Most of those companies are pasting into ChatGPT, not running specialized software - and so are the applicants.
The published pass rates:
| Tool used to write resume | Average ATS pass rate |
|---|---|
| ChatGPT (generic LLM) | 29% |
| Dedicated AI resume tool | 71% |
That is a 42-point gap on the same postings. The reason is the same reason the 55% vanish: generic LLMs generate in their own habitual output vocabulary, which is exactly the vocabulary other generic LLMs are least surprised by, and therefore least likely to rank. Dedicated tools tokenize the JD first, then generate. Refolk sits on the candidate side of this asymmetry: it reads the posting, writes your resume against it, drafts the cover letter, and scores the fit before you spend a slot on the application.
Human recruiter inter-rater reliability sits near κ ≈ 0.49. The models delivered roughly half that, with twice the volatility.
A 30-minute fix for the resume that never shortlists
If your resume has never shortlisted for a role you clearly fit, spend 30 minutes on the four changes most likely to move you out of the vanish cohort. In order of impact:
- Replace your creative job titles with the JD's exact strings. "Growth Marketing Lead" becomes "Senior Marketing Manager" if that is what the posting says.
- Rewrite your top three bullets in the JD's verb grammar. Same verb families, same object shapes, your numbers.
- Flatten the layout. Single column, standard fonts, dates as "Month YYYY," no icons, no tables inside the resume itself.
- Generate one resume per posting. Not one master resume. Not three variants. One per JD, mirrored.
You can do all of it by hand. Most people will not, which is why they stay in the 55%.
FAQ
Does the 55% vanish rate apply to all jobs or just HRBP roles?
The Hair Ventures test used HRBP resumes for a Meta role, so the exact 55% number is specific to that pool. The mechanism, though - rationale recycling, tokenizer overlap, phrasing outside the model's habitual vocabulary - is model behavior, not role behavior. Expect a similar shape wherever recruiters paste resumes into ChatGPT, Gemini, or Grok with a generic prompt. The specific percentage will vary by title, seniority, and how templated the applicant pool is.
Should I still tailor my resume if AI screeners are unstable?
Yes, more than before. Rank volatility (±2.5 places) hurts everyone equally, but structural invisibility (the 55%) hits only the un-tailored. Tailoring moves you from "never eligible" to "eligible and rolling the dice," which is the game you actually want to be in. The re-run then becomes worth doing, because a mirrored resume shortens the distribution around your rank.
If I am in the EU and get rejected by an AI screener, what do I actually do?
Reply to the rejection in writing and cite Article 86 of the EU AI Act, asking for the specific rationale for the decision and confirmation of whether AI-assisted screening was used. You do not need a lawyer for this first step. Since Aug 2, 2026, employers have transparency obligations for high-risk AI use in hiring, and a written request creates a compliance record that is expensive to ignore.
Why do ChatGPT-written resumes have lower ATS pass rates than dedicated tools?
Because generic LLMs generate in the same habitual output vocabulary that other generic LLMs use to screen, which produces text the reader model is least surprised by, and therefore least likely to rank as distinctive. Dedicated resume tools tokenize the job description first and generate against it, which mirrors the JD's grammar into your bullets. The published gap is 29% versus 71% ATS pass rate on the same postings, which is why the tool you choose on the candidate side matters more than the model the recruiter picked.