You picked a chatbot in 2023 and never switched. That casual default is now costing you interviews, because in June 2026 i10X Research measured a 42-percentage-point swing in hire recommendations on identical candidates based purely on which model wrote the resume.
The 42-point tax, in one paragraph
Claude Sonnet 4.6, acting as a screener, recommended "hire" for 84% of resumes written by Claude but only 42% of resumes written by GPT-5.4. Same candidates, same jobs, same rubric. i10X Research (Singapore) generated four resume versions of 100 synthetic profiles across 12 industries and 4 career levels, then had GPT-5.4, Claude Sonnet 4.6, Gemini 3 Pro, and Grok 4.3 blind-score every version. Out of 1,600 possible evaluations, 1,576 were valid. The pattern was not noise. On one identical document, GPT and Claude diverged by 29 score points.
Claude approves 84% of Claude-written resumes and 42% of GPT-written ones, per i10X Research, June 2026.
This matters right now because Stanford HAI reported in 2026 that roughly 90% of U.S. employers are running AI screening tools, "with most relying on the same few third-party vendors." Entry-level application volumes are running near 3x their 2022 levels. The screener stack is concentrated on three or four frontier models, so the model that writes your resume is now an unpriced determinant of your interview rate.
Which AI actually writes the best resume
Gemini. Across every evaluator in the i10X study, Gemini-written resumes averaged a 94.5% hire recommendation rate, the highest of any authoring model. That is not Gemini scoring Gemini highly. That is GPT, Claude, and Grok all preferring Gemini's output over their own.
The mechanism matters, because "just use the best model" is the wrong frame. GPT-5.4 penalizes its own writing by 15 points on the i10X data. It recognizes itself and downgrades. Gemini 3 Pro's default output does not trigger the same recognition in the other frontier screeners.
The comparison, from the i10X dataset:
| Authoring model | Hire rate under all evaluators | Notable pattern |
|---|---|---|
| Gemini 3 Pro | 94.5% avg | Highest score under every screener |
| Claude Sonnet 4.6 | 90% under Gemini, 84% under Claude | Strong only when Claude reads it |
| GPT-5.4 | 42% under Claude | Penalized 15 pts by GPT itself |
| Grok 4.3 | Middle of pack | No strong self-bias reported |
If you have been asking "chatgpt vs claude resume, which one?", the honest answer on this dataset is that both lose to Gemini, and the gap is not small. Choosing Gemini over GPT when the screener is Claude is a 2.25x lift on identical qualifications (94.5% / 42%).
Why AI resume screening bias is identity, not quality
The bias is about who wrote it, not how well it reads. Saraf et al. (August 2025) proved this directly by mislabeling identical blog posts: tagging the same text as "Claude" versus "Gemini" produced up to 50-point voting-outcome swings in LLM evaluators. The "Claude" label added points to prose Claude never wrote. The "Gemini" label depressed scores on prose Gemini never wrote.
That kills the intuition that a well-written resume is a well-written resume. The screener is not reading for craft. It is reading for stylistic fingerprints, and it is rewarding whichever fingerprint it recognizes as its own kin. Xu, Li and Jiang (UMD Smith School, NUS, Ohio State), in their working paper "AI Self-preferencing in Algorithmic Hiring" (arXiv 2509.00462), replicated this across 2,245 resumes and three models (GPT-4o, Llama 3.3-70B, DeepSeek-V3). Self-preference rates ran 67% to 82%. Candidates whose resume was written by the same LLM as the screener were 23% to 60% more likely to be shortlisted.
The screener is not reading for craft. It is reading for stylistic fingerprints and rewarding its own kin.
Worse, self-preference gets stronger with capability. The UMD paper explicitly notes the strongest self-recognizers show the strongest preference. Frontier-model upgrades will amplify the tax, not shrink it. This is not a wait-it-out problem. The two known mitigations from that paper - instructing the evaluator to ignore authorship, and ensembling the main model with smaller models that lack self-recognition - reduced bias by 17% to 63% in relative terms, but both sit on the employer side. Neither is under your control.
The dataset behind the tax
Every number in this piece is sourced. Here is the spine, in one table:
| # | Segment | Number | Source |
|---|---|---|---|
| 1 | Claude hire rate on Claude-written resumes | 84% | i10X 2026 |
| 2 | Claude hire rate on GPT-written resumes | 42% | i10X 2026 |
| 3 | Avg hire rate on Gemini-written resumes (all evaluators) | 94.5% | i10X 2026 |
| 4 | GPT self-penalty on its own writing style | -15 pts | i10X 2026 |
| 5 | Shortlist lift when writer model matches screener | +23% to +60% | UMD Smith / arXiv 2509.00462 |
| 6 | U.S. Software Engineers vs AI/ML Engineers | 348,505 : 14,869 | Refolk's index |
| 7 | U.S. Recruiters | 90,224 | Refolk's index |
Row 6 is the sneaky one. In Refolk's index, there are roughly 348,505 people holding Software Engineer titles in the U.S. against only 14,869 with AI/ML Engineer titles. That is a 23.4-to-1 ratio. AI-native engineers over-index on tool literacy and will figure this out. The other 23 out of 24 are the ones defaulting to whichever chatbot they signed up for first. They are paying the tax and do not know it.
Who is actually running the screener
Assume the screener is a frontier LLM sitting on top of a Workday, Greenhouse, or iCIMS pipeline. The ATS itself does not do the language work; the LLM does, usually via an agentic-hiring vendor.
The concentration numbers on enterprise ATS share:
- Workday: ~30% to 32%
- Greenhouse: ~18% (Airbnb, Datadog, Palo Alto Networks, Block, and Veeva run on it)
- Lever: ~12%
- iCIMS: ~10%
- Ashby: ~5%
The top five cover ~77% of U.S. enterprise postings. On top of those sit Eightfold (talent-intelligence, high-volume) and Paradox / Olivia (Workday-owned, 100+ languages, dominant in hourly). This is where the third-party LLM screener actually lives. When Pymetrics and Stanford analyzed 4 million applications (Fortune, May 26, 2026), they surfaced the 330-day score-reuse problem: one bad model score can lock you out across every employer using that vendor for nearly a year.
You do not get to know which model is grading you. You do get to control which model wrote the input. That is the leverage.
A rewrite strategy that neutralizes model self-bias
Rewrite in Gemini, strip the tells, and match the posting's exact keyword surface. Three steps, in order:
- Draft or rewrite in Gemini 3 Pro. If your current resume was written by GPT or Claude, paste it in and ask Gemini to rewrite each bullet in its own voice against the target job description. On the i10X data, this alone moves you from a 42% floor to a 94.5% ceiling.
- Kill the fingerprint words. GPT-5.4 defaults to "leveraged," "spearheaded," "orchestrated," "streamlined," "actionable insights," "cross-functional collaboration," and triadic openings ("Designed, built, and shipped..."). Claude defaults to hedged qualifiers ("helped drive," "contributed to," "partnered with"). Replace with concrete verbs and numbers. "Cut p95 latency 38% by rewriting the auth cache" beats any adjective.
- Mirror the posting's noun phrases exactly. Screeners weight literal keyword matches heavily. If the JD says "distributed systems," do not write "large-scale infrastructure." Match the string.
This is exactly the work Refolk takes off you. Paste the posting, and Refolk writes your resume from your own history, tailored to that specific JD, and scores how well you actually fit before you spend an evening applying to a role you will not clear.
What to do if you have already applied with a GPT resume
Three moves:
- Do not reapply to the same req; the screener will remember the fingerprint.
- Rewrite the resume in Gemini and apply to a different team at the same company. Most enterprise ATS instances score at the req level, not the candidate level.
- For roles that closed with no response, send a fresh application against a new posting with a Gemini-authored resume and a manually-edited cover letter.
Refolk handles the tailoring loop per posting, which is where the marginal minute actually pays.
The one industry cut worth knowing
Business-facing roles pay the highest tax. The UMD paper's Figure 7 shows the largest self-preference gaps in accounting, sales, and finance. Agriculture, arts, and automotive show the smallest. Mechanism: business-function resumes rely more on prose bullets and less on hard-coded technicals (versions, frameworks, deploy counts), which gives the screener more surface to sniff for authorship. Engineers with dense technical bullets are partially insulated. Finance analysts writing about "driving alignment across stakeholders" are wide open.
If you are in sales, accounting, or finance and you have been getting silent rejections, the model you used to write your resume is a more likely culprit than your qualifications. Rewrite in Gemini, cut the abstraction, put numbers on every bullet.
From Refolk's index (348,505 vs 14,869). The AI-native cohort is only 4.3% the size of the general engineering pool.
The bottom line
Ninety percent of U.S. employers run AI screeners. Those screeners run on three or four frontier models. Those models prefer their own writing by 23 to 60 percentage points, and Claude in particular will cut your hire rate almost in half if it detects GPT authorship. Gemini-written resumes score highest under every evaluator on the i10X data. This is the single largest unpriced variable in the 2026 job search, and it is fixed by switching the tool you draft in.
FAQ
Does it actually matter whether I use ChatGPT, Claude, or Gemini to write my resume?
Yes, by up to 42 percentage points on hire recommendation for identical candidates, per the i10X June 2026 study. Gemini-written resumes averaged a 94.5% hire rate across all four screener models tested; GPT-written resumes dropped to 42% under Claude. If you have to pick one authoring tool without knowing the screener, pick Gemini.
Why does Claude penalize GPT-written resumes so hard?
Claude has the strongest self-recognition and the strictest evaluation profile in the i10X data. It detects GPT-5.4's stylistic fingerprint and downgrades. This is identity-based, not quality-based: Saraf et al. showed mislabeling identical text as "Claude" versus "Gemini" swings scores up to 50 points, so the bias is triggered by perceived authorship, not writing quality.
Will this bias go away as models improve?
No, it gets worse. The UMD Smith School paper (arXiv 2509.00462) explicitly finds the strongest self-recognizers show the strongest self-preference. Frontier upgrades amplify the tax rather than shrinking it. The two known mitigations sit on the employer side, not the candidate side.
What is the single most useful move I can make this week?
Rewrite your resume in Gemini 3 Pro against each specific posting, strip GPT and Claude fingerprint phrases, and mirror the JD's exact noun phrases. If you are applying to more than five roles a week, use Refolk to do the per-posting tailoring, cover letter, and fit score in one pass.