If you drafted your resume in ChatGPT and it landed in a Claude-powered screener, you may have just cut your callback rate roughly in half without changing a single word of your history. A June 2, 2026 study from i10X Research put four frontier models on both sides of a hiring pipeline and found the drafting model changed hire rates by up to 42 percentage points on identical facts. That is not a rounding error. That is the difference between a phone screen and a form-letter reject.
The 42-point gap, in one paragraph
Claude Sonnet 4.6 recommended hiring 84% of the resumes Claude itself wrote, but only 42% of resumes written by GPT-5.4, on the same underlying candidates. i10X Research ran 1,576 valid evaluations across 100 personas, generating four resume versions per persona (GPT-5.4, Claude Sonnet 4.6, Gemini 3 Pro, xAI Grok 4.3) with identical facts but different language and structure, then had every model score every resume blind. The spread was not subtle: on one identical document, GPT and Claude diverged by 29 score points, which the study frames as the difference between a borderline "maybe" and a clear reject.
Claude hired 84% of Claude-written resumes and only 42% of GPT-written ones, i10X Research, June 2 2026.
The point for a job seeker is narrow and practical. AI resume screening bias is no longer just about names, schools, or gaps. It now includes a stylistic fingerprint left by whichever model drafted the document, and that fingerprint is legible to other models.
Why the drafting model leaks through
Models recognize their own prose the way a copy editor recognizes a colleague's tics, and they reward or punish it accordingly. LLMs trained on similar data develop predictable syntactic habits: sentence rhythms, hedge phrases, transition words, bullet cadence. When another model scores that output, those habits act as a soft signal about "who wrote this," which then gets tangled up with quality judgments.
Two forces amplify the effect:
- Experiential bias. Researchers at Princeton and the University of Chicago, in an ICML 2026 Seoul paper covered by MIT Technology Review on July 20, 2026, showed that LLMs form biases from experience, not only from training data. Co-author Ryan Liu notes that models trained on math and code tasks "settle on a hunch too early" and stereotype job applicants more than humans do.
- Monoculture on the screening side. Stanford HAI reports roughly 90% of U.S. employers use AI screening tools, and most rely on the same few third-party vendors. Applicants can diversify their drafting tool. They cannot diversify the screener.
That asymmetry is the leverage point. You get to pick which AI touches your resume first. The employer does not consult you about which vendor grades it.
The full scoreboard
Gemini 3 Pro is the surprise winner: its resumes averaged a 94.5% hire rate across every evaluator, not just its own. Here is the compact version of the i10X dataset, along with the Refolk index figures that explain why the "maybe" bucket is so deadly this year.
| Metric | Figure | Source |
|---|---|---|
| Claude hire rate, Claude-written resumes | 84% | i10X Research |
| Claude hire rate, GPT-written resumes | 42% | i10X Research |
| GPT self-penalty on GPT-written resumes | -15 pts | i10X Research |
| Gemini-written resumes, avg hire rate across all evaluators | 94.5% | i10X Research |
| Max GPT vs Claude score divergence, identical document | 29 pts | i10X Research |
| Valid data points in the study | 1,576 of 1,600 (98.5%) | i10X Research |
| U.S. Software Engineers | ~347,679 | Refolk's index |
| U.S. Marketing Managers | ~83,555 | Refolk's index |
| U.S. Recruiter / TA professionals | ~91,805 | Refolk's index |
| Engineers per recruiter (derived) | ~3.8 : 1 | Refolk's index |
Two things jump out. First, Gemini is the safe ghostwriter regardless of the screener on the other side. Second, GPT punishes its own prose by 15 points, which is the first empirical evidence I have seen of a model detecting and downgrading its own slop.
GPT is now penalizing its own writing style
GPT-5.4 rates GPT-written resumes 15 percentage points lower than the same facts written by other models, which is the first sign that "ChatGPT voice" has become a negative signal even to ChatGPT. Something about the phrasing (the hedge stacks, the parallel triads, the "leveraged cross-functional initiatives to drive measurable impact" cadence) is now common enough in training data that the model treats it as low-signal boilerplate.
Practical implications for anyone writing in ChatGPT right now:
- Strip the giveaway phrases. "In today's competitive landscape," "results-driven professional," and "proven track record" all read as GPT tell-tales.
- Break the triad habit. GPT loves lists of three. Human resumes vary.
- Kill the em-dashes and semicolons GPT scatters through bullets. Short periods read as human.
- Rewrite any bullet that opens with "Spearheaded," "Orchestrated," or "Leveraged." Those verbs are the AI-slop tell.
If you do not want to hand-edit 40 bullets across 12 applications, that is the exact work Refolk takes off you: it drafts your resume from your own history, then rewrites it for each posting in language that keeps the facts and drops the giveaway phrasing.
Which AI to write your resume in, by screener
Match your drafting model to what you know about the target employer, and default to Gemini when you do not know. That is the short version of the guidance the i10X data supports.
- You know the employer uses Claude-based screening. Draft in Claude Sonnet 4.6. The 42-point self-preference is yours to exploit.
- You know the employer uses GPT-based screening. Draft in Gemini, not GPT. GPT's 15-point self-penalty makes it the worst pick for a GPT screener.
- You do not know the screener (the usual case). Draft in Gemini 3 Pro. It scored 94.5% on average across every evaluator in the study, meaning its prose style satisfies every rubric roughly equally.
- You are applying to a very small company that probably still has a human reading. Model choice matters less. Focus on fit and specificity.
The Claude vs GPT resume decision is not really about which model writes better English. Both are fine. It is about which stylistic fingerprint the screener on the other end has been trained to like.
The employer does not consult you about which vendor grades your resume. You still pick the pen.
The "maybe = reject" multiplier
A "maybe" from an AI screener is a rejection, because no human ever reads it. In real applicant tracking systems, the i10X team notes, a maybe verdict effectively ends the candidate's journey. That is why a 15 or 29 point stylistic drift is not a nuisance. It is a hiring catastrophe.
The math gets worse when you layer in volume. Stanford HAI reports companies are now seeing nearly three times as many applications for entry-level roles as in 2022. In Refolk's index of professional profiles, there are roughly 347,679 U.S. Software Engineers and only ~91,805 recruiters and talent acquisition professionals total, a 3.8-to-1 ratio, and the recruiter pool is barely bigger than the entire U.S. Marketing Manager population (~83,555). Recruiter concentration inside that pool skews heavily toward a few large shops like Robert Half, RCM Health Care Services, and Forrester Research.
~347,679 Software Engineers vs ~91,805 recruiter / TA pros in Refolk's index. Nobody is reading the "maybe" pile.
No human is triaging that. AI self-preference effects compound because the human tiebreaker has been outsourced to a queue nobody works.
What actually changes your outcome
Change three things this week: the model you draft in, the phrases that give your draft away, and how tightly the resume matches each posting. That is the whole playbook.
- Draft in Gemini or the target screener's family. Default Gemini.
- Kill AI-slop phrases. No "spearheaded," "leveraged," "results-driven," or triads.
- Tailor per posting. Not the summary. The bullets. The verbs. The metrics you choose to keep in versus cut.
- Score your fit before you send. Do not spray. A 60% fit resume in a 3x volume market gets binned by any screener regardless of who wrote it.
That last bit matters more than model choice. Refolk writes your resume from your own history, tailors it to every posting you paste in, drafts the cover letter, and scores how well you actually fit the job before you press send. That fit score is the honest signal to spend your next hour on this posting or skip it.
What the Colorado AI Act does and does not fix
The Colorado AI Act, effective June 2026, requires developers and users of AI hiring tools to use reasonable care to prevent algorithmic discrimination, but it does not touch stylistic bias. The law targets protected-class outcomes: race, gender, disability, age. A 42-point gap based on whether Claude or GPT held the pen is not a protected class question, and no regulator is going to litigate it.
Prior academic work (Wilson and Caliskan, 2025) already showed LLM resume screeners disadvantaged Black and female-associated names. Rishi Bommasani and colleagues at Stanford, in their 2026 "Algorithmic Monocultures in Hiring" paper, documented how vendor concentration multiplies any single model's biases across the whole market. The stylistic-fingerprint issue is the newer, uglier variant: a bias applicants cannot see, cannot appeal, and are not legally protected against.
The only defense on the applicant side is tool choice. Pick your drafting model with the same care you pick which jobs to apply to.
The shelf life on all of this
Every number in this article has a shelf life measured in months, because models retrain and screeners drift. The Princeton and Chicago finding, that LLMs accumulate biases from experience, means today's Gemini advantage may not hold in six months. Screener vendors will update. Drafting models will update. The specific "GPT loves triads, Claude loves crisp verbs" tells will shift.
Two habits survive the model churn:
- Draft in a different family than the screener you expect, or default to Gemini. The direction of self-preference is stable even if the magnitude drifts.
- Stop letting one model write and grade your resume by itself. Get a second model, or a human, to read it before you send. The 29-point divergence on identical documents is your warning that no single model is calibrated.
If you are applying seriously in the next 90 days, treat which AI drafts your resume as a real tactical decision. It is not aesthetic. It is a 42-point swing on identical facts.
FAQ
Should I ever draft my resume in GPT-5.4 given the 15-point self-penalty?
Only if you know the screener is Claude or Gemini, and even then Gemini is usually a cleaner choice. GPT's 15-point self-penalty means a GPT-drafted resume going into a GPT-based screener is the worst combination in the i10X study. If you are already deep in ChatGPT for other work, at minimum paste the final draft into Gemini and ask it to rewrite in its own voice before you submit.
How do I find out which AI screener a company uses?
You usually cannot, directly. Signals: the ATS name on the job posting URL, any "AI-assisted screening" disclosure in the application flow, and Colorado AI Act notices for roles based in that state once the law takes effect in June 2026. When you cannot tell, default to Gemini 3 Pro because its 94.5% cross-evaluator hire rate in the study makes it the safest pick against an unknown screener.
Does this mean human recruiters no longer see my resume at all?
Not entirely, but the "maybe" bucket almost never reaches them. Stanford HAI's data on 3x application volume since 2022, combined with the 3.8-to-1 engineer-to-recruiter ratio in Refolk's index, means humans triage only the top of the AI-scored stack. A stylistic drift that pushes you from "hire" to "maybe" functionally removes you from the human pipeline, which is why a 15 or 29 point model bias translates directly into a rejection.
Is it worth rewriting old applications I already sent in ChatGPT?
For roles still open, yes, if you can reapply cleanly through a new posting, different req, or a referral path. For roles where you already got a rejection, no, that signal has been logged. Focus the effort forward: pick one drafting model, strip the AI-slop phrases, and tailor per posting. Refolk handles the tailoring and fit scoring so you spend your time on the 5 postings worth applying to instead of the 50 that were never going to score.