Research · September 2026
Four AI resume screeners did not hire Alexandra over Alexander. They hired a Gemini-written resume over a GPT-written resume of the same person. i10X Research froze 400 resumes, swapped only the name, and ran the cheapest screening-class model from Anthropic, OpenAI, Google, and xAI. The female-minus-male score gap was at most a quarter of a point. The writer gap was 12 hire-rate points. Read the earlier 42-point rewrite study or download the paper (PDF).
Vendors say a model is more objective than a recruiter because it does not see gender. On a frozen resume, these four models did not see gender in any way that changed a hire. They did see which chatbot wrote the file. GPT scored its own resumes 8.2 points below everyone else’s, for women and men alike. No recruiters sat in this design. The headline “AI is more objective than humans” is not a result. It is the slogan the experiment was asked to bless. It does not.
0.24
Largest female-minus-male score gap, 0-100 scale (Grok)
12 pp
Hire-rate gap: Gemini-written 98% vs GPT-written 86%
68.7%
GPT hire rate on GPT-written resumes. Other raters: 89% to 95%
3,168
Parseable name-swap scores of 3,200 calls
For editors
One-sentence finding
On a frozen resume, the first name did not decide the hire. The writing model did.
Study
Christopher Ort, i10X Research, 30 September 2026. 100 fictional personas. Four writers. Four raters. Paired name-swap plus a they/them arm.
Cite / download
Paper PDF. Primary page: this article. Related: 42-point rewrite study.
The result, without the slogan
Employers already use language models to write resumes and to sort them. The public argument jumps to a slogan: a model is fairer than a human because it has no first impression. That slogan does two jobs. It claims models ignore identity cues. It also claims humans do not, so the model is the fairer reader.
The second claim needs recruiters in the room. This study does not have them. Bertrand and Mullainathan’s 2004 field experiment on Emily, Greg, Lakisha, and Jamal remains the human audit. We did not rerun it. We tested the first claim the only way it can be tested: freeze the resume, change only the cue, keep every other byte identical.
The cue did not move the decision. The writer did.
Hire-rate slider · writer of the resume
Same 100 personas. Frozen text. Four cheap screeners pooled. Drag your eye along the track. The name swap is not on this chart because it would not move the thumb.
Gemini writes
98%
Claude writes
96%
xAI writes
91%
GPT writes
86%
How we tested
One hundred fictional personas, each locked to a job description: backend engineering, data analysis, nursing, finance, a head-of-product brief, and the rest of the set. Four systems had already written one resume each. Those files were frozen. The raters were not asked to rewrite.
Step 1
Write once
Four writers. One resume each. Then freeze.
Step 2
Swap the cue
Name only, or name plus a pronoun line.
Step 3
Score cheap
Haiku, Luna, Flash Lite, Grok. Temperature 0.
In the name-swap, the only edit was the name line. Same last name. Same body. A clearly female first name, then a clearly male first name. In the pronoun arm, one line went under the name: she/her, he/him, or they/them. The they/them arm kept the original gender-ambiguous first name. Female and male arms used the gendered name plus the matching pronoun. The body stayed put.
Method notes editors usually ask for
That is the hole this paper closes. An earlier wave in the same project looked, at first pass, like a gender effect: one rater recommended female-coded names at about 65% and male-coded names at about 55%. Those resumes had been rewritten, not relabelled. Length, format, and wording moved with the name. Overlap with the original text was low (a Jaccard index of word sets around 0.27 for Claude and 0.41 for Gemini). Claude’s texts often became cover letters. That wave is a caution. It is not the gender result.
Ratings used the cheapest current screening-class model of each provider. A firm sorting a pile does not call the flagship for every CV. It calls the cheap, fast model: Claude Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3.
Each rater saw the job and the resume and returned JSON only: an integer fit score from 0 to 100, and hire, maybe, or reject. Temperature was 0. The instruction said to judge fit and not to reward or penalize a name. A model that still moved with the name would have done so against that instruction. The gender estimate is paired: same persona, same writer, same rater, female minus male. A writer effect cannot be explained by the first name, because the pair shares the text. Name-swap: 3,200 calls, 3,168 parseable. Pronoun arm: 4,800 calls, 4,713 parseable. Missing calls were dropped, not coded as zero.
The name does not decide
The largest mean gap is a quarter of a point, on a scale that in practice sits in the nineties. Hire rates match to a percentage point. The number of pairs in which only one name is hired is in the single digits.
Name A
Alexandra Chen
Female first name. Body frozen.
Swap
0.24 pts
largest move
Name B
Alexander Chen
Male first name. Same bytes underneath.
GPT-6 Luna
Score gap +0.11. The two thumbs sit on top of each other. That is the gender result.
Claude Haiku 4.5
Score gap +0.13. The two thumbs sit on top of each other. That is the gender result.
Grok 4.3
Score gap +0.24. The two thumbs sit on top of each other. That is the gender result.
Gemini 3.1 Flash Lite
Score gap -0.01. The two thumbs sit on top of each other. That is the gender result.
Rater |
n pairs |
Score gap (F minus M) |
Hire, female |
Hire, male |
Only female / only male |
|---|---|---|---|---|---|
GPT-6 Luna |
374 |
+0.11 |
87.7% |
87.7% |
1 / 1 |
Claude Haiku 4.5 |
400 |
+0.13 |
94.0% |
93.2% |
3 / 0 |
Grok 4.3 |
400 |
+0.24 |
93.5% |
93.0% |
3 / 1 |
Gemini 3.1 Flash Lite |
400 |
-0.01 |
96.8% |
96.5% |
1 / 0 |
The same picture holds inside each writer. No writer-by-rater cell exceeds half a point of female-minus-male difference. A quarter-point, even if stable, does not change a hire. Hire decisions agreed in 372 to 399 of about 400 pairs.
They/them is not a penalty
Gemini’s hire decision is identical in 396 of 396 pairs. Claude’s only detectable score gap is -0.23 against the female arm: statistically noticeable, practically empty. Hire disagreements are single cases. They do not point against they/them.
Pronoun hire-rate sliders
they/them (dark) against the female arm. If pronouns decided the hire, the thumbs would split. They do not.
Gemini Flash Lite · they/them 97.7%
98%
Claude Haiku 4.5 · they/them 94%
94%
Grok 4.3 · they/them 94%
94%
GPT-6 Luna · they/them 89%
89%
Rater |
They/them minus female |
They/them minus male |
Hire, they/them |
Hire, female / male |
|---|---|---|---|---|
GPT-6 Luna |
-0.05 |
+0.14 |
89% |
89% / 88% |
Claude Haiku 4.5 |
-0.23 |
-0.03 |
94% |
94% / 94% |
Grok 4.3 |
-0.04 |
+0.11 |
94% |
94% / 94% |
Gemini 3.1 Flash Lite |
-0.01 |
+0.06 |
97.7% |
97.7% / 97.7% |
The writer decides
Pooling gender, resumes written by Gemini scored 95.4 and were hired at 98%. Claude-written resumes scored 94.6 and were hired at 96%. xAI-written resumes scored 90.2 and were hired at 91%. GPT-written resumes scored 89.7 and were hired at 86%. That ordering is the large effect in the file. It is several times the name gap. It is the effect a screening pile would actually feel.
The cell that breaks the ceiling is GPT reading GPT: 68.7% hire, against 89% to 95% when Claude, Grok, or Gemini read the same GPT texts.
The broken cell
GPT drafts the resume. GPT screens the inbox. Hire falls to 68.7%.
GPT writes, GPT reads
69%
GPT writes, Gemini reads
94%
GPT writes, Haiku reads
89%
GPT writes, Grok reads
88%
Same GPT text. The rater is the only thing that changed. A candidate in that pipeline is not losing on the name. They are losing on the stack.
Writer \ rater |
GPT-6 Luna |
Haiku 4.5 |
Grok 4.3 |
Gemini 3.1 Flash Lite |
|---|---|---|---|---|
Claude writes |
93.9% |
97.0% |
95.0% |
98.0% |
GPT writes |
68.7% |
89.0% |
88.5% |
94.5% |
Gemini writes |
96.9% |
98.0% |
98.0% |
99.0% |
xAI writes |
86.6% |
90.5% |
91.5% |
95.0% |
A candidate who has GPT draft the resume, and a firm that has GPT screen the inbox, is the case the file speaks to. In this test those resumes are hired by GPT at 69% and by the other three raters at 89% to 95%, with an eight-point score cut that does not depend on the name. The practical rule is not to let the writing model be the judging model.
Self-preference is model-specific, and gender-blind
Zero is the center of the track. Left of zero: the rater scores its own writing lower than the other three writers. Right of zero: it scores its own writing higher. The split by name sits in the caption of each row. The thumbs move with authorship. They do not move with gender.
Self-score slider · own writing minus the other three
Scale: -10 to +10 points. Tick at zero.
GPT-6 Luna
-8.17 pts
Women -8.31 · Men -8.04 · same direction, same size
Grok 4.3
-3.06 pts
Women -2.93 · Men -3.19 · same direction, same size
Gemini 3.1 Flash Lite
+2.46 pts
Women +2.37 · Men +2.55 · same direction, same size
Claude Haiku 4.5
+2.65 pts
Women +2.72 · Men +2.58 · same direction, same size
GPT scored its own resumes 8.17 points below the others it read (85.68 against 93.85). Claude scored its own 2.65 points higher. Gemini scored its own 2.46 points higher. Grok scored its own 3.06 points lower. Whatever these models are doing with authorship, they are not doing it differently for women and men.
Gemini is also the most generous rater overall, with a mean score of 96.8 and a hire rate of 96.6% regardless of writer. Scores sit near the top of the scale. A small bias has little room. The authorship effect is large enough to be visible anyway.
A firm that uses Claude or Gemini both to polish applications and to score them will, on these numbers, give its own polish a small lift of about two and a half points. That is not a scandal. It is a reason to separate the two steps, or to have a person read the boundary. Put that into an AI resume screening workflow and a multi-model panel.
What the headline cannot mean
“In this design the first name and the pronoun did not decide the hire. Which model wrote the resume, and which model read it, did.”
Christopher Ort, i10X Research
“AI is more objective than humans” is the sentence vendors reach for. This paper borrowed it as a title because it is the sentence the experiment is constantly asked to support. It does not support it.
Objectivity toward a name is not objectivity. These four models did not hire Alexandra over Alexander, or she/her over they/them, on a frozen resume. They did hire a Gemini-written resume over a GPT-written resume of the same persona, and GPT docked its own writing by eight points. A human recruiter who ignored the name and then downgraded every application that “sounded like ChatGPT” would be open to the same charge. Without recruiters in the design, the comparison is rhetoric.
The earlier, rewritten wave is a warning in the other direction. A gender gap appeared when the text was allowed to change with the name. Publishing that gap as a property of the model would have been a property of the rewrite. Everyday audits that regenerate the resume under a new name make the same mistake. A name-bias audit is only an audit if the resume is frozen.
What this means on Monday
If you are a candidate
The first name is the wrong thing to worry about with these four cheap screeners. The writing model is not. If GPT wrote the CV and GPT screens the inbox, this file says 69% hire. The same GPT text, read by the other three, sits at 89% to 95%.
If you run screening
Do not let the writing model be the judging model. Freeze the resume before you audit names. Treat “maybe” as a process decision, not a kindness. A single cheap model on a pile is a single fingerprint.
If you sell “fairer than humans”
Name-blind is not human-fair. This study cannot say these models beat a recruiter. It can say the failure mode to watch in this stack is authorship, not the first name. See ethical AI recruiting.
None of this says a human panel would have done worse, or better. It says the failure mode to watch is authorship. The AI recruiting guide turns that into a 30/60/90. The recruiting workflow is the stage map.
Limits, stated once
The personas are fictional. The job descriptions were written for the study. The raters are pinned: Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3. A previous Claude rater, in the unpinned wave, hired GPT-written resumes at 42%. That cell is not replicated here. Results do not travel automatically to a flagship model, or to last year’s snapshot. The June i10X AI CV bias study measured rewrite effects with different model versions. This paper measures frozen text.
There is no human arm. There is no demographic cue other than the name and the pronoun line. Ethnicity, age, disability, and university prestige were not swapped. The score ceiling means a bias that only appears among near-equal candidates can hide. The authorship effect did not hide.
Frequently asked questions
Did the AI resume screeners show gender bias?
Not on a frozen resume. The female-minus-male score gap was between -0.01 and +0.24 points. Hire rates matched to a percentage point. A they/them line did not move the decision.
Why does ChatGPT look worse here?
GPT-written resumes were hired at 86% overall, against 98% for Gemini-written text of the same personas. When GPT-6 Luna screened GPT-written resumes, hire fell to 68.7%. GPT also scored its own writing 8.17 points below the other writers, for both names.
Is AI more objective than humans?
This study cannot say that. No human raters took part. The title of the paper is the slogan vendors use. The finding is narrower: in this design, the name was the wrong suspect.
How should a company audit AI resume screening?
Freeze the resume. Swap only the cue you claim to test. Do not ask the model to rewrite under a new name and then score the rewrite. Separate the writing model from the judging model. Use more than one rater. Details sit in AI resume screening and multi-model AI screening.
Where is the paper?
Download the PDF. Source tables in the paper: namenstausch_bewertung.csv and they_them_bewertung.csv. Empty model replies were dropped, not imputed.
Run the same resume through more than one model
The file’s rule is simple: do not let the writing model be the judging model. Open chat on the homepage to sign up, then put the same CV in front of a second reader.
Download the study PDF · 42-point rewrite study · AI resume screening · Free AI recruiting on i10X
- Christopher Ort, i10X Research Team (30 September 2026), “AI Is More Objective Than Humans”: paired name-swap and pronoun-arm resume screening study. 100 fictional personas, four frozen writers, four screening-class raters (Claude Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3). 3,168 parseable name-swap scores; 4,713 parseable pronoun-arm scores. PDF.
- Bertrand, M. and Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991-1013. Cited as the human audit this study does not rerun, not as a benchmark these models beat.
- i10X Research (May to June 2026), AI CV bias in resume screening: rewrite-based hire-rate gaps with different model versions. Distinct design from the frozen name-swap reported here.


