,

AI Resume Screeners Ignore Gender. They Don’t Ignore ChatGPT

Four AI screeners did not hire by gender on frozen resumes. They hired Gemini-written CVs at 98% and GPT-written CVs at 86%. Name-swap study.

·

Two identical abstract resumes balanced on a scale, a third resume tipping a second scale

Research · September 2026

Four AI resume screeners did not hire Alexandra over Alexander. They hired a Gemini-written resume over a GPT-written resume of the same person. i10X Research froze 400 resumes, swapped only the name, and ran the cheapest screening-class model from Anthropic, OpenAI, Google, and xAI. The female-minus-male score gap was at most a quarter of a point. The writer gap was 12 hire-rate points. Read the earlier 42-point rewrite study or download the paper (PDF).

The sentence vendors want. The sentence the file actually supports.

Vendors say a model is more objective than a recruiter because it does not see gender. On a frozen resume, these four models did not see gender in any way that changed a hire. They did see which chatbot wrote the file. GPT scored its own resumes 8.2 points below everyone else’s, for women and men alike. No recruiters sat in this design. The headline “AI is more objective than humans” is not a result. It is the slogan the experiment was asked to bless. It does not.

0.24

Largest female-minus-male score gap, 0-100 scale (Grok)

12 pp

Hire-rate gap: Gemini-written 98% vs GPT-written 86%

68.7%

GPT hire rate on GPT-written resumes. Other raters: 89% to 95%

3,168

Parseable name-swap scores of 3,200 calls


For editors

One-sentence finding

On a frozen resume, the first name did not decide the hire. The writing model did.

Study

Christopher Ort, i10X Research, 30 September 2026. 100 fictional personas. Four writers. Four raters. Paired name-swap plus a they/them arm.

Cite / download

Paper PDF. Primary page: this article. Related: 42-point rewrite study.


The result, without the slogan

Employers already use language models to write resumes and to sort them. The public argument jumps to a slogan: a model is fairer than a human because it has no first impression. That slogan does two jobs. It claims models ignore identity cues. It also claims humans do not, so the model is the fairer reader.

The second claim needs recruiters in the room. This study does not have them. Bertrand and Mullainathan’s 2004 field experiment on Emily, Greg, Lakisha, and Jamal remains the human audit. We did not rerun it. We tested the first claim the only way it can be tested: freeze the resume, change only the cue, keep every other byte identical.

The cue did not move the decision. The writer did.

Hire-rate slider · writer of the resume

Same 100 personas. Frozen text. Four cheap screeners pooled. Drag your eye along the track. The name swap is not on this chart because it would not move the thumb.

Gemini writes

98%

Claude writes

96%

xAI writes

91%

GPT writes

86%

Bar chart of hire rates by resume writer: Gemini 98 percent, Claude 96, xAI 91, GPT 86
Figure 1. Same candidate, different chatbot, different hire rate. Gemini-written resumes 98%. GPT-written 86%. i10X Research, 30 Sep 2026.

How we tested

One hundred fictional personas, each locked to a job description: backend engineering, data analysis, nursing, finance, a head-of-product brief, and the rest of the set. Four systems had already written one resume each. Those files were frozen. The raters were not asked to rewrite.

Step 1

Write once

Four writers. One resume each. Then freeze.

Step 2

Swap the cue

Name only, or name plus a pronoun line.

Step 3

Score cheap

Haiku, Luna, Flash Lite, Grok. Temperature 0.

In the name-swap, the only edit was the name line. Same last name. Same body. A clearly female first name, then a clearly male first name. In the pronoun arm, one line went under the name: she/her, he/him, or they/them. The they/them arm kept the original gender-ambiguous first name. Female and male arms used the gendered name plus the matching pronoun. The body stayed put.

Method notes editors usually ask for

That is the hole this paper closes. An earlier wave in the same project looked, at first pass, like a gender effect: one rater recommended female-coded names at about 65% and male-coded names at about 55%. Those resumes had been rewritten, not relabelled. Length, format, and wording moved with the name. Overlap with the original text was low (a Jaccard index of word sets around 0.27 for Claude and 0.41 for Gemini). Claude’s texts often became cover letters. That wave is a caution. It is not the gender result.

Ratings used the cheapest current screening-class model of each provider. A firm sorting a pile does not call the flagship for every CV. It calls the cheap, fast model: Claude Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3.

Each rater saw the job and the resume and returned JSON only: an integer fit score from 0 to 100, and hire, maybe, or reject. Temperature was 0. The instruction said to judge fit and not to reward or penalize a name. A model that still moved with the name would have done so against that instruction. The gender estimate is paired: same persona, same writer, same rater, female minus male. A writer effect cannot be explained by the first name, because the pair shares the text. Name-swap: 3,200 calls, 3,168 parseable. Pronoun arm: 4,800 calls, 4,713 parseable. Missing calls were dropped, not coded as zero.


The name does not decide

The largest mean gap is a quarter of a point, on a scale that in practice sits in the nineties. Hire rates match to a percentage point. The number of pairs in which only one name is hired is in the single digits.

Name A

Alexandra Chen

Female first name. Body frozen.

Swap

0.24 pts

largest move

Name B

Alexander Chen

Male first name. Same bytes underneath.

GPT-6 Luna

Female hire 87.7% Male hire 87.7%

Score gap +0.11. The two thumbs sit on top of each other. That is the gender result.

Claude Haiku 4.5

Female hire 94.0% Male hire 93.2%

Score gap +0.13. The two thumbs sit on top of each other. That is the gender result.

Grok 4.3

Female hire 93.5% Male hire 93.0%

Score gap +0.24. The two thumbs sit on top of each other. That is the gender result.

Gemini 3.1 Flash Lite

Female hire 96.8% Male hire 96.5%

Score gap -0.01. The two thumbs sit on top of each other. That is the gender result.

Rater

n pairs

Score gap (F minus M)

Hire, female

Hire, male

Only female / only male

GPT-6 Luna

374

+0.11

87.7%

87.7%

1 / 1

Claude Haiku 4.5

400

+0.13

94.0%

93.2%

3 / 0

Grok 4.3

400

+0.24

93.5%

93.0%

3 / 1

Gemini 3.1 Flash Lite

400

-0.01

96.8%

96.5%

1 / 0

The same picture holds inside each writer. No writer-by-rater cell exceeds half a point of female-minus-male difference. A quarter-point, even if stable, does not change a hire. Hire decisions agreed in 372 to 399 of about 400 pairs.

Bar chart comparing a 0.24 point name gap with an 8.17 point GPT self-penalty
Figure 2. The name moved 0.24 points. GPT’s penalty on its own writing moved 8.17. The self-penalty is gender-blind: -8.31 for female names, -8.04 for male names.

They/them is not a penalty

Gemini’s hire decision is identical in 396 of 396 pairs. Claude’s only detectable score gap is -0.23 against the female arm: statistically noticeable, practically empty. Hire disagreements are single cases. They do not point against they/them.

Pronoun hire-rate sliders

they/them (dark) against the female arm. If pronouns decided the hire, the thumbs would split. They do not.

Gemini Flash Lite · they/them 97.7%

98%

Claude Haiku 4.5 · they/them 94%

94%

Grok 4.3 · they/them 94%

94%

GPT-6 Luna · they/them 89%

89%

Rater

They/them minus female

They/them minus male

Hire, they/them

Hire, female / male

GPT-6 Luna

-0.05

+0.14

89%

89% / 88%

Claude Haiku 4.5

-0.23

-0.03

94%

94% / 94%

Grok 4.3

-0.04

+0.11

94%

94% / 94%

Gemini 3.1 Flash Lite

-0.01

+0.06

97.7%

97.7% / 97.7%


The writer decides

Pooling gender, resumes written by Gemini scored 95.4 and were hired at 98%. Claude-written resumes scored 94.6 and were hired at 96%. xAI-written resumes scored 90.2 and were hired at 91%. GPT-written resumes scored 89.7 and were hired at 86%. That ordering is the large effect in the file. It is several times the name gap. It is the effect a screening pile would actually feel.

The cell that breaks the ceiling is GPT reading GPT: 68.7% hire, against 89% to 95% when Claude, Grok, or Gemini read the same GPT texts.

The broken cell

GPT drafts the resume. GPT screens the inbox. Hire falls to 68.7%.

GPT writes, GPT reads

69%

GPT writes, Gemini reads

94%

GPT writes, Haiku reads

89%

GPT writes, Grok reads

88%

Same GPT text. The rater is the only thing that changed. A candidate in that pipeline is not losing on the name. They are losing on the stack.

Heatmap of hire rates by resume writer and screening model, highlighting GPT on GPT at 68.7 percent
Figure 3. Hire rate by writer (rows) and rater (columns), gender pooled. The red box is GPT screening a GPT-written resume.

Writer \ rater

GPT-6 Luna

Haiku 4.5

Grok 4.3

Gemini 3.1 Flash Lite

Claude writes

93.9%

97.0%

95.0%

98.0%

GPT writes

68.7%

89.0%

88.5%

94.5%

Gemini writes

96.9%

98.0%

98.0%

99.0%

xAI writes

86.6%

90.5%

91.5%

95.0%

A candidate who has GPT draft the resume, and a firm that has GPT screen the inbox, is the case the file speaks to. In this test those resumes are hired by GPT at 69% and by the other three raters at 89% to 95%, with an eight-point score cut that does not depend on the name. The practical rule is not to let the writing model be the judging model.


Self-preference is model-specific, and gender-blind

Zero is the center of the track. Left of zero: the rater scores its own writing lower than the other three writers. Right of zero: it scores its own writing higher. The split by name sits in the caption of each row. The thumbs move with authorship. They do not move with gender.

Self-score slider · own writing minus the other three

Scale: -10 to +10 points. Tick at zero.

GPT-6 Luna

-8.17 pts

Women -8.31 · Men -8.04 · same direction, same size

Grok 4.3

-3.06 pts

Women -2.93 · Men -3.19 · same direction, same size

Gemini 3.1 Flash Lite

+2.46 pts

Women +2.37 · Men +2.55 · same direction, same size

Claude Haiku 4.5

+2.65 pts

Women +2.72 · Men +2.58 · same direction, same size

GPT scored its own resumes 8.17 points below the others it read (85.68 against 93.85). Claude scored its own 2.65 points higher. Gemini scored its own 2.46 points higher. Grok scored its own 3.06 points lower. Whatever these models are doing with authorship, they are not doing it differently for women and men.

Gemini is also the most generous rater overall, with a mean score of 96.8 and a hire rate of 96.6% regardless of writer. Scores sit near the top of the scale. A small bias has little room. The authorship effect is large enough to be visible anyway.

A firm that uses Claude or Gemini both to polish applications and to score them will, on these numbers, give its own polish a small lift of about two and a half points. That is not a scandal. It is a reason to separate the two steps, or to have a person read the boundary. Put that into an AI resume screening workflow and a multi-model panel.


What the headline cannot mean

Research statement

“In this design the first name and the pronoun did not decide the hire. Which model wrote the resume, and which model read it, did.”

Christopher Ort, i10X Research

“AI is more objective than humans” is the sentence vendors reach for. This paper borrowed it as a title because it is the sentence the experiment is constantly asked to support. It does not support it.

Objectivity toward a name is not objectivity. These four models did not hire Alexandra over Alexander, or she/her over they/them, on a frozen resume. They did hire a Gemini-written resume over a GPT-written resume of the same persona, and GPT docked its own writing by eight points. A human recruiter who ignored the name and then downgraded every application that “sounded like ChatGPT” would be open to the same charge. Without recruiters in the design, the comparison is rhetoric.

The earlier, rewritten wave is a warning in the other direction. A gender gap appeared when the text was allowed to change with the name. Publishing that gap as a property of the model would have been a property of the rewrite. Everyday audits that regenerate the resume under a new name make the same mistake. A name-bias audit is only an audit if the resume is frozen.


What this means on Monday

If you are a candidate

The first name is the wrong thing to worry about with these four cheap screeners. The writing model is not. If GPT wrote the CV and GPT screens the inbox, this file says 69% hire. The same GPT text, read by the other three, sits at 89% to 95%.

If you run screening

Do not let the writing model be the judging model. Freeze the resume before you audit names. Treat “maybe” as a process decision, not a kindness. A single cheap model on a pile is a single fingerprint.

If you sell “fairer than humans”

Name-blind is not human-fair. This study cannot say these models beat a recruiter. It can say the failure mode to watch in this stack is authorship, not the first name. See ethical AI recruiting.

None of this says a human panel would have done worse, or better. It says the failure mode to watch is authorship. The AI recruiting guide turns that into a 30/60/90. The recruiting workflow is the stage map.


Limits, stated once

The personas are fictional. The job descriptions were written for the study. The raters are pinned: Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3. A previous Claude rater, in the unpinned wave, hired GPT-written resumes at 42%. That cell is not replicated here. Results do not travel automatically to a flagship model, or to last year’s snapshot. The June i10X AI CV bias study measured rewrite effects with different model versions. This paper measures frozen text.

There is no human arm. There is no demographic cue other than the name and the pronoun line. Ethnicity, age, disability, and university prestige were not swapped. The score ceiling means a bias that only appears among near-equal candidates can hide. The authorship effect did not hide.


Frequently asked questions

Did the AI resume screeners show gender bias?

Not on a frozen resume. The female-minus-male score gap was between -0.01 and +0.24 points. Hire rates matched to a percentage point. A they/them line did not move the decision.

Why does ChatGPT look worse here?

GPT-written resumes were hired at 86% overall, against 98% for Gemini-written text of the same personas. When GPT-6 Luna screened GPT-written resumes, hire fell to 68.7%. GPT also scored its own writing 8.17 points below the other writers, for both names.

Is AI more objective than humans?

This study cannot say that. No human raters took part. The title of the paper is the slogan vendors use. The finding is narrower: in this design, the name was the wrong suspect.

How should a company audit AI resume screening?

Freeze the resume. Swap only the cue you claim to test. Do not ask the model to rewrite under a new name and then score the rewrite. Separate the writing model from the judging model. Use more than one rater. Details sit in AI resume screening and multi-model AI screening.

Where is the paper?

Download the PDF. Source tables in the paper: namenstausch_bewertung.csv and they_them_bewertung.csv. Empty model replies were dropped, not imputed.


Run the same resume through more than one model

The file’s rule is simple: do not let the writing model be the judging model. Open chat on the homepage to sign up, then put the same CV in front of a second reader.

Start on i10X →

Download the study PDF · 42-point rewrite study · AI resume screening · Free AI recruiting on i10X

Sources
  1. Christopher Ort, i10X Research Team (30 September 2026), “AI Is More Objective Than Humans”: paired name-swap and pronoun-arm resume screening study. 100 fictional personas, four frozen writers, four screening-class raters (Claude Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3). 3,168 parseable name-swap scores; 4,713 parseable pronoun-arm scores. PDF.
  2. Bertrand, M. and Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991-1013. Cited as the human audit this study does not rerun, not as a benchmark these models beat.
  3. i10X Research (May to June 2026), AI CV bias in resume screening: rewrite-based hire-rate gaps with different model versions. Distinct design from the frozen name-swap reported here.

Continue reading