{"id":1000,"date":"2026-10-06T08:38:41","date_gmt":"2026-10-06T08:38:41","guid":{"rendered":"https:\/\/i10x.ai\/blog\/?p=1000"},"modified":"2026-10-06T08:38:41","modified_gmt":"2026-10-06T08:38:41","slug":"ai-resume-screening-gender-chatbot-bias","status":"draft","type":"post","link":"https:\/\/i10x.ai\/blog\/?p=1000","title":{"rendered":"AI Resume Screeners Ignore Gender. They Don&#8217;t Ignore ChatGPT"},"content":{"rendered":"\n<!--\nTITLE: AI Resume Screeners Ignore Gender. They Don't Ignore ChatGPT\nEXCERPT: Four AI screeners did not hire by gender on frozen resumes. They hired Gemini-written CVs at 98% and GPT-written CVs at 86%. Name-swap study.\nSLUG: ai-resume-screening-gender-chatbot-bias\nCATEGORY: AI\nPRIMARY_KW: AI resume screening gender bias\nDATA_CHECKED: 2026-09-30\n-->\n\n<div class=\"i10x-article\">\n\n<p class=\"i10x-pill\">Research \u00b7 September 2026<\/p>\n\n<p class=\"i10x-lead\">\nFour AI resume screeners did not hire Alexandra over Alexander. They hired a Gemini-written resume over a GPT-written resume of the same person. i10X Research froze 400 resumes, swapped only the name, and ran the cheapest screening-class model from Anthropic, OpenAI, Google, and xAI. The female-minus-male score gap was at most a quarter of a point. The writer gap was 12 hire-rate points. Read the\n<a href=\"https:\/\/i10x.ai\/blog\/ai-cv-bias\">earlier 42-point rewrite study<\/a>\nor\n<a href=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/i10x_resume_bias_study.pdf\" rel=\"noopener\" target=\"_blank\">download the paper (PDF)<\/a>.\n<\/p>\n\n<div class=\"i10x-callout\">\n<strong>The sentence vendors want. The sentence the file actually supports.<\/strong>\n<p>Vendors say a model is more objective than a recruiter because it does not see gender. On a frozen resume, these four models did not see gender in any way that changed a hire. They did see which chatbot wrote the file. GPT scored its own resumes 8.2 points below everyone else&#8217;s, for women and men alike. No recruiters sat in this design. The headline &#8220;AI is more objective than humans&#8221; is not a result. It is the slogan the experiment was asked to bless. It does not.<\/p>\n<\/div>\n\n<div class=\"i10x-highlight-stats\">\n<table class=\"i10x-table\">\n<tbody>\n<tr>\n<td><p><strong>0.24<\/strong><\/p><\/td>\n<td><p>Largest female-minus-male score gap, on a 0-100 scale (Grok)<\/p><\/td>\n<\/tr>\n<tr>\n<td><p><strong>12 pp<\/strong><\/p><\/td>\n<td><p>Hire-rate gap between Gemini-written (98%) and GPT-written (86%) resumes<\/p><\/td>\n<\/tr>\n<tr>\n<td><p><strong>68.7%<\/strong><\/p><\/td>\n<td><p>GPT hire rate on GPT-written resumes. The other three raters: 89% to 95%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p><strong>3,168<\/strong><\/p><\/td>\n<td><p>Parseable name-swap scores (of 3,200 calls). Pronoun arm: 4,713 of 4,800<\/p><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n\n<hr>\n\n<h2 id=\"for-editors\">For editors<\/h2>\n<div style=\"display:flex;flex-wrap:wrap;gap:0.75rem;margin:0.85rem 0 1.2rem\">\n<div style=\"flex:1 1 240px;padding:1rem 1.1rem;border:1px solid #e6eaf4;border-radius:1.05rem;background:linear-gradient(180deg,#f6f8ff 0%,#ffffff 100%)\">\n<p style=\"margin:0 0 0.35rem;font-size:0.78rem;font-weight:700;letter-spacing:.04em;text-transform:uppercase;color:#0052f5\">One-sentence finding<\/p>\n<p style=\"margin:0\">On a frozen resume, the first name did not decide the hire. The writing model did.<\/p>\n<\/div>\n<div style=\"flex:1 1 240px;padding:1rem 1.1rem;border:1px solid #e6eaf4;border-radius:1.05rem;background:linear-gradient(180deg,#f6f8ff 0%,#ffffff 100%)\">\n<p style=\"margin:0 0 0.35rem;font-size:0.78rem;font-weight:700;letter-spacing:.04em;text-transform:uppercase;color:#0052f5\">Study<\/p>\n<p style=\"margin:0\">Christopher Ort, i10X Research, 30 September 2026. 100 fictional personas. Four writers. Four raters. Paired name-swap plus a they\/them arm.<\/p>\n<\/div>\n<div style=\"flex:1 1 240px;padding:1rem 1.1rem;border:1px solid #e6eaf4;border-radius:1.05rem;background:linear-gradient(180deg,#f6f8ff 0%,#ffffff 100%)\">\n<p style=\"margin:0 0 0.35rem;font-size:0.78rem;font-weight:700;letter-spacing:.04em;text-transform:uppercase;color:#0052f5\">Cite \/ download<\/p>\n<p style=\"margin:0\"><a href=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/i10x_resume_bias_study.pdf\" rel=\"noopener\" target=\"_blank\">Paper PDF<\/a>. Primary page: this article. Related:\n<a href=\"https:\/\/i10x.ai\/blog\/ai-cv-bias\">42-point rewrite study<\/a>.<\/p>\n<\/div>\n<\/div>\n\n<hr>\n\n<h2 id=\"the-result\">The result, without the slogan<\/h2>\n<p>Employers already use language models to write resumes and to sort them. The public argument jumps to a slogan: a model is fairer than a human because it has no first impression. That slogan does two jobs. It claims models ignore identity cues. It also claims humans do not, so the model is the fairer reader.<\/p>\n<p>The second claim needs recruiters in the room. This study does not have them. Bertrand and Mullainathan&#8217;s 2004 field experiment on Emily, Greg, Lakisha, and Jamal remains the human audit. We did not rerun it. We tested the first claim the only way it can be tested: freeze the resume, change only the cue, keep every other byte identical.<\/p>\n<p>The cue did not move the decision. The writer did.<\/p>\n\n<figure class=\"i10x-figure\">\n<img fetchpriority=\"high\" decoding=\"async\" src=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png\" alt=\"Bar chart of hire rates by resume writer: Gemini 98 percent, Claude 96, xAI 91, GPT 86\" width=\"1600\" height=\"900\" loading=\"eager\">\n<figcaption><strong>Figure 1.<\/strong> Same candidate, different chatbot, different hire rate. Gemini-written resumes 98%. GPT-written 86%. i10X Research, 30 Sep 2026.<\/figcaption>\n<\/figure>\n\n<hr>\n\n<h2 id=\"how-we-tested\">How we tested<\/h2>\n<p>One hundred fictional personas, each locked to a job description: backend engineering, data analysis, nursing, finance, a head-of-product brief, and the rest of the set. Four systems had already written one resume each. Those files were frozen. The raters were not asked to rewrite.<\/p>\n<p>In the name-swap, the only edit was the name line. Same last name. Same body. A clearly female first name, then a clearly male first name. In the pronoun arm, one line went under the name: she\/her, he\/him, or they\/them. The they\/them arm kept the original gender-ambiguous first name. Female and male arms used the gendered name plus the matching pronoun. The body stayed put.<\/p>\n<p>That is the hole this paper closes. An earlier wave in the same project looked, at first pass, like a gender effect: one rater recommended female-coded names at about 65% and male-coded names at about 55%. Those resumes had been rewritten, not relabelled. Length, format, and wording moved with the name. Overlap with the original text was low (a Jaccard index of word sets around 0.27 for Claude and 0.41 for Gemini). Claude&#8217;s texts often became cover letters. That wave is a caution. It is not the gender result.<\/p>\n<p>Ratings used the cheapest current screening-class model of each provider. A firm sorting a pile does not call the flagship for every CV. It calls the cheap, fast model:<\/p>\n<ul>\n<li><strong>Claude Haiku 4.5<\/strong><\/li>\n<li><strong>GPT-6 Luna<\/strong><\/li>\n<li><strong>Gemini 3.1 Flash Lite<\/strong><\/li>\n<li><strong>Grok 4.3<\/strong><\/li>\n<\/ul>\n<p>Each rater saw the job and the resume and returned JSON only: an integer fit score from 0 to 100, and hire, maybe, or reject. Temperature was 0. The instruction said to judge fit and not to reward or penalize a name. A model that still moved with the name would have done so against that instruction.<\/p>\n<p>The gender estimate is paired: same persona, same writer, same rater, female minus male. A writer effect cannot be explained by the first name, because the pair shares the text.<\/p>\n\n<hr>\n\n<h2 id=\"the-name-does-not-decide\">The name does not decide<\/h2>\n<p>The largest mean gap is a quarter of a point, on a scale that in practice sits in the nineties. Hire rates match to a percentage point. The number of pairs in which only one name is hired is in the single digits.<\/p>\n<table class=\"i10x-table\">\n<tbody>\n<tr>\n<th><p>Rater<\/p><\/th>\n<th><p>n pairs<\/p><\/th>\n<th><p>Score gap (F minus M)<\/p><\/th>\n<th><p>Hire, female<\/p><\/th>\n<th><p>Hire, male<\/p><\/th>\n<th><p>Only female \/ only male<\/p><\/th>\n<\/tr>\n<tr>\n<td><p>GPT-6 Luna<\/p><\/td>\n<td><p>374<\/p><\/td>\n<td><p>+0.11<\/p><\/td>\n<td><p>87.7%<\/p><\/td>\n<td><p>87.7%<\/p><\/td>\n<td><p>1 \/ 1<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Claude Haiku 4.5<\/p><\/td>\n<td><p>400<\/p><\/td>\n<td><p>+0.13<\/p><\/td>\n<td><p>94.0%<\/p><\/td>\n<td><p>93.2%<\/p><\/td>\n<td><p>3 \/ 0<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Grok 4.3<\/p><\/td>\n<td><p>400<\/p><\/td>\n<td><p>+0.24<\/p><\/td>\n<td><p>93.5%<\/p><\/td>\n<td><p>93.0%<\/p><\/td>\n<td><p>3 \/ 1<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Gemini 3.1 Flash Lite<\/p><\/td>\n<td><p>400<\/p><\/td>\n<td><p>-0.01<\/p><\/td>\n<td><p>96.8%<\/p><\/td>\n<td><p>96.5%<\/p><\/td>\n<td><p>1 \/ 0<\/p><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The same picture holds inside each writer. No writer-by-rater cell exceeds half a point of female-minus-male difference. A quarter-point, even if stable, does not change a hire. Hire decisions agreed in 372 to 399 of about 400 pairs.<\/p>\n\n<figure class=\"i10x-figure\">\n<img decoding=\"async\" src=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig3-name-vs-authorship.png\" alt=\"Bar chart comparing a 0.24 point name gap with an 8.17 point GPT self-penalty\" width=\"1600\" height=\"900\" loading=\"lazy\">\n<figcaption><strong>Figure 2.<\/strong> The name moved 0.24 points. GPT&#8217;s penalty on its own writing moved 8.17. The self-penalty is gender-blind: -8.31 for female names, -8.04 for male names.<\/figcaption>\n<\/figure>\n\n<hr>\n\n<h2 id=\"they-them\">They\/them is not a penalty<\/h2>\n<p>Gemini&#8217;s hire decision is identical in 396 of 396 pairs. Claude&#8217;s only detectable score gap is -0.23 against the female arm: statistically noticeable, practically empty. Hire disagreements are single cases. They do not point against they\/them.<\/p>\n<table class=\"i10x-table\">\n<tbody>\n<tr>\n<th><p>Rater<\/p><\/th>\n<th><p>They\/them minus female<\/p><\/th>\n<th><p>They\/them minus male<\/p><\/th>\n<th><p>Hire, they\/them<\/p><\/th>\n<th><p>Hire, female \/ male<\/p><\/th>\n<\/tr>\n<tr>\n<td><p>GPT-6 Luna<\/p><\/td>\n<td><p>-0.05<\/p><\/td>\n<td><p>+0.14<\/p><\/td>\n<td><p>89%<\/p><\/td>\n<td><p>89% \/ 88%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Claude Haiku 4.5<\/p><\/td>\n<td><p>-0.23<\/p><\/td>\n<td><p>-0.03<\/p><\/td>\n<td><p>94%<\/p><\/td>\n<td><p>94% \/ 94%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Grok 4.3<\/p><\/td>\n<td><p>-0.04<\/p><\/td>\n<td><p>+0.11<\/p><\/td>\n<td><p>94%<\/p><\/td>\n<td><p>94% \/ 94%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Gemini 3.1 Flash Lite<\/p><\/td>\n<td><p>-0.01<\/p><\/td>\n<td><p>+0.06<\/p><\/td>\n<td><p>97.7%<\/p><\/td>\n<td><p>97.7% \/ 97.7%<\/p><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n\n<hr>\n\n<h2 id=\"the-writer-decides\">The writer decides<\/h2>\n<p>Pooling gender, resumes written by Gemini scored 95.4 and were hired at 98%. Claude-written resumes scored 94.6 and were hired at 96%. xAI-written resumes scored 90.2 and were hired at 91%. GPT-written resumes scored 89.7 and were hired at 86%. That ordering is the large effect in the file. It is several times the name gap. It is the effect a screening pile would actually feel.<\/p>\n<p>The cell that breaks the ceiling is GPT reading GPT: <strong>68.7% hire<\/strong>, against 89% to 95% when Claude, Grok, or Gemini read the same GPT texts.<\/p>\n<figure class=\"i10x-figure\">\n<img decoding=\"async\" src=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig2-writer-rater-heatmap.png\" alt=\"Heatmap of hire rates by resume writer and screening model, highlighting GPT on GPT at 68.7 percent\" width=\"1600\" height=\"900\" loading=\"lazy\">\n<figcaption><strong>Figure 3.<\/strong> Hire rate by writer (rows) and rater (columns), gender pooled. The red box is GPT screening a GPT-written resume.<\/figcaption>\n<\/figure>\n<table class=\"i10x-table\">\n<tbody>\n<tr>\n<th><p>Writer \\ rater<\/p><\/th>\n<th><p>GPT-6 Luna<\/p><\/th>\n<th><p>Haiku 4.5<\/p><\/th>\n<th><p>Grok 4.3<\/p><\/th>\n<th><p>Gemini 3.1 Flash Lite<\/p><\/th>\n<\/tr>\n<tr>\n<td><p>Claude writes<\/p><\/td>\n<td><p>93.9%<\/p><\/td>\n<td><p>97.0%<\/p><\/td>\n<td><p>95.0%<\/p><\/td>\n<td><p>98.0%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>GPT writes<\/p><\/td>\n<td><p><strong>68.7%<\/strong><\/p><\/td>\n<td><p>89.0%<\/p><\/td>\n<td><p>88.5%<\/p><\/td>\n<td><p>94.5%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>Gemini writes<\/p><\/td>\n<td><p>96.9%<\/p><\/td>\n<td><p>98.0%<\/p><\/td>\n<td><p>98.0%<\/p><\/td>\n<td><p>99.0%<\/p><\/td>\n<\/tr>\n<tr>\n<td><p>xAI writes<\/p><\/td>\n<td><p>86.6%<\/p><\/td>\n<td><p>90.5%<\/p><\/td>\n<td><p>91.5%<\/p><\/td>\n<td><p>95.0%<\/p><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>A candidate who has GPT draft the resume, and a firm that has GPT screen the inbox, is the case the file speaks to. In this test those resumes are hired by GPT at 69% and by the other three raters at 89% to 95%, with an eight-point score cut that does not depend on the name. The practical rule is not to let the writing model be the judging model.<\/p>\n\n<hr>\n\n<h2 id=\"self-preference\">Self-preference is model-specific, and gender-blind<\/h2>\n<p>GPT scored its own resumes 8.17 points below the others it read (85.68 against 93.85). The split is -8.31 for female names and -8.04 for male names. Claude scored its own 2.65 points higher (+2.72 female, +2.58 male). Gemini scored its own 2.46 points higher (+2.37 female, +2.55 male). Grok scored its own 3.06 points lower (-2.93 female, -3.19 male). Whatever these models are doing with authorship, they are not doing it differently for women and men.<\/p>\n<p>Gemini is also the most generous rater overall, with a mean score of 96.8 and a hire rate of 96.6% regardless of writer. Scores sit near the top of the scale. A small bias has little room. The authorship effect is large enough to be visible anyway.<\/p>\n<p>A firm that uses Claude or Gemini both to polish applications and to score them will, on these numbers, give its own polish a small lift of about two and a half points. That is not a scandal. It is a reason to separate the two steps, or to have a person read the boundary. Put that into an\n<a href=\"https:\/\/i10x.ai\/blog\/ai-resume-screening\">AI resume screening<\/a>\nworkflow and a\n<a href=\"https:\/\/i10x.ai\/blog\/multi-model-ai-screening\">multi-model panel<\/a>.<\/p>\n\n<hr>\n\n<h2 id=\"what-the-headline-cannot-mean\">What the headline cannot mean<\/h2>\n<div class=\"i10x-callout i10x-callout--quote\">\n<strong>Research statement<\/strong>\n<p><em>&#8220;In this design the first name and the pronoun did not decide the hire. Which model wrote the resume, and which model read it, did.&#8221;<\/em><\/p>\n<p>Christopher Ort, i10X Research<\/p>\n<\/div>\n<p>&#8220;AI is more objective than humans&#8221; is the sentence vendors reach for. This paper borrowed it as a title because it is the sentence the experiment is constantly asked to support. It does not support it.<\/p>\n<p>Objectivity toward a name is not objectivity. These four models did not hire Alexandra over Alexander, or she\/her over they\/them, on a frozen resume. They did hire a Gemini-written resume over a GPT-written resume of the same persona, and GPT docked its own writing by eight points. A human recruiter who ignored the name and then downgraded every application that &#8220;sounded like ChatGPT&#8221; would be open to the same charge. Without recruiters in the design, the comparison is rhetoric.<\/p>\n<p>The earlier, rewritten wave is a warning in the other direction. A gender gap appeared when the text was allowed to change with the name. Publishing that gap as a property of the model would have been a property of the rewrite. Everyday audits that regenerate the resume under a new name make the same mistake. A name-bias audit is only an audit if the resume is frozen.<\/p>\n\n<hr>\n\n<h2 id=\"monday\">What this means on Monday<\/h2>\n<div style=\"display:flex;flex-wrap:wrap;gap:0.75rem;margin:0.85rem 0 1.2rem\">\n<div style=\"flex:1 1 240px;padding:1rem 1.1rem;border:1px solid #e6eaf4;border-radius:1.05rem;background:linear-gradient(180deg,#f6f8ff 0%,#ffffff 100%)\">\n<p style=\"margin:0 0 0.35rem;font-size:0.78rem;font-weight:700;letter-spacing:.04em;text-transform:uppercase;color:#0052f5\">If you are a candidate<\/p>\n<p style=\"margin:0\">The first name is the wrong thing to worry about with these four cheap screeners. The writing model is not. If GPT wrote the CV and GPT screens the inbox, this file says 69% hire. The same GPT text, read by the other three, sits at 89% to 95%.<\/p>\n<\/div>\n<div style=\"flex:1 1 240px;padding:1rem 1.1rem;border:1px solid #e6eaf4;border-radius:1.05rem;background:linear-gradient(180deg,#f6f8ff 0%,#ffffff 100%)\">\n<p style=\"margin:0 0 0.35rem;font-size:0.78rem;font-weight:700;letter-spacing:.04em;text-transform:uppercase;color:#0052f5\">If you run screening<\/p>\n<p style=\"margin:0\">Do not let the writing model be the judging model. Freeze the resume before you audit names. Treat &#8220;maybe&#8221; as a process decision, not a kindness. A single cheap model on a pile is a single fingerprint.<\/p>\n<\/div>\n<div style=\"flex:1 1 240px;padding:1rem 1.1rem;border:1px solid #e6eaf4;border-radius:1.05rem;background:linear-gradient(180deg,#f6f8ff 0%,#ffffff 100%)\">\n<p style=\"margin:0 0 0.35rem;font-size:0.78rem;font-weight:700;letter-spacing:.04em;text-transform:uppercase;color:#0052f5\">If you sell &#8220;fairer than humans&#8221;<\/p>\n<p style=\"margin:0\">Name-blind is not human-fair. This study cannot say these models beat a recruiter. It can say the failure mode to watch in this stack is authorship, not the first name. See\n<a href=\"https:\/\/i10x.ai\/blog\/ethical-ai-recruiting\">ethical AI recruiting<\/a>.<\/p>\n<\/div>\n<\/div>\n<p>None of this says a human panel would have done worse, or better. It says the failure mode to watch is authorship. The\n<a href=\"https:\/\/i10x.ai\/blog\/ai-recruiting-guide\">AI recruiting guide<\/a>\nturns that into a 30\/60\/90. The\n<a href=\"https:\/\/i10x.ai\/blog\/ai-recruiting-workflow\">recruiting workflow<\/a>\nis the stage map.<\/p>\n\n<hr>\n\n<h2 id=\"limits\">Limits, stated once<\/h2>\n<p>The personas are fictional. The job descriptions were written for the study. A real pile has messier text, employment gaps, and names the model has seen in training. One rating was taken per cell, at temperature 0. A second draw could move a score by more than the name gap. It is unlikely to invent an eight-point authorship effect, and unlikely to erase one.<\/p>\n<p>The writer texts come from an earlier generation pass whose model versions were not pinned in the source files. The raters are pinned: Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3. A previous Claude rater, in the unpinned wave, hired GPT-written resumes at 42%. That cell is not replicated here. Results do not travel automatically to a flagship model, or to last year&#8217;s snapshot. The June i10X\n<a href=\"https:\/\/i10x.ai\/blog\/ai-cv-bias\">AI CV bias study<\/a>\nmeasured rewrite effects with different model versions. This paper measures frozen text.<\/p>\n<p>There is no human arm. There is no demographic cue other than the name and the pronoun line. Ethnicity, age, disability, and university prestige were not swapped. The score ceiling means a bias that only appears among near-equal candidates can hide. The authorship effect did not hide.<\/p>\n\n<hr>\n\n<h2 id=\"faq\">Frequently asked questions<\/h2>\n<div style=\"margin:0.85rem 0;padding:0.95rem 1.05rem;border:1px solid #e6eaf4;border-radius:0.95rem;background:#fafbff\">\n<h3 id=\"faq-gender\" style=\"margin:0 0 0.4rem\">Did the AI resume screeners show gender bias?<\/h3>\n<p style=\"margin:0\">Not on a frozen resume. The female-minus-male score gap was between -0.01 and +0.24 points. Hire rates matched to a percentage point. A they\/them line did not move the decision.<\/p>\n<\/div>\n<div style=\"margin:0.85rem 0;padding:0.95rem 1.05rem;border:1px solid #e6eaf4;border-radius:0.95rem;background:#fafbff\">\n<h3 id=\"faq-chatgpt\" style=\"margin:0 0 0.4rem\">Why does ChatGPT look worse here?<\/h3>\n<p style=\"margin:0\">GPT-written resumes were hired at 86% overall, against 98% for Gemini-written text of the same personas. When GPT-6 Luna screened GPT-written resumes, hire fell to 68.7%. GPT also scored its own writing 8.17 points below the other writers, for both names.<\/p>\n<\/div>\n<div style=\"margin:0.85rem 0;padding:0.95rem 1.05rem;border:1px solid #e6eaf4;border-radius:0.95rem;background:#fafbff\">\n<h3 id=\"faq-humans\" style=\"margin:0 0 0.4rem\">Is AI more objective than humans?<\/h3>\n<p style=\"margin:0\">This study cannot say that. No human raters took part. The title of the paper is the slogan vendors use. The finding is narrower: in this design, the name was the wrong suspect.<\/p>\n<\/div>\n<div style=\"margin:0.85rem 0;padding:0.95rem 1.05rem;border:1px solid #e6eaf4;border-radius:0.95rem;background:#fafbff\">\n<h3 id=\"faq-audit\" style=\"margin:0 0 0.4rem\">How should a company audit AI resume screening?<\/h3>\n<p style=\"margin:0\">Freeze the resume. Swap only the cue you claim to test. Do not ask the model to rewrite under a new name and then score the rewrite. Separate the writing model from the judging model. Use more than one rater. Details sit in\n<a href=\"https:\/\/i10x.ai\/blog\/ai-resume-screening\">AI resume screening<\/a>\nand\n<a href=\"https:\/\/i10x.ai\/blog\/multi-model-ai-screening\">multi-model AI screening<\/a>.<\/p>\n<\/div>\n<div style=\"margin:0.85rem 0;padding:0.95rem 1.05rem;border:1px solid #e6eaf4;border-radius:0.95rem;background:#fafbff\">\n<h3 id=\"faq-paper\" style=\"margin:0 0 0.4rem\">Where is the paper?<\/h3>\n<p style=\"margin:0\"><a href=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/i10x_resume_bias_study.pdf\" rel=\"noopener\" target=\"_blank\">Download the PDF<\/a>. Source tables in the paper: namenstausch_bewertung.csv and they_them_bewertung.csv. Empty model replies were dropped, not imputed.<\/p>\n<\/div>\n\n<hr>\n\n<div class=\"i10x-cta\">\n<h3 id=\"try-i10x\">Run the same resume through more than one model<\/h3>\n<p>The file&#8217;s rule is simple: do not let the writing model be the judging model. Open chat on the homepage to sign up, then put the same CV in front of a second reader.<\/p>\n<p><a href=\"https:\/\/i10x.ai\/\" rel=\"noopener\" target=\"_blank\">Start on i10X \u2192<\/a><\/p>\n<p><a href=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/i10x_resume_bias_study.pdf\" rel=\"noopener\" target=\"_blank\">Download the study PDF<\/a> \u00b7\n<a href=\"https:\/\/i10x.ai\/blog\/ai-cv-bias\">42-point rewrite study<\/a> \u00b7\n<a href=\"https:\/\/i10x.ai\/blog\/ai-resume-screening\">AI resume screening<\/a> \u00b7\n<a href=\"https:\/\/i10x.ai\/tools\/category\/business-management\/free-ai-recruiting\">Free AI recruiting on i10X<\/a><\/p>\n<\/div>\n\n<div class=\"i10x-sources\">\n<strong>Sources<\/strong>\n<ol>\n<li>Christopher Ort, i10X Research Team (30 September 2026), &#8220;AI Is More Objective Than Humans&#8221;: paired name-swap and pronoun-arm resume screening study. 100 fictional personas, four frozen writers, four screening-class raters (Claude Haiku 4.5, GPT-6 Luna, Gemini 3.1 Flash Lite, Grok 4.3). 3,168 parseable name-swap scores; 4,713 parseable pronoun-arm scores. <a href=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/i10x_resume_bias_study.pdf\">PDF<\/a>.<\/li>\n<li>Bertrand, M. and Mullainathan, S. (2004). Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination. American Economic Review, 94(4), 991-1013. Cited as the human audit this study does not rerun, not as a benchmark these models beat.<\/li>\n<li>i10X Research (May to June 2026), <a href=\"https:\/\/i10x.ai\/blog\/ai-cv-bias\">AI CV bias in resume screening<\/a>: rewrite-based hire-rate gaps with different model versions. Distinct design from the frozen name-swap reported here.<\/li>\n<\/ol>\n<\/div>\n\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Four AI screeners did not hire by gender on frozen resumes. They hired Gemini-written CVs at 98% and GPT-written CVs at 86%. Name-swap study.<\/p>\n","protected":false},"author":5,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_yoast_wpseo_focuskw":"","_yoast_wpseo_title":"","_yoast_wpseo_metadesc":"","footnotes":""},"categories":[10],"tags":[],"class_list":["post-1000","post","type-post","status-draft","format-standard","hentry","category-ai"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.8 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>AI Resume Screeners Ignore Gender. They Don&#039;t Ignore ChatGPT - i10X Blog<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/i10x.ai\/blog\/?p=1000\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"AI Resume Screeners Ignore Gender. They Don&#039;t Ignore ChatGPT - i10X Blog\" \/>\n<meta property=\"og:description\" content=\"Four AI screeners did not hire by gender on frozen resumes. They hired Gemini-written CVs at 98% and GPT-written CVs at 86%. Name-swap study.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/i10x.ai\/blog\/?p=1000\" \/>\n<meta property=\"og:site_name\" content=\"i10X Blog\" \/>\n<meta property=\"article:published_time\" content=\"-0001-11-30T00:00:00+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png\" \/>\n<meta name=\"author\" content=\"Christopher Ort\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Christopher Ort\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"11 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000\"},\"author\":{\"name\":\"Christopher Ort\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/#\\\/schema\\\/person\\\/c5af13ca4e2bbda197660fab76672060\"},\"headline\":\"AI Resume Screeners Ignore Gender. They Don&#8217;t Ignore ChatGPT\",\"datePublished\":\"-0001-11-30T00:00:00+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000\"},\"wordCount\":2294,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/fig1-hire-by-writer.png\",\"articleSection\":[\"AI\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000\",\"url\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000\",\"name\":\"AI Resume Screeners Ignore Gender. They Don't Ignore ChatGPT - i10X Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/fig1-hire-by-writer.png\",\"datePublished\":\"-0001-11-30T00:00:00+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/#\\\/schema\\\/person\\\/c5af13ca4e2bbda197660fab76672060\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#primaryimage\",\"url\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/fig1-hire-by-writer.png\",\"contentUrl\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/fig1-hire-by-writer.png\",\"width\":2560,\"height\":1440,\"caption\":\"Figure 1. Same candidate, different chatbot, different hire rate.\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?p=1000#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/i10x.ai\\\/blog\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"AI Resume Screeners Ignore Gender. They Don&#8217;t Ignore ChatGPT\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/\",\"name\":\"i10X Blog\",\"description\":\"Model comparisons, workspace guides, and practical ideas on AI productivity, agents, and multi-model work.\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/#\\\/schema\\\/person\\\/c5af13ca4e2bbda197660fab76672060\",\"name\":\"Christopher Ort\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/ea95f3291658df6863df50e0ba53ddde5c83538e2079f4b3b9b548cc92d90cca?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/ea95f3291658df6863df50e0ba53ddde5c83538e2079f4b3b9b548cc92d90cca?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/ea95f3291658df6863df50e0ba53ddde5c83538e2079f4b3b9b548cc92d90cca?s=96&d=mm&r=g\",\"caption\":\"Christopher Ort\"},\"url\":\"https:\\\/\\\/i10x.ai\\\/blog\\\/author\\\/christopher-ort\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"AI Resume Screeners Ignore Gender. They Don't Ignore ChatGPT - i10X Blog","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/i10x.ai\/blog\/?p=1000","og_locale":"en_US","og_type":"article","og_title":"AI Resume Screeners Ignore Gender. They Don't Ignore ChatGPT - i10X Blog","og_description":"Four AI screeners did not hire by gender on frozen resumes. They hired Gemini-written CVs at 98% and GPT-written CVs at 86%. Name-swap study.","og_url":"https:\/\/i10x.ai\/blog\/?p=1000","og_site_name":"i10X Blog","article_published_time":"-0001-11-30T00:00:00+00:00","og_image":[{"url":"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png","type":"","width":"","height":""}],"author":"Christopher Ort","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Christopher Ort","Est. reading time":"11 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/i10x.ai\/blog\/?p=1000#article","isPartOf":{"@id":"https:\/\/i10x.ai\/blog\/?p=1000"},"author":{"name":"Christopher Ort","@id":"https:\/\/i10x.ai\/blog\/#\/schema\/person\/c5af13ca4e2bbda197660fab76672060"},"headline":"AI Resume Screeners Ignore Gender. They Don&#8217;t Ignore ChatGPT","datePublished":"-0001-11-30T00:00:00+00:00","mainEntityOfPage":{"@id":"https:\/\/i10x.ai\/blog\/?p=1000"},"wordCount":2294,"commentCount":0,"image":{"@id":"https:\/\/i10x.ai\/blog\/?p=1000#primaryimage"},"thumbnailUrl":"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png","articleSection":["AI"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/i10x.ai\/blog\/?p=1000#respond"]}]},{"@type":"WebPage","@id":"https:\/\/i10x.ai\/blog\/?p=1000","url":"https:\/\/i10x.ai\/blog\/?p=1000","name":"AI Resume Screeners Ignore Gender. They Don't Ignore ChatGPT - i10X Blog","isPartOf":{"@id":"https:\/\/i10x.ai\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/i10x.ai\/blog\/?p=1000#primaryimage"},"image":{"@id":"https:\/\/i10x.ai\/blog\/?p=1000#primaryimage"},"thumbnailUrl":"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png","datePublished":"-0001-11-30T00:00:00+00:00","author":{"@id":"https:\/\/i10x.ai\/blog\/#\/schema\/person\/c5af13ca4e2bbda197660fab76672060"},"breadcrumb":{"@id":"https:\/\/i10x.ai\/blog\/?p=1000#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/i10x.ai\/blog\/?p=1000"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/i10x.ai\/blog\/?p=1000#primaryimage","url":"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png","contentUrl":"https:\/\/i10x.ai\/blog\/wp-content\/uploads\/2026\/10\/fig1-hire-by-writer.png","width":2560,"height":1440,"caption":"Figure 1. Same candidate, different chatbot, different hire rate."},{"@type":"BreadcrumbList","@id":"https:\/\/i10x.ai\/blog\/?p=1000#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/i10x.ai\/blog"},{"@type":"ListItem","position":2,"name":"AI Resume Screeners Ignore Gender. They Don&#8217;t Ignore ChatGPT"}]},{"@type":"WebSite","@id":"https:\/\/i10x.ai\/blog\/#website","url":"https:\/\/i10x.ai\/blog\/","name":"i10X Blog","description":"Model comparisons, workspace guides, and practical ideas on AI productivity, agents, and multi-model work.","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/i10x.ai\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/i10x.ai\/blog\/#\/schema\/person\/c5af13ca4e2bbda197660fab76672060","name":"Christopher Ort","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/ea95f3291658df6863df50e0ba53ddde5c83538e2079f4b3b9b548cc92d90cca?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/ea95f3291658df6863df50e0ba53ddde5c83538e2079f4b3b9b548cc92d90cca?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/ea95f3291658df6863df50e0ba53ddde5c83538e2079f4b3b9b548cc92d90cca?s=96&d=mm&r=g","caption":"Christopher Ort"},"url":"https:\/\/i10x.ai\/blog\/author\/christopher-ort"}]}},"_links":{"self":[{"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/posts\/1000","targetHints":{"allow":["GET","POST","PUT","PATCH","DELETE"]}}],"collection":[{"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/users\/5"}],"replies":[{"embeddable":true,"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/comments?post=1000"}],"version-history":[{"count":0,"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/posts\/1000\/revisions"}],"wp:attachment":[{"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/media?parent=1000"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/categories?post=1000"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/i10x.ai\/blog\/wp-json\/wp\/v2\/tags?post=1000"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}