Guide · August 2026
An AI job description is only as good as the intake behind it. When hiring managers dump vague notes into a model, you get confident fiction: inflated years, copy-pasted culture slogans, and screen criteria that reject the people you actually need. This guide gives a full Hiring Intake Agent questionnaire, a JD to scorecard to screen criteria chain with a worked mid backend engineer example, an inclusive language red-flag table with rewrites, prestige proxy traps, must-have vs nice-to-have weighting, and a copy-paste intake agent prompt. Run the pattern with i10X Free AI Recruiting and keep fairness controls from the i10X CV bias study in view when screening starts.
43% |
Organizations using AI in HR tasks (SHRM, 2025), up from 26% in 2024 |
~44 days |
Median time-to-fill, non-executive (SHRM, 2025) |
42 pp |
Max hire-rate gap, same candidate, different AI resume style (i10X Research) |
Chain |
Intake answers → JD draft → scorecard → screen criteria (never reverse) |
Why AI job descriptions fail (bad intake, not bad models)
Teams blame the model when the JD is wrong. The failure usually starts earlier: incomplete intake. The hiring manager wants “a senior generalist who can wear many hats,” the model invents five years of Kubernetes and a computer science degree, and the screening agent ranks applicants against a role that does not exist.
AI job description generation is popular because drafting is tedious. SHRM reports 43% of organizations used AI in HR tasks in 2025, up from 26% in 2024. Drafting is a good use of that adoption. Unreviewed drafts as source of truth are not.
Three failure modes show up constantly:
- Priority inversion: nice-to-haves written as must-haves, so AI screeners over-filter.
- Proxy language: “culture fit,” “native English,” “young and energetic,” pedigree cues that invite bias and legal risk.
- Stack fiction: tools listed because they sound modern, not because the job uses them weekly.
Median time-to-fill sits near 44 days for non-executive roles (SHRM, 2025), with SHRM 2026 executives benchmarking materials citing about 39 days median non-executive in that context. Every week spent interviewing against a false JD is a week you do not get back. Cost-per-hire averages (SHRM 2025: about $5,475 non-exec, $35,879 exec) make re-opened searches painful.
For lifecycle context, see the AI recruiting guide and AI recruiting workflow.
Why a bad JD poisons AI matching
Screening systems, whether keyword ATS rules or LLM rankers, treat the JD or its derived criteria as ground truth. Garbage in still produces ranked garbage out, only faster and with more confidence. The poison path is simple:
- Manager gives incomplete notes.
- Model invents requirements to sound complete.
- Humans lightly edit tone but leave false must-haves.
- AI screen criteria inherit the fiction.
- Strong transfer candidates fail keyword or tenure filters.
- Interviews start late or with the wrong bar; time-to-fill drifts toward multi-week medians.
JD defect |
What AI screening tends to do |
Business result |
|---|---|---|
Laundry list of tools |
Keyword-matches tool names over impact |
Strong problem-solvers ranked low |
Inflated years (“10+ years React” for a 4-year-old API) |
Hard filters on tenure |
False rejects of capable mid-levels |
Vague rockstar / ninja / ownership slogans |
Rewards flashy resume language |
Style bias; see i10X study gaps |
Degree required without job need |
Overweights education section |
Misses self-taught and bootcamp talent |
Culture slogans without behaviors |
Invented soft-score rationales |
Inconsistent advances; hard to audit |
Copy-paste from another company |
Matches a different org’s reality |
Interview surprise and early attrition risk |
Prestige proxies (top school, logo employers only) |
Scores pedigree as skill |
Homogeneous shortlists; missed operators |
i10X Research shows evaluation is already unstable across writing styles: up to a 42 percentage-point hire-rate gap for the same person by resume-writer model (100 profiles, 1,576 evaluation points), and a 29 point evaluator gap on identical documents. A sloppy JD multiplies that instability. Fair screening design: AI resume screening and AI CV bias study.
Never let the model invent must-haves. The hiring manager locks dealbreakers in intake. The model drafts language and structure. The scorecard, not the marketing JD alone, drives AI screen criteria.
Framework 1: Hiring Intake Agent pattern
A Hiring Intake Agent is a structured workflow (human + AI) that turns a messy manager conversation into three artifacts: a publishable JD, a hiring scorecard, and machine-usable screen criteria. It is not an unsupervised bot that emails candidates. It is an intake interview with memory and a fixed output schema.
Think of the agent as a disciplined recruiter who always asks the same hard questions and refuses to draft until answers are complete.
Full Hiring Intake Agent questionnaire (use every time)
Ask all blocks. Mark unknowns as blockers, not guesses. Fifteen questions cover outcomes, bar, tools, constraints, and fairness.
# |
Question |
Why AI needs it |
|---|---|---|
1 |
What will this person deliver in 30 / 90 / 180 days? What does success look like in a demo or metric? |
Grounds JD bullets in work, not titles |
2 |
Which skills are true dealbreakers? What evidence counts (project, stack, domain)? Cap at 3 to 6. |
Prevents average-away scoring |
3 |
What is helpful but trainable? Assign weight 1 to 5 for each nice-to-have. |
Stops laundry-list filters |
4 |
IC vs lead? Scope of decisions? Who do they influence or manage? |
Aligns seniority language |
5 |
What tools and systems are used weekly vs monthly vs nice if known? |
Avoids stack fiction |
6 |
Location, remote rules, work authorization, travel, on-call, language needs that are truly required? |
Legal and operational clarity |
7 |
Compensation band, bonus, equity, benefits highlights you will publish or share on first call? |
Reduces ghosting and wasted screens |
8 |
What must screeners ignore (school prestige, photo, age cues, hobbies, logo employers alone)? |
Bias control for AI prompts |
9 |
Interview stages, owners, and what each stage uniquely tests? |
Prevents duplicate questions and bar drift |
10 |
Where have great people come from? Communities, companies, portfolios? (hints only) |
Feeds sourcing without inventing contacts |
11 |
What would make a candidate a clear no after 15 minutes of screen? |
Sharpens early must-have tests |
12 |
What is often overweighted by managers historically that we should de-weight? |
Surfaces prestige and vibe traps |
13 |
Any legal or credential requirements that are truly mandatory (license, clearance, degree for regulated work)? |
Separates real credentials from habit |
14 |
Who approves the final must-have list and scorecard version? |
Creates accountability before AI ranks people |
15 |
What will we publish vs keep internal (comp, headcount, roadmap secrets)? |
Prevents oversharing and under-informing |
Agent steps (copy into your workspace)
- Collect: run the questionnaire live or async. Mark unknowns as blockers, not guesses.
- Normalize: AI rewrites answers into structured fields (must / nice / evidence / constraints / non-criteria / weights).
- Confirm: hiring manager approves the structured fields before any public JD.
- Draft JD: AI produces a candidate-facing JD from approved fields only.
- Draft scorecard: AI produces interviewer-facing scorecard with weights and anchors.
- Derive screen criteria: AI produces machine-usable pass/fail and ranking rules for resume screening.
- Inclusive pass: run the red-flag table and edit.
- Version and store: scorecard ID plus prompt version on the requisition.
- Only then: post the job and turn on AI-assisted ranking.
Operationalize these steps inside i10X Free AI Recruiting or as an AI recruiting agent pattern with human confirmation gates.
Must-have vs nice-to-have weighting
Weighting is where most intakes fail. Managers say everything is critical. AI then enforces a fantasy bar.
Class |
Definition |
Weight rule |
AI behavior |
|---|---|---|---|
Must-have |
Without this, the person cannot do the job in a reasonable ramp |
Binary pass/fail; 3 to 6 max |
Fail any must-have = do not rank high; still human-review rejects |
Nice-to-have |
Speeds ramp or covers secondary work; trainable |
Weight 1 to 5; cannot override a must-have fail |
Adds ranking points only among must-have passers |
Narrative only |
Context for attraction, not evaluation |
Weight 0 |
Never becomes a screen filter |
Non-criteria |
Explicit ignore list |
N/A |
Prompt bans scoring on these |
Forced cut exercise: If the manager lists 12 must-haves, ask: “If only three remain, which three close the business risk this quarter?” Put the rest in nice-to-have or drop them. LinkedIn Future of Recruiting 2025 notes about 37% of talent professionals integrating or experimenting with generative AI (up from 27%), with ~20% workweek savings context for gen AI users. That time savings is wasted if the model is filtering on a wish list.
Prestige proxy traps
Prestige proxies are stand-ins for skill that feel rigorous and measure poorly. AI is excellent at amplifying them because logos and school names are easy tokens.
- School rank as ability: unless a credential is legally required, prefer work evidence.
- Logo employers only: large brand experience can help or hurt; require transferable outcomes.
- Years as mastery: ten years of shallow use loses to three years of ownership.
- Keyword density: resumes stuffed with tools outrank quieter high performers.
- Writing polish as intelligence: i10X’s style-driven gaps show this risk clearly.
- Culture fit as sameness: replace with observable collaboration and reliability behaviors.
Put prestige proxies in the non-criteria list during intake. If a manager insists, demand a job-related reason in writing before it becomes a must-have.
Framework 2: JD → scorecard → screen criteria chain
Keep the chain one-directional. Do not reverse it (do not invent scorecard must-haves from a marketing JD the model embellished).
Artifact 1: candidate-facing JD
Purpose: attract and set expectations. Tone: clear, specific, respectful. Include mission of the role, outcomes for the first two quarters, responsibilities as verbs tied to real work, requirements split into required vs preferred, logistics you will not negotiate away, and comp range when you can publish it.
Artifact 2: hiring scorecard
Purpose: human interview consistency. Include dimensions mapped to must-haves and critical nice-to-haves, evidence examples for strong / mixed / weak, explicit non-criteria, and interview stage coverage so each loop tests something new.
Artifact 3: AI screen criteria
Purpose: first-pass ranking with auditability. Include binary must-have checks with acceptable evidence forms, weighted nice-to-haves that cannot override a failed must-have, instructions to quote evidence from the materials, a ban on demographic inference and style worship, and a human gate before reject.
Worked scenario: mid backend engineer chain
Intake snapshot (approved): 90-day outcomes include shipping one production service improvement and owning on-call for that service with clean handoffs. Must-haves: production backend service ownership; weekly relational SQL; debugging production issues with logs and metrics. Nice-to-haves: Kafka (weight 3), mentoring 1 to 2 engineers (weight 2). Non-criteria: school brand, FAANG-only logos, photo, age cues. Stack weekly: Go or Java OK if transfer clear; Postgres daily. Comp band approved for first-call share. Stages: screen → technical → manager → values.
Field |
JD language |
Scorecard dimension |
Screen criterion |
|---|---|---|---|
Core skill |
Build and operate services in our primary stack |
Service ownership: design, ship, on-call hygiene |
Must: evidence of shipping backend services in production (listed stack or clear transfer) |
Data |
We use Postgres daily; Kafka is a plus |
Data stores: practical SQL depth |
Must: relational DB experience; Nice (3): streaming/event systems |
Level |
You may mentor one to two engineers |
Mentorship: feedback and unblocking |
Nice (2): mentorship examples; not a hard filter on IC track |
Ops |
Participate in on-call with runbooks |
Production debugging with metrics and logs |
Must: production incident or on-call evidence |
Non-criterion |
(omit prestige language) |
Ignore school brand unless credential truly required |
Prompt: do not score on university rank or logo employers alone |
Chain checklist
- Every must-have in screen criteria appears in intake approvals.
- Every preferred line in the JD has a weight or is marked narrative-only.
- Scorecard dimensions map to interview questions already assigned to owners.
- Version numbers match across ATS, prompt library, and intake notes.
- Inclusive red-flag pass completed before posting.
Inclusive language red-flag table (with rewrites)
Use this table when reviewing AI-drafted JDs. Replace flags with job-related wording. This is practical editing guidance, not legal advice.
Red flag in draft |
Why it is a problem |
Prefer instead |
|---|---|---|
Rockstar, ninja, guru, superhero |
Vague, culturally loaded, rewards swagger over evidence |
Name the skill and outcome (ships production APIs with tests) |
Young and energetic / digital native |
Age-coded language |
Describe pace and tools actually used |
Culture fit |
Often a proxy for sameness |
Culture add plus specific behaviors (feedback, reliability, collaboration) |
Native English speaker |
Can unlawfully screen by national origin in many contexts |
Define communication needs of the role (e.g. client-facing writing) without nativeness |
Must have degree from top school |
Pedigree proxy; rarely job-related |
Required credential only if legally or truly necessary; else skills evidence |
No career gaps |
Penalizes caregiving, health, layoffs |
Evaluate recent, relevant evidence of skill |
Work hard / play hard, always on |
Signals unsustainable norms; can deter caregivers |
State on-call or travel honestly with boundaries |
Strong man / salesman archetypes |
Gendered coding in many languages and metaphors |
Neutral role nouns; skill-based adjectives |
Unlimited overtime expected |
May conflict with wage/hour reality and inclusion goals |
Peak seasons and time-off norms stated clearly |
Copy-pasted diversity paragraph only |
Empty if requirements still exclude |
Inclusive requirements design plus EEO statement your counsel approves |
Aggressive / thrives under pressure only |
Can code for dominance culture and deter capable quiet performers |
Describe real pressure sources (deadlines, incidents) and support norms |
Must be a culture carrier from day one |
Vague and often biased toward insiders |
List onboarding outcomes and collaboration behaviors |
Inclusive edit checklist
- Required section is short enough that a strong candidate can say yes to all of it.
- Preferred section is clearly optional.
- No jokes, memes, or beer-centric culture bait as requirements.
- Accommodation and application accessibility notes exist where you can offer them.
- AI draft compared against intake non-criteria list before publish.
Copy-paste template: intake agent prompt
You are a Hiring Intake Agent. Do not invent requirements. STEP 1: Ask any missing questions from this list until complete: [paste questionnaire 1-15] STEP 2: Normalize answers into fields: - Role title, level, team - Outcomes 30/90/180 - Must-haves (max 6) with evidence examples - Nice-to-haves with weights 1-5 - Weekly tools vs monthly vs nice-if-known - Constraints (location, auth, on-call, travel, language if required) - Comp band (publish vs first-call) - Non-criteria (explicit ignore list) - Interview stages and owners - Sourcing hints (no invented contacts) - Open [GAP] items STEP 3: Stop and ask the hiring manager to APPROVE the structured fields. Do not draft a JD until approval is explicit. STEP 4 (after approval only): Draft: A) Candidate-facing JD (Required vs Preferred; no new tools/years/degrees) B) Scorecard (dimensions, strong/mixed/weak anchors, non-criteria) C) AI screen criteria (binary must-haves, weighted nice-to-haves, quote evidence, ban demographic inference, ban style-only scoring, human gate before reject) If intake is incomplete, output [GAP] instead of guessing. Version label: scorecard-vX / prompt-vY.
Related prompt stubs
Draft JD: “Using only the approved intake fields below, write a job description. Do not add tools, years, degrees, or soft requirements that are not in the fields. Split Required vs Preferred. Write outcomes for 90 days. Flag any intake gap as [GAP].”
Scorecard: “Convert approved must-haves and nice-to-haves into a scorecard with dimensions, weights, and strong/mixed/weak evidence. List non-criteria. Do not introduce new must-haves.”
Screen criteria: “Produce first-pass screen criteria: binary must-haves with evidence rules, weighted nice-to-haves, output schema (pass/fail, quotes, risks, rationale). Ban demographic inference. Ban scoring on writing polish alone. Human review required before reject.”
Version these prompts. When the i10X study shows large swings from style and evaluator choice, stable criteria and human gates are how you stay defensible. Ethics deep dive: ethical AI recruiting.
20-minute hiring manager intake script
- Minutes 0 to 3: “What does success look like at day 90?”
- Minutes 3 to 8: must-haves vs trainables (force a cut list and weights).
- Minutes 8 to 12: weekly tools and real constraints (location, auth, on-call).
- Minutes 12 to 16: interview loop ownership and what each stage tests.
- Minutes 16 to 20: non-criteria, comp band, prestige proxies to ban, posting date. Confirm AI will not invent fields.
After the meeting, run the agent steps the same day while context is fresh. Delayed intake is how JDs rot into generic templates. For small-team operating constraints, pair with AI recruiting for startups.
Anti-patterns and failure modes
- Reverse chain: polishing a marketing JD first, then inventing scorecard must-haves to match the hype.
- Wish-list must-haves: twelve dealbreakers that no human on the team would meet.
- Copy-paste JD from a unicorn careers page: different product, different constraints, same fiction.
- Skipping non-criteria: model scores school brand and logo density by default.
- Years inflation: requiring decade-long tenure on young tools.
- Inclusive paragraph without inclusive requirements: brand risk and empty trust.
- Async intake without confirmation: manager never locks fields; AI keeps guessing.
- Screen criteria broader than scorecard: candidates pass interviews on different rules than the resume gate.
- No version IDs: three prompt variants ranking the same req with different bars.
- Publishing before inclusive pass: red-flag language goes live and into AI filters.
When NOT to use AI for this step
- Locking must-haves: the hiring manager decides dealbreakers; AI only structures them.
- Legal language finalization: EEO statements, disclaimers, and regulated credential text need counsel-approved templates.
- Compensation policy decisions: AI can format a band you already approved; it should not invent pay ranges.
- When scope is still political: if two leaders disagree on the role, facilitate humans first.
- Auto-posting JDs without human inclusive pass and field approval.
- Deriving rejects from an unapproved draft JD.
- Translating vague culture slogans into scored soft skills without behavioral anchors.
Governance notes (brief, not legal advice)
Under the EU AI Act, Annex III treats certain recruitment and selection AI (including targeted job ads, filtering applications, and evaluating candidates) as high-risk use cases, with obligations ramping through 2026 to 2027. A clean intake chain helps: you can show job-relatedness, versioned criteria, and human oversight when AI later ranks people. Drafting a JD with a human publisher is a different risk posture than auto-filtering applicants, but the quality of the JD still drives fairer downstream automation.
LinkedIn Future of Recruiting 2025 finds about 37% of talent professionals integrating or experimenting with generative AI (up from 27%), with ~20% workweek savings context for gen AI users, and +9% quality of hire for heavy AI-Assisted Messaging users in LinkedIn’s messaging research. Better intake multiplies those gains because outreach and screening point at the same role definition.
For startups and scaled TA teams
Startups: run the questionnaire in a doc, generate artifacts once, and refuse to post until the manager signs the must-have list. See AI recruiting for startups and free AI recruiting tools 2026.
Scaled teams: store intake fields in the ATS, block req approval without scorecard ID, and sample JDs quarterly for red flags. Agents can prefill from past similar roles, but each req still needs confirmation. Workflow design: AI recruiting workflow. Architecture when AI sits beside the ATS: AI recruiting vs ATS.
End-to-end AI job description checklist
- Intake questionnaire complete; no critical [GAP] left inventable.
- Hiring manager signed must-haves, weights, and non-criteria.
- JD draft uses only approved fields; inclusive red-flag pass done.
- Scorecard and screen criteria generated and versioned.
- Interview stages mapped to dimensions.
- Prestige proxies explicitly banned in prompts.
- Human gate defined before any AI reject path.
- Sourcing notes captured for outbound (no invented emails).
- Artifacts stored on the requisition; prompt library updated.
- Only then: post, screen, message.
Frequently asked questions
What is an AI job description?
A job description drafted or refined with generative AI. Quality depends on intake quality. The model should structure and phrase approved facts, not invent requirements.
Can AI write the JD by itself?
It can draft. It should not invent must-haves, years, tools, or credentials. Without a completed intake and human approval, AI-written JDs poison matching downstream.
How long should intake take?
A focused live session can finish in about 20 minutes if the manager is prepared. Async questionnaires work if unknowns stay marked as [GAP] and someone chases blockers the same day. Delaying intake to “save time” usually costs weeks in false filters.
Should AI write the JD or the scorecard first?
Neither first. Intake fields first. Then draft JD and scorecard from the same approved fields. Screen criteria come last in the chain.
What legal language should AI put in a JD?
Use counsel-approved EEO and local disclaimer templates. AI can place them; humans own legal accuracy. This guide is not legal advice.
Should salary be in the JD?
When you can publish a band, do it. It reduces mismatch and wasted screens. If policy forbids public comp, train first-call scripts with the approved band and keep it in intake fields for recruiters.
Why do AI matches look wrong even with a polished JD?
Polished prose can still encode wrong must-haves. Screening AI optimizes for the criteria you gave it. Fix the chain. Also watch style bias when ranking resumes (
i10X study).
How detailed should must-haves be?
Few, testable, and truly disqualifying. If you have twelve must-haves, you probably have a wish list. AI will enforce your over-filtering at scale.
Can we reuse one JD template for every eng role?
Templates for structure yes. Content no. Reusing another team’s stack list is how fiction enters the pipeline.
How does this relate to ethical AI recruiting?
Job-related, versioned criteria and inclusive language are fairness controls. See
ethical AI recruiting
for scorecards and EU-oriented checklists.
Where should small teams run the intake agent?
In one approved workspace so CVs and notes do not scatter. Try
Free AI Recruiting
with a fixed questionnaire prompt pack.
Does better intake reduce time-to-fill?
It reduces wasted loops and false filters, which is how teams lose weeks inside a ~40 day median fill context (SHRM). Intake is not delay. It is acceleration with fewer do-overs.
How do prestige proxies sneak into AI screening?
Through JD language and unstated manager preferences. Put them in non-criteria, ban them in screen prompts, and sample low scores for pedigree-only rejects.
What is the difference between nice-to-have weight 5 and a must-have?
A must-have fails the candidate for the role. A weight-5 nice-to-have ranks higher among people who already pass must-haves. If weight 5 is truly required, promote it to must-have and cut something else.
AI job descriptions work when a Hiring Intake Agent forces complete answers, a one-way JD to scorecard to screen criteria chain prevents invented must-haves, weights separate dealbreakers from trainables, prestige proxies are banned explicitly, and an inclusive red-flag table cleans language before AI ranking begins. Bad JDs create bad AI matches at speed. Fix intake first.
Run intake to shortlist in one workspace
Turn manager notes into structured criteria, then screen and shortlist with human-ready outputs instead of a pile of disconnected prompts.
Launch Free AI Recruiting →- SHRM (2025): AI in HR (43%, up from 26% in 2024); cost-per-hire averages (~$5,475 non-executive, ~$35,879 executive); time-to-fill (~44 days median non-executive).
- SHRM 2026 executives benchmarking materials: non-executive median time-to-fill context (~39 days).
- LinkedIn Future of Recruiting (2025): gen AI integrate/experiment (37%, from 27%); ~20% workweek savings context; AI-Assisted Messaging heavy users (+9% quality of hire).
- i10X Research: CV bias study (100 profiles, 1,576 evaluation points; up to 42 pp hire-rate gap by resume-writer model; 29 pt evaluator gap). Study.
- EU AI Act Annex III: high-risk AI for recruitment and selection (targeted job ads, analysing and filtering applications, evaluating candidates). Obligations ramp 2026 to 2027. Not legal advice.
- i10X Blog: AI recruiting guide, AI resume screening, AI recruiting workflow, AI recruiting agents, AI candidate sourcing, Ethical AI recruiting, AI recruiting for startups, AI recruiting vs ATS, Free AI recruiting tools 2026.



