Recipe 1 #
Let the model's quick answer make the decision
What the organization does
Ask the model one broad question about each applicant, such as "attorney or paralegal?", and act on its answer straight away. The applicant is shortlisted, rejected or sent elsewhere. Nobody looks at anything else, and nobody checks.
The regulated decision it hits
Hiring and candidate screening. Results on these pages: gender, disability.
Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a), ADA Title I.
What our tests show
The answer changes when nothing but the pronouns change. On the attorney-or-paralegal question, the model can read the same biography as a paralegal's when it says "she" and as an attorney's when it says "he". The example below shows one such biography, with the changed word marked.
- Laya changes its paralegal or attorney answer when only the pronouns change: on 17.85 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.00 of every 100. We are 95% sure the true figure is between 16.15 and 19.55, from 2,000 bios (a clear effect). When its answer changed, it moved toward “paralegal” for the version that read as a woman 100.0 times in 100. See one real biography, both ways.
- Jev changes its paralegal or attorney answer when only the pronouns change: on 3.90 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.40 of every 100. We are 95% sure the true figure is between 3.10 and 4.85, from 2,000 bios (a clear effect). When its answer changed, it moved toward “paralegal” for the version that read as a woman 93.8 times in 100. See one real biography, both ways.
- Laya changes how sure it is of its architect or interior designer answer when a bio says “a wheelchair user”: by 2.84 percentage points. That is already measured against a harmless control edit, text by text. We are 95% sure the true figure is between 2.51 and 3.15, from 1,071 bios (a clear effect). See one real biography, both ways.
Do this instead
Treat the model's answer as one piece of evidence, not the decision. Ask it factual questions about the job as well. Before you rely on the combined decision, run the one-word test on it: the same text asked twice, with one detail changed. Keep a named person accountable for the outcome.
Guidance: test the model on your own texts before you use it