Guidance · regulated decisions
Before you let AI judge people
An AI model can read a biography and help decide who reaches a recruiter. But does its answer change when the same person is described with a different gender, name or religion? Our tests examine that problem. This guide explains what to check before using such a model to make decisions about people, with links to the evidence. Common ways to go wrong shows how these problems can enter a real process.
0.79Asked twice about each bio, once as written and once with the pronouns swapped, then averaged, Laya put 38.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 49.4 of every 100 men. Without averaging, the shortlist ratio (women's shortlist rate divided by men's; 1.0 is equal) was 0.48. U.S. hiring guidance treats a ratio under 0.80 as a warning sign.
This is not legal advice. It connects what these models did in our tests to the rules that govern decisions a screening tool might use them for. Whether a real use creates legal liability depends on the facts, the jurisdiction and your lawyers. Each risk below links to the test result behind it, and each rule links to its source.
Before deploying: a checklist
- Name the decision, the rules that govern it, and every characteristic they protect. Why
- Run the one-word test on your own texts and question: change one detail, such as a pronoun or a name, and count how often the answer changes. Test one characteristic at a time. Why
- Compare each effect with an equally harmless edit, including simply asking twice. Why
- Count each group's shortlist rate at your real cut-off, with a range of likely values, and compare the ratio with the four-fifths line. Why
- Ask with the two answers in both orders. Use the tool in the order you audited. Why
- Accept a fix or a new question only if the shortlist ratio stays the same or improves. Why
- Fix the model version, and repeat all of the above on every new one. Why
- Show reviewers the people who were cut, not only the shortlist. Why
- Publish every "no clear effect" result with its range and the number of texts tested. Why
- Keep questions about character out of the decision. Why
Test the model on your own texts before you use it #
Take your own texts and your own question. Change one protected detail, such as a pronoun or a name, and nothing else, then ask again. Count how often the decision changes and by how much. That is the one-word test. Run it for every characteristic the rules name, on every model and decision you will use. A result for one does not carry over to another.
The evidence
- Laya changes its paralegal or attorney answer when only the pronouns change: on 17.85 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.00 of every 100. We are 95% sure the true figure is between 16.15 and 19.55, from 2,000 bios (a clear effect). When its answer changed, it moved toward “paralegal” for the version that read as a woman 100.0 times in 100. See one real biography, both ways.
- Jev changes how sure it is of its surgeon or physician answer when a white-sounding full name becomes a Black-sounding one: by 0.35 percentage points. By comparison, after a harmless control edit of the same size, its confidence moves 0.06 percentage points. We are 95% sure the true figure is between 0.19 and 0.50, from 500 bios (a clear effect). See one real biography, both ways.
- Laya changes how sure it is of its architect or interior designer answer when a bio says “a wheelchair user”: by 2.84 percentage points. That is already measured against a harmless control edit, text by text. We are 95% sure the true figure is between 2.51 and 3.15, from 1,071 bios (a clear effect). See one real biography, both ways.
How the one-word test works · How to fail: let the model's quick answer make the decision · How to fail: sell every employer the same rented model · How to fail: test for gender only, and assume the rest behave the same
Compare every effect with a harmless edit #
Some edits change nothing protected: asking the same question twice, a second name from the same group, or a hobby instead of a religion. They show how much the model moves anyway, for no good reason. Count an effect only when its range of likely values sits clearly above that.
The evidence
- Laya changes its surgeon or physician answer when only a white-sounding first name becomes a Black-sounding one: on 3.18 of every 100 bios. By comparison, after a harmless control edit of the same size, it changes its answer on 2.36 of every 100. We are 95% sure the true figure is between 2.36 and 4.14, from 1,571 bios (no clear effect). See one real biography, both ways.
- Jev changes its journalist or professor answer when only the pronouns change: on 0.80 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.60 of every 100 (measured on another decision). We are 95% sure the true figure is between 0.40 and 1.20, from 2,000 bios (no clear effect). When its answer changed, it moved toward “journalist” for the version that read as a woman 80.0 times in 100. See one real biography, both ways.
How we rank the results · How to fail: audit with no harmless edit to compare against
Check who makes the shortlist, not only each answer #
Rank applicants the way your tool will rank them, and cut the list where it will cut. Divide each group's shortlist rate by the rate of the group doing best, such as women's rate by men's. Compare that ratio, with its range of likely values, to the four-fifths line: U.S. hiring guidance treats a ratio under four-fifths as evidence of adverse impact. Try several cut-offs, because the ratio depends on where the list is cut.
The evidence
- Laya put 28.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 60.1 of every 100 men: a shortlist ratio of 0.48 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.40 and 0.55. Counting applicants the model scored equally as a group, not in file order, gives 0.48. Laya got the role right for 72.1 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 82 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
- Jev put 44.4 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 52.1 of every 100 men: a shortlist ratio of 0.85 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.73 and 0.96. Counting applicants the model scored equally as a group, not in file order, gives 0.85. Jev got the role right for 85.5 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 15 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
- Jev put 22.2 of every 100 women attorneys on a shortlist of the top 250 of 2,000 bios, against 25.8 of every 100 men: a shortlist ratio of 0.86 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.66 and 1.08. Counting applicants the model scored equally as a group, not in file order, gives 0.85. Jev got the role right for 85.5 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 5 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
- Laya put 66.6 of every 100 women attorneys on a shortlist of the top 1,000 of 2,000 bios, against 88.1 of every 100 men: a shortlist ratio of 0.76 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.70 and 0.82. Counting applicants the model scored equally as a group, not in file order, gives 0.76. Laya got the role right for 72.1 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 73 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
How to fail: never count who makes the shortlist · How to fail: decide a model is fair because it rarely changes its answer
Ask twice with the pronouns swapped, and average the answers #
Make a second copy of each text with only the gendered words swapped: she becomes he, her becomes his, Ms becomes Mr. Ask the model about both copies and average its two answers, so the pronoun cannot tip the result either way. In our shortlist test this moved women's share of the shortlist much closer to men's for both models, and it also made them match the dataset's own job labels a little less often, because in this data the pronoun carries some real information about the job and we removed it. It is not a guarantee. On the nurse and physician bios Laya's averaged ranking went too far and favoured women, which suggests it reads women physicians' bios as more physician-like once the pronouns are neutral. In an experiment where we trained a small decision layer on top of Jev, averaging moved its ratio the wrong way. So check the shortlist again after any fix.
The evidence
- Asked twice about each bio, once as written and once with the pronouns swapped, then averaged, Laya put 38.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 49.4 of every 100 men: a shortlist ratio of 0.79 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.69 and 0.91. Counting applicants the model scored equally as a group, not in file order, gives 0.79. Laya got the role right for 66.5 of every 100 bios. The test had 419 women and 581 men who really were attorneys.
- Asked twice about each bio, once as written and once with the pronouns swapped, then averaged, Jev put 47.3 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 50.6 of every 100 men: a shortlist ratio of 0.93 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.82 and 1.05. Counting applicants the model scored equally as a group, not in file order, gives 0.91. Jev got the role right for 84.2 of every 100 bios. The test had 419 women and 581 men who really were attorneys.
- Asked twice about each bio, once as written and once with the pronouns swapped, then averaged, Laya put 54.5 of every 100 women physicians on a shortlist of the top 500 of 2,000 bios, against 44.1 of every 100 men: a shortlist ratio of 1.24 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 1.10 and 1.42. Counting applicants the model scored equally as a group, not in file order, gives 1.25. Laya got the role right for 79.7 of every 100 bios. The test had 501 women and 499 men who really were physicians.
- The shortlist ratio when we trained a small decision layer on top of Jev and averaged each text with its pronoun-swapped copy: 0.7169. Reported in Can You Fix It?.
Judge every change by the shortlist it produces #
A question can pass the pronoun test, with answers that stay the same when the pronouns are swapped, and still line up with gender through what the text says. Accept a new question, a new input or a retrained model only when the shortlist ratio at your real cut-off stays the same or improves.
The evidence
- How often the answer changed when only the pronouns were swapped, after we trained a small decision layer on top of Jev and let it add only questions whose answers held steady when the pronouns were swapped, compared with Jev alone: promoted seeds 1.70%, 2.15% (mean 1.93%, almost exactly half of 3.90%); unpromoted seed 3 2.95%. Reported in Can You Fix It?.
- How the shortlist ratio at the top 500 changed when we trained a small decision layer on top of Jev and let it add only questions whose answers held steady when the pronouns were swapped, compared with Jev alone: promoted seeds 0.6562, 0.6488 -- both below J0; unpromoted seed 3 0.8512 (=J0). Reported in Can You Fix It?.
How to fail: screen out questions that react to pronouns, and call the model fixed
Offer the two answers in both orders #
Ask each question with the two possible answers in one order, then in the other, and average the results. Next to every bias result, report how many answers change with the order. Use the tool in the order you audited.
The evidence
- Laya changes its teacher or professor answer when the two answer options swap places: on 7.80 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.00 of every 100. We are 95% sure the true figure is between 6.70 and 9.06, from 2,000 bios (a clear effect). See one real biography, both ways.
- Jev changes its teacher or professor answer when the two answer options swap places: on 2.80 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.60 of every 100. We are 95% sure the true figure is between 1.68 and 4.64, from 500 bios (a clear effect). See one real biography, both ways.
- Asked the same teacher or professor question twice about the same unchanged bio, Laya changes its answer: on 0.00 of every 100 bios, out of 500 tested. That is how much it moves for no reason at all.
Results for the order of the answers · How to fail: always offer the two answers in the same order, and never test the other
Fix the version: a new version is a new model #
Record the exact model and version behind every decision. Repeat the one-word test and the shortlist count on each new version before it decides anything. When a new test gives the same numbers as the old one, say so, as we did when we ran Laya two ways, an Apple MLX build and the original PyTorch build, and the headline results matched.
The evidence
- Laya changes its paralegal or attorney answer when only the pronouns change: on 17.85 of every 100 bios. By comparison, when we simply ask again about the same unchanged text, it changes its answer on 0.00 of every 100. We are 95% sure the true figure is between 16.15 and 19.55, from 2,000 bios (a clear effect). When its answer changed, it moved toward “paralegal” for the version that read as a woman 100.0 times in 100. See one real biography, both ways.
- How the shortlist ratio at the top 500 changed when we trained a small decision layer on top of Jev and let it add only questions whose answers held steady when the pronouns were swapped, compared with Jev alone: promoted seeds 0.6562, 0.6488 -- both below J0; unpromoted seed 3 0.8512 (=J0). Reported in Can You Fix It?.
Put the reviewer where the harm happens #
A person who reviews only the shortlist does not see the people the model left off it. If someone reviews, they need to see who was cut, and why, not only who was kept.
The evidence
- Laya put 28.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 60.1 of every 100 men: a shortlist ratio of 0.48 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.40 and 0.55. Counting applicants the model scored equally as a group, not in file order, gives 0.48. Laya got the role right for 72.1 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 82 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
Report "no clear effect" with its range #
"No clear effect" describes one test of one size. Publish the range of likely values and how many texts were tested. Then make the next test large enough to rule out an effect that would matter.
The evidence
- Jev changes its surgeon or physician answer when a bio gives the age as 61, not 34: on 1.30 of every 100 bios. By comparison, after a harmless control edit of the same size, it changes its answer on 0.89 of every 100. We are 95% sure the true figure is between 0.73 and 1.95, from 1,231 bios (no clear effect). See one real biography, both ways.
- Laya changes its surgeon or physician answer when a bio gives the age as 61, not 34: on 0.97 of every 100 bios. By comparison, after a harmless control edit of the same size, it changes its answer on 0.49 of every 100. We are 95% sure the true figure is between 0.49 and 1.54, from 1,231 bios (no clear effect). See one real biography, both ways.
- Jev changes its surgeon or physician answer when only a white-sounding first name becomes a Black-sounding one: on 0.89 of every 100 bios. By comparison, after a harmless control edit of the same size, it changes its answer on 0.57 of every 100. We are 95% sure the true figure is between 0.51 and 1.40, from 1,571 bios (no clear effect). See one real biography, both ways.
Do not ask a model about character #
Questions about character invite the stereotype. Ask about job-related facts the text states, and run the one-word test on those too.
The evidence
- Laya leans toward calling a person “honest” when a bio says “Christian”: by 10.92 percentage points. That is compared with the other groups, so a change every group shares is left out. We are 95% sure the true figure is between 10.69 and 11.15, from 2,000 bios (a clear effect).
- Laya leans toward calling a person “greedy” when a bio says “Jewish”: by 0.74 percentage points. That is compared with the other groups, so a change every group shares is left out. We are 95% sure the true figure is between 0.58 and 0.91, from 2,000 bios (a clear effect).
The rules cited on this site
Each rule below is cited for a kind of decision we measured. Each links to its official text.
- Title VII of the Civil Rights Act of 1964: race, color, religion, sex and national origin in employment.
Read the text
“to fail or refuse to hire or to discharge any individual, or otherwise to discriminate against any individual with respect to his compensation, terms, conditions, or privileges of employment, because of such individual's race, color, religion, sex, or national origin”
“to limit, segregate, or classify his employees or applicants for employment in any way which would deprive or tend to deprive any individual of employment opportunities or otherwise adversely affect his status as an employee, because of such individual's race, color, religion, sex, or national origin”
- The four-fifths rule of the Uniform Guidelines on Employee Selection Procedures (29 CFR 1607.4(D)): selection rates by race, sex or ethnic group.
Read the text
“A selection rate for any race, sex, or ethnic group which is less than four-fifths ( 4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact, while a greater than four-fifths rate will generally not be regarded by Federal enforcement agencies as evidence of adverse impact.”
- Age Discrimination in Employment Act of 1967: age, for people aged forty and over, in employment.
Read the text
“to fail or refuse to hire or to discharge any individual or otherwise discriminate against any individual with respect to his compensation, terms, conditions, or privileges of employment, because of such individual's age”
“individuals who are at least 40 years of age”
- Americans with Disabilities Act of 1990, Title I: disability in employment.
Read the text
“No covered entity shall discriminate against a qualified individual on the basis of disability in regard to job application procedures, the hiring, advancement, or discharge of employees, employee compensation, job training, and other terms, conditions, and privileges of employment.”
“Using qualification standards, employment tests or other selection criteria that screen out or tend to screen out an individual with a disability or a class of individuals with disabilities unless the standard, test or other selection criteria, as used by the covered entity, is shown to be job-related for the position in question and is consistent with business necessity”
- New York City Local Law 144 (automated employment decision tools): automated tools that screen candidates or employees in New York City, and the bias audit they need, which reports results by sex and by race or ethnicity.
Read the text
“to screen candidates for employment or employees for promotion within the city”
“the tool has been subject to a bias audit within one year of the use of the tool”
- Regulation (EU) 2024/1689 (the AI Act), Annex III, point 4(a): high-risk AI systems for recruitment and selection.
Read the text
“AI systems intended to be used for the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications, and to evaluate candidates”
- Regulation (EU) 2024/1689 (the AI Act), Annex III, point 4(b): high-risk AI systems for promotion, termination and evaluating workers.
Read the text
“AI systems intended to be used to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships, to allocate tasks based on individual behaviour or personal traits or characteristics or to monitor and evaluate the performance and behaviour of persons in such relationships”
Rules we name, for decisions we have not tested yet
- Hiring, where a bio mentions military service (USERRA): planned, not yet measured (veteran status). Our notes.
- Approving a line of credit (ECOA and Regulation B): planned, not yet measured (regulated screening decisions). Our notes.
- Approving an application to rent (Fair Housing Act): planned, not yet measured (regulated screening decisions). Our notes.
- Prioritising a medical appointment (ACA Section 1557): planned, not yet measured (regulated screening decisions). Our notes.
- Escalating a consumer complaint (UDAAP): surveyed, not yet measured. Our notes.
- Scoring performance reviews or promotion readiness (Title VII, NYC Local Law 144 and EU AI Act 4(b), for evaluations): planned, not yet measured (gendered language in evaluations). Our notes.