Guidance · regulated decisions

Before you let AI judge people

An AI model can read a biography and help decide who reaches a recruiter. But does its answer change when the same person is described with a different gender, name or religion? Our tests examine that problem. This guide explains what to check before using such a model to make decisions about people, with links to the evidence. Common ways to go wrong shows how these problems can enter a real process.

0.79Asked twice about each bio, once as written and once with the pronouns swapped, then averaged, Laya put 38.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 49.4 of every 100 men. Without averaging, the shortlist ratio (women's shortlist rate divided by men's; 1.0 is equal) was 0.48. U.S. hiring guidance treats a ratio under 0.80 as a warning sign.

This is not legal advice. It connects what these models did in our tests to the rules that govern decisions a screening tool might use them for. Whether a real use creates legal liability depends on the facts, the jurisdiction and your lawyers. Each risk below links to the test result behind it, and each rule links to its source.

Before deploying: a checklist

  1. Name the decision, the rules that govern it, and every characteristic they protect. Why
  2. Run the one-word test on your own texts and question: change one detail, such as a pronoun or a name, and count how often the answer changes. Test one characteristic at a time. Why
  3. Compare each effect with an equally harmless edit, including simply asking twice. Why
  4. Count each group's shortlist rate at your real cut-off, with a range of likely values, and compare the ratio with the four-fifths line. Why
  5. Ask with the two answers in both orders. Use the tool in the order you audited. Why
  6. Accept a fix or a new question only if the shortlist ratio stays the same or improves. Why
  7. Fix the model version, and repeat all of the above on every new one. Why
  8. Show reviewers the people who were cut, not only the shortlist. Why
  9. Publish every "no clear effect" result with its range and the number of texts tested. Why
  10. Keep questions about character out of the decision. Why

Test the model on your own texts before you use it #

Take your own texts and your own question. Change one protected detail, such as a pronoun or a name, and nothing else, then ask again. Count how often the decision changes and by how much. That is the one-word test. Run it for every characteristic the rules name, on every model and decision you will use. A result for one does not carry over to another.

The evidence

Compare every effect with a harmless edit #

Some edits change nothing protected: asking the same question twice, a second name from the same group, or a hobby instead of a religion. They show how much the model moves anyway, for no good reason. Count an effect only when its range of likely values sits clearly above that.

The evidence

Check who makes the shortlist, not only each answer #

Rank applicants the way your tool will rank them, and cut the list where it will cut. Divide each group's shortlist rate by the rate of the group doing best, such as women's rate by men's. Compare that ratio, with its range of likely values, to the four-fifths line: U.S. hiring guidance treats a ratio under four-fifths as evidence of adverse impact. Try several cut-offs, because the ratio depends on where the list is cut.

The evidence

Ask twice with the pronouns swapped, and average the answers #

Make a second copy of each text with only the gendered words swapped: she becomes he, her becomes his, Ms becomes Mr. Ask the model about both copies and average its two answers, so the pronoun cannot tip the result either way. In our shortlist test this moved women's share of the shortlist much closer to men's for both models, and it also made them match the dataset's own job labels a little less often, because in this data the pronoun carries some real information about the job and we removed it. It is not a guarantee. On the nurse and physician bios Laya's averaged ranking went too far and favoured women, which suggests it reads women physicians' bios as more physician-like once the pronouns are neutral. In an experiment where we trained a small decision layer on top of Jev, averaging moved its ratio the wrong way. So check the shortlist again after any fix.

The evidence

Judge every change by the shortlist it produces #

A question can pass the pronoun test, with answers that stay the same when the pronouns are swapped, and still line up with gender through what the text says. Accept a new question, a new input or a retrained model only when the shortlist ratio at your real cut-off stays the same or improves.

The evidence

Offer the two answers in both orders #

Ask each question with the two possible answers in one order, then in the other, and average the results. Next to every bias result, report how many answers change with the order. Use the tool in the order you audited.

The evidence

Fix the version: a new version is a new model #

Record the exact model and version behind every decision. Repeat the one-word test and the shortlist count on each new version before it decides anything. When a new test gives the same numbers as the old one, say so, as we did when we ran Laya two ways, an Apple MLX build and the original PyTorch build, and the headline results matched.

The evidence

Put the reviewer where the harm happens #

A person who reviews only the shortlist does not see the people the model left off it. If someone reviews, they need to see who was cut, and why, not only who was kept.

The evidence

Report "no clear effect" with its range #

"No clear effect" describes one test of one size. Publish the range of likely values and how many texts were tested. Then make the next test large enough to rule out an effect that would matter.

The evidence

Do not ask a model about character #

Questions about character invite the stereotype. Ask about job-related facts the text states, and run the one-word test on those too.

The evidence

The rules cited on this site

Each rule below is cited for a kind of decision we measured. Each links to its official text.

Rules we name, for decisions we have not tested yet

Articles about these results