Turn it around · regulated decisions

How AI screening can go wrong

Imagine using AI to rank job applicants, then showing a recruiter only the top names. If the AI judges the same work history differently after a pronoun changes, that difference can decide who gets seen. This page walks through ways to make that problem worse, the evidence behind each, and what to do instead.

“Invert, always invert.”

Charlie Munger, quoting the mathematician Carl Jacobi, in his commencement speech to the Harvard School, June 13, 1986. Transcript.

0.48Laya put 28.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 60.1 of every 100 men. Divide the first rate by the second and you get 0.48, the shortlist ratio, for a shortlist built the first two ways below. A ratio of 1.0 means women and men make the list at the same rate. U.S. hiring guidance treats a ratio under 0.80 as a warning sign.

This is not legal advice. It connects what these models did in our tests to the rules that govern decisions a screening tool might use them for. Whether a real use creates legal liability depends on the facts, the jurisdiction and your lawyers. Each risk below links to the test result behind it, and each rule links to its source.

Recipe 1 #

Let the model's quick answer make the decision

What the organization does

Ask the model one broad question about each applicant, such as "attorney or paralegal?", and act on its answer straight away. The applicant is shortlisted, rejected or sent elsewhere. Nobody looks at anything else, and nobody checks.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: gender, disability.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a), ADA Title I.

What our tests show

The answer changes when nothing but the pronouns change. On the attorney-or-paralegal question, the model can read the same biography as a paralegal's when it says "she" and as an attorney's when it says "he". The example below shows one such biography, with the changed word marked.

Do this instead

Treat the model's answer as one piece of evidence, not the decision. Ask it factual questions about the job as well. Before you rely on the combined decision, run the one-word test on it: the same text asked twice, with one detail changed. Keep a named person accountable for the outcome.

Recipe 2 #

Never count who makes the shortlist

What the organization does

Rank applicants by the model's confidence, its own probability that each one fits the senior job. Hand the top of the list to a recruiter, and never count how many women and men, or people of each race, made the list.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: gender, race.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a).

What our tests show

A small lean in each answer becomes a large gap in who gets shortlisted. In the example shortlist below, women attorneys make the list at a fraction of the rate of men. Swapping only the pronouns shows that the model is the cause, not the biographies. Some women make the list only when the model reads them as men. In this test, no man made it only when read as a woman.

  • Laya put 28.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 60.1 of every 100 men: a shortlist ratio of 0.48 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.40 and 0.55. Counting applicants the model scored equally as a group, not in file order, gives 0.48. Laya got the role right for 72.1 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 82 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
  • Jev put 44.4 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 52.1 of every 100 men: a shortlist ratio of 0.85 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.73 and 0.96. Counting applicants the model scored equally as a group, not in file order, gives 0.85. Jev got the role right for 85.5 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 15 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.
  • Laya put 66.6 of every 100 women attorneys on a shortlist of the top 1,000 of 2,000 bios, against 88.1 of every 100 men: a shortlist ratio of 0.76 (the women's rate divided by the men's; 1.00 is equal). We are 95% sure the true figure is between 0.70 and 0.82. Counting applicants the model scored equally as a group, not in file order, gives 0.76. Laya got the role right for 72.1 of every 100 bios. The test had 419 women and 581 men who really were attorneys. 73 of the 419 women attorneys made the list only when their bio was read as a man's. 0 men made it only when read as a woman's.

Do this instead

Count each group's shortlist rate at the exact cut-off you will use, before you start and again on live applicants. Divide women's rate by men's to get the shortlist ratio, and report its range of likely values. U.S. hiring guidance, the four-fifths rule, treats a ratio under four-fifths as evidence of adverse impact. Treat a ratio under that line, or a range that reaches it, as a reason to stop.

Recipe 3 #

Decide a model is fair because it rarely changes its answer

What the organization does

Run the one-word test, see that the model rarely changes its answer when the pronouns change, and conclude that it is fit to screen people with.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: gender.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a).

What our tests show

Few changed answers can still mean a real gap on the shortlist. The model that rarely changed its answer still put women attorneys on the list at a lower rate than men. On the shorter shortlists, the range of likely ratios reaches down to the four-fifths line.

Do this instead

Measure what the decision actually does: the rate at which each group makes your shortlist, at your cut-off. How often single answers change is not enough.

Recipe 4 #

Sell every employer the same rented model

What the organization does

Build every screening product on the same rented model. Every employer that buys one gets the same answers, for the same reasons.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: gender.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a).

What our tests show

The model's lean is not random: the same biographies lose their place every time. The women attorneys who make the list only when read as men would be turned away the same way at every firm using this model. Yet each firm's audit looks only at its own applicants.

Do this instead

Test your own use of the model on your own applicants. Prefer vendors that publish one-word-test results for each version. Do not assume that a model many firms use has been checked by any of them.

Recipe 5 #

Screen out questions that react to pronouns, and call the model fixed

What the organization does

Let the system add a new question only if its answers stay the same when the pronouns are swapped. Then declare the system fair because every question passed.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: gender.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a).

What our tests show

The screening did cut how often the answer changed. But it let through a question whose answers still line up with gender through what the biography says. The shortlist then got worse for women than if nothing had been done at all.

Do this instead

Judge a change by the shortlist, not by the pronoun test. Accept it only if the shortlist ratio at your real cut-off stays the same or improves.

Recipe 6 #

Let the model change without testing it again

What the organization does

Accept every vendor update, retraining or new input as it arrives, relying on the audit you ran on the old version.

The regulated decision it hits

Hiring and candidate screening; Judging a person's character.

What our tests show

One change passed its own check and still pushed the shortlist ratio well below the original model's. An audit describes the version it tested, and nothing after it.

Do this instead

Fix the model version you use, and treat any change as a new model. Repeat the one-word test and the shortlist count before the new version decides anything.

Recipe 7 #

Always offer the two answers in the same order, and never test the other

What the organization does

Build the tool to list the two possible answers in one order, audit it in that order, and never ask which decisions depend on the order.

The regulated decision it hits

Any screening decision above. Results on these pages: option order.

What our tests show

Swapping the order of the two options changes the model's answer more often than simply asking the same question twice. So the audit describes one arbitrary order, and the decisions that depend on it are arbitrary too.

Do this instead

Ask in both orders and average the answers. Next to every bias result, report how often the order changes the answer. Then set the tool to the order you audited.

Recipe 8 #

Audit with no harmless edit to compare against

What the organization does

Report how often changing a protected detail, such as a name, changes the answer. Never measure how often a harmless change of the same size does.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: race, religion.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a), EU AI Act, Annex III 4(b).

What our tests show

Without that comparison, noise looks like bias and bias hides in noise. One result that looks like a race effect is mostly what swapping one white-sounding name for another already does, and its range of likely values reaches that level. The same arithmetic can also hide a real effect.

Do this instead

Compare every protected edit with an equally harmless edit to the same text. Ask the same question twice, swap a name for another from the same group, or change a hobby.

Recipe 9 #

Treat "no clear effect" as "no bias"

What the organization does

Run a test too small to find the effect, see nothing clear, and file it as proof that the model does not discriminate.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: race, age.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a), ADEA.

What our tests show

No clear effect in a test this size means only that this test could not see one. The age test's range of likely values still includes an effect as large as the one a harmless edit causes. So it neither clears the model nor convicts it.

Do this instead

With every result that shows no clear effect, report the range of likely values and how many texts were tested. Make the test large enough to rule out an effect you would care about.

Recipe 10 #

Test for gender only, and assume the rest behave the same

What the organization does

Audit the model for gender, find a number you can live with, and assume race, disability and religion will behave the same way.

The regulated decision it hits

Hiring and candidate screening. Results on these pages: race, age, disability, religion, nationality, sexuality, veteran status.

Could create legal risk under Title VII, EEOC four-fifths rule, NYC Local Law 144, EU AI Act, Annex III 4(a), ADEA, ADA Title I, EU AI Act, Annex III 4(b).

What our tests show

Each characteristic behaves differently, in each model and each decision. A model that barely moves on one can move on another. A full name, a wheelchair and a religion each moved these models on their own.

Do this instead

Test every characteristic that the rules for your decision name, on your own decision, each against its own harmless control edit.

Recipe 11 #

Ask the model about a candidate's character

What the organization does

Screen people with questions about their character, such as whether they are honest, hardworking or greedy, and let the answers count toward the decision.

The regulated decision it hits

Judging a person's character. Results on these pages: religion, nationality.

Could create legal risk under Title VII, EU AI Act, Annex III 4(a), EU AI Act, Annex III 4(b).

What our tests show

The model's answers move with the religion or nationality a biography names, in the direction documented stereotypes predict.

Do this instead

Do not ask a model about character. Ask about job-related facts the text states, and run the one-word test on even those.