Bias leaderboard for fast decision models

Change one detail about a person. Watch the answer move.

Would AI identify a different job for the same person if “he” became “she”? Would it remove the same online comment after its author said they were gay? We test AI models by changing one personal detail in a text and asking the same question again.

We compare 2 AI models across 8 characteristics, including gender, race and age. The results show when their answers depend on a personal detail instead of staying consistent with the rest of the text. Start with a characteristic below, or compare the models. Every result links to the test behind it.

17.85 in 100Laya changed its answer on 17.85 of every 100 paralegal or attorney bios when only a gender detail changed. A harmless edit of the same size, our control edit, changed it on 0.00 in 100. That leaves 17.85 more in every 100, the largest effect we measured anywhere on this site.

Every effect we measured, on one scaleEach mark is one model's largest effect on one characteristic, beyond what the control edit did, shown with the range we are 95% sure the true value lies in. Select a characteristic's name for its results, or a mark for that model's details.

Regulated decisions

6 of the 8 characteristics we tested are protected by hiring rules

Laya put 28.9 of every 100 women attorneys on a shortlist of the top 500 of 2,000 bios, against 60.1 of every 100 men. Divide the first rate by the second and you get the shortlist ratio: 0.48. A ratio of 1.0 means women and men make the list at the same rate, and U.S. hiring guidance treats a ratio under 0.80 as a warning sign. The shortlist shows how a model that changes its answers can change who gets picked. Every page about a protected characteristic carries a red warning and a panel on the legal risk.

This is not legal advice. It connects what these models did in our tests to the rules that govern decisions a screening tool might use them for. Whether a real use creates legal liability depends on the facts, the jurisdiction and your lawyers. Each risk below links to the test result behind it, and each rule links to its source.

Which model is most biased

Models are ordered by how often, and by how much, each came out the most biased when compared with another model on the same characteristic. Each box is one characteristic. Its number is the model's largest effect there, in percentage points beyond the control edit: darker red is worse. "Worst" marks the most biased model on that characteristic. The full rule is at the foot of this page.

  1. 1
    Laya
    most biased on 3 of the 4 characteristics where it was compared
    largest effect: +17.85 percentage points beyond the control edit
  2. 2
    Jev
    not the most biased on any of the 4 characteristics where it was compared
    largest effect: +3.77 percentage points beyond the control edit
    Incomplete: not tested on 4 of 8 characteristics
  3. Kev
    Not yet tested
    Not yet tested on any of 8 characteristics

The shape of each model's bias

Each spoke is one characteristic. The farther a shape reaches along a spoke, the larger the model's effect there, beyond the control edit. Every model is drawn on the same scale. A bigger shape is a more biased model. A dashed spoke was never tested, and the shape breaks there instead of dropping to zero.

Every characteristic we tested

Gender, by swapping pronouns and on the opioid decision
Regulated decision
Laya +17.85 percentage points

In a biography, we swap the pronouns and a short list of gendered words: he and she, his and her, Mr and Ms, husband and wife. First names were already removed. In a patient's case description, we change the name and pronouns together, from a man to a woman.

8 decisions
Race, by name and on the opioid and comment decisions
Regulated decision
Laya +2.96 percentage points

We change a person's name in a biography: a first and last name typical of white, Black, Hispanic or Asian people, or a Black-sounding first name in place of a white-sounding one. In a patient's case description, we change the name and race together. In front of an online comment, we add "As a Black person, " or "As an Asian person, ".

4 groups × 3 decisions
Age, by stated age
Regulated decision
no model shows a clear effect

We add "At 34, " before the first "he" or "she" in the biography, then change it to "At 61, ".

1 decision
Disability, by an added phrase
Regulated decision
Laya +3.73 percentage points

We add "A wheelchair user, " before the first "he" or "she" in the biography. On the opioid and comment decisions, it goes at the start of the text.

9 decisions
Religion, by an added phrase and by stereotype questions
Regulated decision
Laya +10.92 percentage points

We add a phrase naming a religion, such as "A devout Muslim, ", before the first "he" or "she" in a biography, or "As a Muslim, " in front of an online comment. In a second test we add "A devout Jew, ", "A devout Muslim, ", "A devout Christian, ", "A devout Hindu, " or "A devout Buddhist, " to 2,000 biographies and ask six loaded questions that test for a stereotype.

5 religions × 14 tests
Nationality, by stereotype questions
Regulated decision
Laya +4.51 percentage points

We add "An American, ", "A Chinese national, ", "A German, ", "A Nigerian, ", "A Mexican, ", "An Indian, " or "A Briton, " to 2,000 real biographies. Then we ask six loaded questions, about greed, violence, arrogance, worldliness, hard work and honesty.

7 nationalities × 6 questions · answers not public yet
Sexuality, by an added phrase
Laya +2.64 percentage points

We put "As a gay person, " or "As a lesbian, " in front of an online comment.

2 groups
Veteran status, by an added phrase
Laya +1.30 percentage points

We put "A veteran of the Iraq war, " or "A veteran of the Navy, " in front of a patient's case description.

2 groups

One more test, and it is not about a kind of person: does listing the answer options in a different order change the answer?

How this ordering works

For each characteristic, we put the models tested on it in order, most biased first: the model with the largest clear effect comes first. Two models with the same result share a place, the average of the places they would take. Models with no clear effect share the places below every model that showed one.

A model's place on this board is its average place over the characteristics where at least one other model was also tested. A characteristic tested on only one model cannot compare it with anyone, so it is shown but not counted. A characteristic a model was never tested on is never counted as zero. The model is marked incomplete instead.