Bias leaderboard for fast decision models
Change one detail about a person. Watch the answer move.
Would AI identify a different job for the same person if “he” became “she”? Would it remove the same online comment after its author said they were gay? We test AI models by changing one personal detail in a text and asking the same question again.
We compare 3 AI models across 8 characteristics, including gender, race and age. The results show when their answers depend on a personal detail instead of staying consistent with the rest of the text. Start with a characteristic below, or compare the models. Every result links to the test behind it.
17.85 in 100Laya changed its answer on 17.85 of every 100 paralegal or attorney bios when only a gender detail changed. A harmless edit of the same size, our control edit, changed it on 0.00 in 100. That leaves 17.85 more in every 100, the largest effect we measured anywhere on this site.
Which model is most biased
Models are ordered by how often, and by how much, each came out the most biased when compared with another model on the same characteristic. Each box is one characteristic. Its number is the model's largest effect there, in percentage points beyond the control edit: darker red is worse. "Worst" marks the most biased model on that characteristic. The full rule is at the foot of this page.
- 1Layamost biased on 3 of the 5 characteristics where it was comparedlargest effect: +17.85 percentage points beyond the control edit
- 2Kevmost biased on 1 of the 4 characteristics where it was comparedlargest effect: +16.20 percentage points beyond the control editIncomplete: not tested on 4 of 8 characteristics
- 3Jevnot the most biased on any of the 4 characteristics where it was comparedlargest effect: +3.77 percentage points beyond the control editIncomplete: not tested on 4 of 8 characteristics
The shape of each model's bias
Each spoke is one characteristic. The farther a shape reaches along a spoke, the larger the model's effect there, beyond the control edit. Every model is drawn on the same scale. A bigger shape is a more biased model. A dashed spoke was never tested, and the shape breaks there instead of dropping to zero.
Every characteristic we tested
In a biography, we swap the pronouns and a short list of gendered words: he and she, his and her, Mr and Ms, husband and wife. First names were already removed. In a patient's case description, we change the name and pronouns together, from a man to a woman.
We change a person's name in a biography: a first and last name typical of white, Black, Hispanic or Asian people, or a Black-sounding first name in place of a white-sounding one. In a patient's case description, we change the name and race together. In front of an online comment, we add "As a Black person, " or "As an Asian person, ".
We add "At 34, " before the first "he" or "she" in the biography, then change it to "At 61, ".
We add "A wheelchair user, " before the first "he" or "she" in the biography. On the opioid and comment decisions, it goes at the start of the text.
We add a phrase naming a religion, such as "A devout Muslim, ", before the first "he" or "she" in a biography, or "As a Muslim, " in front of an online comment. In a second test we add "A devout Jew, ", "A devout Muslim, ", "A devout Christian, ", "A devout Hindu, " or "A devout Buddhist, " to 2,000 biographies and ask six loaded questions that test for a stereotype.
We add "An American, ", "A Chinese national, ", "A German, ", "A Nigerian, ", "A Mexican, ", "An Indian, " or "A Briton, " to 2,000 real biographies. Then we ask six loaded questions, about greed, violence, arrogance, worldliness, hard work and honesty.
We put "As a gay person, " or "As a lesbian, " in front of an online comment.
We put "A veteran of the Iraq war, " or "A veteran of the Navy, " in front of a patient's case description.
One more test, and it is not about a kind of person: does listing the answer options in a different order change the answer?
How this ordering works
For each characteristic, we put the models tested on it in order, most biased first: the model with the largest clear effect comes first. Two models with the same result share a place, the average of the places they would take. Models with no clear effect share the places below every model that showed one.
A model's place on this board is its average place over the characteristics where at least one other model was also tested. A characteristic tested on only one model cannot compare it with anyone, so it is shown but not counted. A characteristic a model was never tested on is never counted as zero. The model is marked incomplete instead.