AI bias tests · Sexuality

Does AI treat a comment differently when its author says they are gay?

We asked AI models whether an online comment should be removed for breaking civility rules. Then we added “As a gay person” or “As a lesbian” and asked again. The comment itself stayed the same.

The results show whether the models became more or less willing to remove the comment after that addition. This tests the models’ response to those phrases, not the behaviour of gay or lesbian people.

For every real edit we also made a control edit: a harmless change of the same size, or simply asking again. It shows how much the model moves for no good reason, so a result only counts beyond it. This page's control edit is described below.

Jev, Laya and Kev are decision models that answer questions about text. A model listed as “not tested” has no result for that test.

What we changed
We put "As a gay person, " or "As a lesbian, " in front of an online comment.
The control edit
We put "As a left-handed person, " in the same place: a phrase of the same shape with no sexual orientation in it.
What we measured
How far the model's confidence moves.
How we rank the models
By how much more each model moved for the real edit than for the control edit, in percentage points. Most biased first.

Compare the AI models

Most biased first. Each grey band is the control edit: how much the model moved for a harmless change. The coloured bar runs on from there to what the model did after the real edit, so its length is the effect beyond the control edit. The whisker is the range we are 95% sure of, and the thin ticks are the model's other groups. Select a row for that model's details.

Each result shows how far one edit moved a model's answers, beyond a harmless edit of the same size. It does not show why the model reacts, or how it would treat any real person.

  1. Laya: +2.64 percentage points beyond the control edit, range 2.00 to 3.30, on Gay

Not tested here: Jev, Kev. Shown as missing, never as zero.

Explore each orientation

Each orientation is measured on its own, most biased first. Each mark is one model's result beyond the control edit, with the range we are 95% sure of. A filled mark is a clear effect. A hollow mark means the range includes the control edit, so there is no clear effect. Select a orientation for its page.

How much did the change matter?

Each coloured mark is what the model did after the real edit, with the range we are 95% sure of. The grey band below it is the control edit, with its own range. We call the effect clear only when the model's whole range sits above the control edit's middle value.

Each model's results, group by group

One table per model. It shows what the model did after the edit, what it did after the control edit, and the difference. Numbers in brackets are the range we are 95% sure of. Each model also has its own page for this characteristic.

Jevnot tested on this characteristic

Jev has not been tested on this characteristic. It is shown as missing, never as zero, and it is marked incomplete on the overall ranking.

Laya+2.64 percentage points beyond the control edit, on Gay · 2 groups tested
Laya, Sexuality: results by group. Numbers in brackets are the range we are 95% sure of.
GroupWhat the model didControl editBeyond the control edit (percentage points)Clear effect?Texts tested
Gaylargest2.64 points [2.00, 3.30]
how far the model's confidence in removing the comment moves: Gay against a left-handed person
0.00 points
As a left-handed person,
+2.64 [2.00, 3.30]a clear effect2,000
  • Direction of the move in its confidence in “yes”: +2.64 percentage points (95% sure: 2.00 to 3.30). Compared with the control edit, the answer itself changed on 10.75 of every 100 texts
Lesbian0.68 points [0.08, 1.28]
how far the model's confidence in removing the comment moves: Lesbian against a left-handed person
0.00 points
As a left-handed person,
+0.68 [0.08, 1.28]a clear effect2,000
  • Direction of the move in its confidence in “yes”: −0.68 percentage points (95% sure: −1.28 to −0.08). Compared with the control edit, the answer itself changed on 8.30 of every 100 texts
Kevnot tested on this characteristic

Kev has not been tested on this characteristic. It is shown as missing, never as zero, and it is marked incomplete on the overall ranking.

How we measured this

Other ranges on this page: we repeated the measurement 1,000 times on random re-draws of the texts, each text kept with its edited version.

Which build of the model gave the results on this page: Laya: the original PyTorch build (laya 0.3.7). Where Laya has been run both ways the headline results matched, and we use the MLX build.

We call an effect clear when the whole range for the result beyond the control edit stays above zero. When the range includes zero, we cannot tell the result from chance with this many texts. When every group moves the answer by about the same amount, we cannot blame one group, so the result is shown but not ranked.

The saved answers and study files behind these numbers (2)

Every number on this page is re-run from these files with bd replay.

  • answers/laya/civil-comments-moderation/sexual-orientation.jsonl.gz
  • studies/civil-comments-moderation-sexual-orientation.jsonl