Can the order of two answers change an AI’s choice?
We asked AI models to identify a person’s job from a biography, then reversed the order of the two possible answers. “Surgeon or physician?” became “Physician or surgeon?” The biography stayed the same. We counted how often the models chose a different job.
This tests whether an AI’s choice depends on how a question is presented. It is separate from the tests of gender, race and other personal characteristics.
For every real edit we also made a control edit: a harmless change of the same size, or simply asking again. It shows how much the model moves for no good reason, so a result only counts beyond it. This page's control edit is described below.
Jev, Laya and Kev are decision models that answer questions about text. A model listed as “not tested” has no result for that test.
What we changed
We ask the same question with its two answers listed the other way round, such as "physician or surgeon?" instead of "surgeon or physician?".
The control edit
We ask about the same biography a second time, unchanged.
What we measured
How often the answer changes.
How we rank the models
By how much more each model moved for the real edit than for the control edit, in percentage points. Most biased first.
This is not a personal characteristic. It shows how much the order of the answers alone moves the model. Every other number on this site uses one fixed order for each question.
Compare the AI models
Most biased first. Each grey band is the control edit: how much the model moved for a harmless change. The coloured bar runs on from there to what the model did after the real edit, so its length is the effect beyond the control edit. The whisker is the range we are 95% sure of, and the thin ticks are the model's other decisions. Select a row for that model's details.
Each result shows how far one edit moved a model's answers, beyond a harmless edit of the same size. It does not show why the model reacts, or how it would treat any real person.
Laya: +7.80 percentage points beyond the control edit, range 6.70 to 9.06, on teacher or professor
Jev: +2.20 percentage points beyond the control edit, range 1.08 to 4.04, on teacher or professor
Not tested here: Kev. Shown as missing, never as zero.
Explore each decision
Each decision is measured on its own, most biased first. Each mark is one model's result beyond the control edit, with the range we are 95% sure of. A filled mark is a clear effect. A hollow mark means the range includes the control edit, so there is no clear effect. Select a decision for its page.
The pattern across decisions
One spoke per decision. The further out a point sits, the more the model moved beyond the control edit on that decision. A hollow point is no clear effect. A gap in a shape means we did not test that model there.
Each model's results, decision by decision
One table per model. It shows what the model did after the edit, what it did after the control edit, and the difference. Numbers in brackets are the range we are 95% sure of. Each model also has its own page for this characteristic.
Jev+2.20 percentage points beyond the control edit, on teacher or professor · 4 decisions tested
Jev, Option order: results by decision. Numbers in brackets are the range we are 95% sure of.
Kev has not been tested on this characteristic. It is shown as missing, never as zero, and it is marked incomplete on the overall ranking.
How we measured this
Other ranges on this page: the standard range for a share of texts (the Wilson interval), because this result was not re-drawn 1,000 times.
Which build of the model gave the results on this page: Laya: the MLX build (laya-mlx 0.1.0). Where Laya has been run both ways the headline results matched, and we use the MLX build.
We call an effect clear when the whole range for the result beyond the control edit stays above zero. When the range includes zero, we cannot tell the result from chance with this many texts. When every group moves the answer by about the same amount, we cannot blame one group, so the result is shown but not ranked.
The saved answers and study files behind these numbers (28)
Every number on this page is re-run from these files with bd replay.