Models
Models on the leaderboard
We test AI models that read a short text and make a quick judgment about a person. Does the model’s answer change when we change a name, a pronoun or another personal detail? Here you can compare the 2 models tested so far and open each one’s results.
1 model is not yet tested.
Models with the strongest bias across comparable tests appear first. A high place is a warning, not an endorsement. How the comparison works.
- 1Layaaverage place 1.12 over the 4 characteristics where it was compared · tested on 8 of 8largest effect: +17.85 percentage points beyond the control edit
Laya is a fast decision model: it answers a yes-or-no question about a text instantly and gives no reasons. Anyone can download it. We ran it two ways: an Apple MLX build (laya-mlx 0.1.0, used wherever we have it) and the original PyTorch build its authors released (laya 0.3.7). Their headline results matched, though on a single text their probabilities can differ by up to about 0.02, so we show one Laya, and each result says which build gave it. The confidence it reports for each answer is its own.
- 2Jevaverage place 1.88 over the 4 characteristics where it was compared · tested on 4 of 8largest effect: +3.77 percentage points beyond the control edit
Jev is a fast decision model: it answers a yes-or-no question about a text instantly and gives no reasons. It runs as an online service, and the confidence it reports for each answer is its own, to two decimal places.
- —KevNot yet tested
Kev is an open-source decision model. It answers questions about text with yes-or-no, choice, or rating answers.