Models

Models on the leaderboard

We test AI models that read a short text and make a quick judgment about a person. Does the model’s answer change when we change a name, a pronoun or another personal detail? Here you can compare the 2 models tested so far and open each one’s results.

1 model is not yet tested.

Models with the strongest bias across comparable tests appear first. A high place is a warning, not an endorsement. How the comparison works.

  1. 1
    Laya
    average place 1.12 over the 4 characteristics where it was compared · tested on 8 of 8
    largest effect: +17.85 percentage points beyond the control edit

    Laya is a fast decision model: it answers a yes-or-no question about a text instantly and gives no reasons. Anyone can download it. We ran it two ways: an Apple MLX build (laya-mlx 0.1.0, used wherever we have it) and the original PyTorch build its authors released (laya 0.3.7). Their headline results matched, though on a single text their probabilities can differ by up to about 0.02, so we show one Laya, and each result says which build gave it. The confidence it reports for each answer is its own.

  2. 2
    Jev
    average place 1.88 over the 4 characteristics where it was compared · tested on 4 of 8
    largest effect: +3.77 percentage points beyond the control edit

    Jev is a fast decision model: it answers a yes-or-no question about a text instantly and gives no reasons. It runs as an online service, and the confidence it reports for each answer is its own, to two decimal places.

  3. Kev
    Not yet tested

    Kev is an open-source decision model. It answers questions about text with yes-or-no, choice, or rating answers.