Every AI judge on this benchmark sees two anonymous texts and picks the better one. Anonymous for the judge, that is. We know who wrote each text, so we can ask a question the judge cannot: when one of the two texts comes from the judge’s own model, does the verdict change?

It does, for some judges. The table below compares, for each judge, the share of points its own model earns when that judge decides, with the share the same model earns under all the other judges. A win counts one point, a tie half a point. Data: every published duel as of 19 September 2026, about 63,000 verdicts per language.

Sol and Opus 5 favour themselves, Fable does not

Judge (own model) English French German Spanish
GPT-5.6 Sol 87.9 % vs 69.0 % 90.1 % vs 74.6 % 93.4 % vs 78.3 % 92.8 % vs 72.1 %
Claude Opus 5 94.0 % vs 82.6 % 99.6 % vs 96.4 % 98.0 % vs 90.5 % 96.9 % vs 92.0 %
Muse Spark 1.3 82.6 % vs 73.5 % 88.2 % vs 84.4 % 84.1 % vs 79.0 % 79.3 % vs 80.4 %
Claude Fable 5 88.1 % vs 89.8 % 90.7 % vs 96.0 % 91.3 % vs 92.8 % 91.0 % vs 93.8 %
Claude Fable 5.1 84.4 % vs 87.8 % 90.2 % vs 95.6 % 90.7 % vs 93.0 % 90.6 % vs 96.5 %

Read each cell as “under its own judge vs under the other judges”. GPT-5.6 Sol is the clearest case: its own texts collect between 15 and 21 points more when Sol is the judge, in every language, over roughly 1,000 self-judged duels per language. Claude Opus 5 shows the same direction with a smaller gap, 3 to 11 points. Muse Spark 1.3 leans the same way in three languages out of four, on a small sample of 135 to 172 self-judged duels.

Claude Fable 5 and Claude Fable 5.1 go the other way. Under their own verdicts their texts earn 2 to 6 points less than under the other judges, in all four languages. Whatever the reason, they are not flattering themselves.

What it does to the rankings

Each judge also publishes its own table, ranked by wins. In that table GPT-5.6 Sol places its own model second in all four languages. In the main ranking, which pools every judge, the same model sits ninth in English, tenth in French, sixth in German and ninth in Spanish. Claude Opus 5 is first or second everywhere, in its own table as in the main one, so its bonus changes little at the top but still inflates its score.

The pooled ranking is where the bonus matters. Claude Opus 5 and GPT-5.6 Sol together supply about 85% of all verdicts, so their self-judged duels are not a footnote: around 800 Opus-on-Opus and 1,000 Sol-on-Sol duels per language enter the Bradley–Terry fit with everyone else’s.

Why we do not simply call it bias

Two caveats keep this from being a verdict. First, the duels are not the same. A judge’s own model meets a different mix of opponents under that judge than under the others, and the mix changes the expected share. The gaps for Sol are large enough that opponent mix is unlikely to explain all of it, but it can explain part. Second, “the other judges” is not a neutral instrument: it is mostly Opus 5 when we look at Sol, and mostly Sol when we look at Opus 5. If one of them is severe with the other, that severity shows up on the same row.

The clean test is a controlled sample: the same pairs judged by every judge. We have started building it.

What changes on the site

Three things. Each judge’s table already lets you read its verdicts separately, and now you know which ones to read with a raised eyebrow. We will publish a variant of the main ranking that excludes every duel in which the judge and one of the candidates share a maker, and compare it with the current one. And when a new judge is added, its self-preference will be measured before its verdicts are pooled.