The ranking pools verdicts from five AI judges. That only makes sense if they mostly agree. Do they?

We measured it the way a reader would: for every pair of models that two judges have both seen at least four times, does each judge give the same model the majority of points? Pairs that end exactly level under either judge are left out. Data as of 19 September 2026.

The two big judges agree nine times out of ten

Claude Opus 5 and GPT-5.6 Sol have judged about 2,750 model pairs in common per language. They side with the same model in 88.5% of them in English, 91.6% in French, 91.6% in German and 91.1% in Spanish.

Two judges built by different companies, on different data, reaching the same verdict on nine pairs out of ten, is the strongest evidence this benchmark has that its ranking measures something real about writing rather than the taste of one model. The tenth pair is where the interesting questions live, and we will publish those disagreements as a list.

The smaller judges

Judge pair English French German Spanish
Claude Fable 5.1 × Claude Opus 5 91.3 % (69) 95.8 % (71) 92.4 % (79) 88.6 % (70)
Claude Fable 5.1 × GPT-5.6 Sol 76.1 % (71) 82.2 % (73) 85.4 % (82) 76.4 % (72)
Claude Fable 5 × Claude Opus 5 75.5 % (49) 91.8 % (49) 91.5 % (47) 95.0 % (40)
Claude Fable 5 × GPT-5.6 Sol 86.0 % (50) 90.2 % (51) 91.3 % (46) 94.9 % (39)
Claude Opus 4.8 × Claude Opus 5 72.4 % (87) 66.7 % (24) 72.1 % (68) 81.1 % (95)
Claude Opus 4.8 × GPT-5.6 Sol 66.3 % (83) 71.4 % (21) 71.0 % (69) 75.0 % (96)
Muse Spark 1.3 × Claude Opus 5 93.5 % (108) 100 % (33) 90.8 % (65)
Muse Spark 1.3 × GPT-5.6 Sol 75.7 % (136) 92.1 % (38) 94.9 % (79)

The number in brackets is the count of shared pairs; below a hundred, read the percentages as indications. Two things stand out. Claude Opus 4.8, the previous generation of Opus, agrees with the current judges only two times in three: its verdicts look like an earlier taste, which is why it is now a retired judge with a small share of the pool. And Claude Fable 5.1 agrees more with Claude Opus 5 than with GPT-5.6 Sol, a family resemblance that a good benchmark should keep an eye on.

Not every judge means the same thing by “better”

Judges also differ in how often they refuse to choose:

Judge Verdicts per language Tie rate (EN / FR / DE / ES)
GPT-5.6 Sol ≈ 25,000 0.5 % / 0.1 % / 0.1 % / 0.2 %
Claude Opus 5 ≈ 28,000 5.4 % / 2.4 % / 3.3 % / 3.3 %
Claude Opus 4.8 400 – 1,100 3.7 % / 0.5 % / 1.5 % / 5.4 %
Muse Spark 1.3 ≈ 6,000 9.7 % / 6.9 % / 9.6 % / 9.7 %
Claude Fable 5 ≈ 1,000 10.2 % / 5.4 % / 4.0 % / 7.7 %
Claude Fable 5.1 ≈ 1,800 15.6 % / 6.6 % / 12.0 % / 9.5 %

GPT-5.6 Sol almost always picks a winner. Claude Fable 5.1 calls a tie for one English duel in six. Neither is wrong: a tie is a legitimate answer when two texts are equally good, and English duels between top models often are. But it means a “win” under Sol is a weaker statement than a “win” under Fable 5.1. This is one reason the main ranking uses a Bradley–Terry fit over verdicts, where a tie counts half a win for each side, rather than averaging the judges’ 0–100 scores.

What we will do with it

Publish the list of pairs where Opus 5 and Sol disagree, with the texts, so readers can be the third judge. Keep Claude Opus 4.8 out of new verdicts. And weight future judges by how well they agree with the pool before their verdicts change the ranking, not after.