How much do the AI judges agree with each other?
Claude Opus 5 and GPT-5.6 Sol, who supply most verdicts, point the same way on about nine model pairs out of ten. The disagreements are elsewhere: Sol almost never calls a tie, Claude Fable 5.1 does so up to one time in six, and Claude Opus 4.8 sides with the others only two times in three.
The ranking pools verdicts from five AI judges. That only makes sense if they mostly agree. Do they?
We measured it the way a reader would: for every pair of models that two judges have both seen at least four times, does each judge give the same model the majority of points? Pairs that end exactly level under either judge are left out. Data as of 19 September 2026.
The two big judges agree nine times out of ten
Claude Opus 5 and GPT-5.6 Sol have judged about 2,750 model pairs in common per language. They side with the same model in 88.5% of them in English, 91.6% in French, 91.6% in German and 91.1% in Spanish.
Two judges built by different companies, on different data, reaching the same verdict on nine pairs out of ten, is the strongest evidence this benchmark has that its ranking measures something real about writing rather than the taste of one model. The tenth pair is where the interesting questions live, and we will publish those disagreements as a list.
The smaller judges
| Judge pair | English | French | German | Spanish |
|---|---|---|---|---|
| Claude Fable 5.1 × Claude Opus 5 | 91.3 % (69) | 95.8 % (71) | 92.4 % (79) | 88.6 % (70) |
| Claude Fable 5.1 × GPT-5.6 Sol | 76.1 % (71) | 82.2 % (73) | 85.4 % (82) | 76.4 % (72) |
| Claude Fable 5 × Claude Opus 5 | 75.5 % (49) | 91.8 % (49) | 91.5 % (47) | 95.0 % (40) |
| Claude Fable 5 × GPT-5.6 Sol | 86.0 % (50) | 90.2 % (51) | 91.3 % (46) | 94.9 % (39) |
| Claude Opus 4.8 × Claude Opus 5 | 72.4 % (87) | 66.7 % (24) | 72.1 % (68) | 81.1 % (95) |
| Claude Opus 4.8 × GPT-5.6 Sol | 66.3 % (83) | 71.4 % (21) | 71.0 % (69) | 75.0 % (96) |
| Muse Spark 1.3 × Claude Opus 5 | 93.5 % (108) | 100 % (33) | 90.8 % (65) | — |
| Muse Spark 1.3 × GPT-5.6 Sol | 75.7 % (136) | 92.1 % (38) | 94.9 % (79) | — |
The number in brackets is the count of shared pairs; below a hundred, read the percentages as indications. Two things stand out. Claude Opus 4.8, the previous generation of Opus, agrees with the current judges only two times in three: its verdicts look like an earlier taste, which is why it is now a retired judge with a small share of the pool. And Claude Fable 5.1 agrees more with Claude Opus 5 than with GPT-5.6 Sol, a family resemblance that a good benchmark should keep an eye on.
Not every judge means the same thing by “better”
Judges also differ in how often they refuse to choose:
| Judge | Verdicts per language | Tie rate (EN / FR / DE / ES) |
|---|---|---|
| GPT-5.6 Sol | ≈ 25,000 | 0.5 % / 0.1 % / 0.1 % / 0.2 % |
| Claude Opus 5 | ≈ 28,000 | 5.4 % / 2.4 % / 3.3 % / 3.3 % |
| Claude Opus 4.8 | 400 – 1,100 | 3.7 % / 0.5 % / 1.5 % / 5.4 % |
| Muse Spark 1.3 | ≈ 6,000 | 9.7 % / 6.9 % / 9.6 % / 9.7 % |
| Claude Fable 5 | ≈ 1,000 | 10.2 % / 5.4 % / 4.0 % / 7.7 % |
| Claude Fable 5.1 | ≈ 1,800 | 15.6 % / 6.6 % / 12.0 % / 9.5 % |
GPT-5.6 Sol almost always picks a winner. Claude Fable 5.1 calls a tie for one English duel in six. Neither is wrong: a tie is a legitimate answer when two texts are equally good, and English duels between top models often are. But it means a “win” under Sol is a weaker statement than a “win” under Fable 5.1. This is one reason the main ranking uses a Bradley–Terry fit over verdicts, where a tie counts half a win for each side, rather than averaging the judges’ 0–100 scores.
What we will do with it
Publish the list of pairs where Opus 5 and Sol disagree, with the texts, so readers can be the third judge. Keep Claude Opus 4.8 out of new verdicts. And weight future judges by how well they agree with the pool before their verdicts change the ranking, not after.