AI judge results
See each AI judge’s detailed ranking and where their results differ.
Models are ranked by the quality of their English writing. Use the filters to narrow the list. For the selected writing categories, ranks are calculated before publisher, size, access, reasoning and quantization filters; hidden models can therefore leave gaps.
Last updated:
| Rank | Model | AI judge score | Human score |
|---|---|---|---|
| 1 | 93.8 | — | |
| 2 | 93.7 | — | |
| 4 | 92.2 | — | |
| 5 | 91.4 | — | |
| 6 | 90.3 | — | |
| 8 | 86.5 | — | |
| 9 | 85.5 | — | |
| 10 | 83.1 | — | |
| 11 | 81.8 | — | |
| 12 | 80.8 | — | |
| 14 | 79.7 | — | |
| 14 | 79.7 | — | |
| 16 | 78.4 | — | |
| 18 | 77.6 | — | |
| 20 | 73.5 | — | |
| 23 | 71 | — | |
| 24 | 70.5 | — | |
| 25 | 67.7 | — | |
| 26 | 65.5 | — | |
| 29 | 63.2 | — | |
| 32 | 61.3 | — | |
| 33 | 61.2 | — | |
| 39 | 55.5 | — | |
| 41 | 54.5 | — | |
| 42 | 54 | — | |
| 43 | 53.3 | — | |
| 44 | 52.2 | — | |
| 46 | 49.6 | — | |
| 51 | 46 | — | |
| 54 | 44.6 | — | |
| 55 | 44.3 | — | |
| 57 | 43.6 | — | |
| 60 | 40.1 | — | |
| 61 | 34.1 | — | |
| 63 | 31.1 | — | |
| 69 | 26.6 | — | |
| 71 | 25.3 | — | |
| 74 | 24 | — | |
| 76 | 23.1 | — | |
| 77 | 21.3 | — | |
| 78 | 20.5 | — | |
| 83 | 15.3 | — | |
| 84 | 14.6 | — | |
| 89 | 12.4 | — |
A model appears after its first published individual evaluation; its ranking score appears after its first eligible duel.
92 models compared63038 published AI verdicts
Method: regularized Bradley–Terry model. The AI-judge score determines rank; the human score is independent and remains secondary until more community votes are available. A score of 50 means an estimated one-in-two chance of beating the average model in this language and category.
Each cell shows the share of available points earned by the row model against the column model: a win is worth 100%, a tie 50%, and a loss 0%. This result matrix uses only the compatible evaluation protocols included in the main ranking.
Directional matrix of points earned in direct duels. Values above 50% favor the row model; values below 50% favor the column model.
0 %100 %
💡 marks a model with reasoning, ↯ one without reasoning. 🏠 marks a model run locally on the site author’s computer.“on*” means reasoning is enabled or disabled (local models with no intermediate levels): the requested fine-grained setting is not supported by the local file and falls back to the default enabled mode. For local models with adjustable reasoning (Qwen3.8, Muse Glimmer, Granite 4.2, Nemotron 3 Super, K2 Horizon, Ling 3.0 Flash), the displayed level (xhigh**, high, full) is the model’s default level: the requested fine-grained setting is not supported by the local file, which falls back to its default behavior.
| Model |
|---|
Each cell shows how many published AI-judged duels across all published evaluation protocols compare the two models for this language and category. The models on both axes are candidates; this matrix does not show which judge produced each verdict.
Symmetric matrix of published duel counts. The diagonal is zero because a model is not compared with itself.
0
💡 marks a model with reasoning, ↯ one without reasoning. 🏠 marks a model run locally on the site author’s computer.“on*” means reasoning is enabled or disabled (local models with no intermediate levels): the requested fine-grained setting is not supported by the local file and falls back to the default enabled mode. For local models with adjustable reasoning (Qwen3.8, Muse Glimmer, Granite 4.2, Nemotron 3 Super, K2 Horizon, Ling 3.0 Flash), the displayed level (xhigh**, high, full) is the model’s default level: the requested fine-grained setting is not supported by the local file, which falls back to its default behavior.
| Model |
|---|
Which model writes the better text? Judge this anonymous pair and contribute to the human ranking. Model names stay hidden during the duel and are revealed only after your choice.
See each AI judge’s detailed ranking and where their results differ.
Provisional results from blinded comparisons in this language.