Last updated:

Writing model ranking based on pairwise results
RankModelAI judge scoreHuman score
1Kimi K3 💡 enabled93.8
2Claude Opus 5 💡 adaptive high93.7
4Claude Fable 5.1 💡 adaptive high92.2
5GPT-6 Astra 💡 high91.4
6GLM 5.3 💡 max90.3
8Claude Sonnet 5 💡 adaptive high86.5
9GPT-5.6 Sol 💡 high85.5
10GLM 5.3 Flash 💡 high83.1
11MiniMax M3 💡 enabled81.8
12Muse Spark 1.3 💡 high80.8
14DeepSeek V4.1 Flash 💡 high79.7
14Hy3 💡 high79.7
16Gemini 3.8 Flash 💡 high78.4
18Qwen3.8 Max 💡 xhigh77.6
20Aion 3.0 💡 mandatory73.5
23GPT-5.6 Terra 💡 high71
24Hy4 Preview 💡 high70.5
25Grok 4.6 💡 high67.7
26GPT-5.6 Luna 💡 high65.5
29DeepSeek V4 Pro 💡 enabled63.2
32Qwen3.8 Flash 💡 high61.3
33Qwen3.8 27B Q8_0 💡 xhigh** 🏠61.2
39Gemma 4 31B Q8_0 💡 on* 🏠55.5
41Qwen3.6 35B A3B Q8_0 💡 on* 🏠54.5
42Qwen3.7 Plus 💡 enabled54
43MiMo V2.5 Pro 💡 enabled53.3
44Inkling 💡 high52.2
46MiMo V2.5 💡 enabled49.6
51Gemma 4 26B A4B Q8_0 💡 on* 🏠46
54Gemma 4 31B Q8_0 🏠44.6
55Gemma 4 12B Q8_0 💡 on* 🏠44.3
57Seed 2.1 Turbo43.6
60Gemma 4 12B Q8_0 🏠40.1
61Granite 4.2 30B Q8_0 💡 full 🏠34.1
63Gemma 4 26B A4B Q8_0 🏠31.1
69Mistral Medium 3.5 💡 high26.6
71Granite 4.2 8B Q8_0 💡 full 🏠25.3
74Qwen3.6 35B A3B Q8_0 🏠24
76Nemotron 3 Nano Omni 30B A3B Q8_0 💡 on* 🏠23.1
77Gemma 4 E2B Q8_0 💡 on* 🏠21.3
78Qwen3.8 27B Q8_0 🏠20.5
83Gemma 4 E4B Q8_0 🏠15.3
84Gemma 4 E4B Q8_0 💡 on* 🏠14.6
89Gemma 4 E2B Q8_0 🏠12.4

92 models compared63038 published AI verdicts

Method: regularized Bradley–Terry model. The AI-judge score determines rank; the human score is independent and remains secondary until more community votes are available. A score of 50 means an estimated one-in-two chance of beating the average model in this language and category.

Direct duel results

Each cell shows the share of available points earned by the row model against the column model: a win is worth 100%, a tie 50%, and a loss 0%. This result matrix uses only the compatible evaluation protocols included in the main ranking.

Directional matrix of points earned in direct duels. Values above 50% favor the row model; values below 50% favor the column model.

0 %100 %

💡 marks a model with reasoning, ↯ one without reasoning. 🏠 marks a model run locally on the site author’s computer.“on*” means reasoning is enabled or disabled (local models with no intermediate levels): the requested fine-grained setting is not supported by the local file and falls back to the default enabled mode. For local models with adjustable reasoning (Qwen3.8, Muse Glimmer, Granite 4.2, Nemotron 3 Super, K2 Horizon, Ling 3.0 Flash), the displayed level (xhigh**, high, full) is the model’s default level: the requested fine-grained setting is not supported by the local file, which falls back to its default behavior.

View as a data table
Points earned in direct duels by model pair
Model

Published duel coverage

Each cell shows how many published AI-judged duels across all published evaluation protocols compare the two models for this language and category. The models on both axes are candidates; this matrix does not show which judge produced each verdict.

Symmetric matrix of published duel counts. The diagonal is zero because a model is not compared with itself.

0

💡 marks a model with reasoning, ↯ one without reasoning. 🏠 marks a model run locally on the site author’s computer.“on*” means reasoning is enabled or disabled (local models with no intermediate levels): the requested fine-grained setting is not supported by the local file and falls back to the default enabled mode. For local models with adjustable reasoning (Qwen3.8, Muse Glimmer, Granite 4.2, Nemotron 3 Super, K2 Horizon, Ling 3.0 Flash), the displayed level (xhigh**, high, full) is the model’s default level: the requested fine-grained setting is not supported by the local file, which falls back to its default behavior.

View as a data table
Published duel counts by model pair
Model

Help elect the best LLM for writing in English

Which model writes the better text? Judge this anonymous pair and contribute to the human ranking. Model names stay hidden during the duel and are revealed only after your choice.

Arena

AI judge results

See each AI judge’s detailed ranking and where their results differ.

Human ranking

Provisional results from blinded comparisons in this language.