The models that write best in English

The AI Writing Benchmark measures how well leading AI models write in English. Responses to a shared corpus of English writing prompts are compared blind, without model names, and the verdicts are combined into a ranking for English.

Leading across all categories

The 10 best models

The ten best models for writing in English
#ModelScore
1Kimi K3 πŸ’‘ enabled93.8
2Claude Opus 5 πŸ’‘ adaptive high93.7
3Claude Fable 5 πŸ’‘ adaptive high93.2
4Claude Fable 5.1 πŸ’‘ adaptive high92.2
5GPT-6 Astra πŸ’‘ high91.4
6GLM 5.3 πŸ’‘ max90.3
7Claude Opus 4.8 πŸ’‘ adaptive high89.5
8Claude Sonnet 5 πŸ’‘ adaptive high86.5
9GPT-5.6 Sol πŸ’‘ high85.5
10GLM 5.3 Flash πŸ’‘ high83.1

This score is not a percentage grade: 50 means an estimated one-in-two chance of beating the average model in English.

Last updated: See the full leaderboard

The 10 best local variants

Estimated VRAM is the approximate graphics memory needed to load the entire model.

The ten best exact variants run locally
#ModelScoreEstimated VRAM
1Muse Glimmer 30B K-Quant Dynamic πŸ’‘ high 🏠77.6β‰ˆ 32 GB
2Qwen3.8 27B Q4_K_M πŸ’‘ xhigh** 🏠65.1β‰ˆ 24 GB
3Muse Glimmer 30B K-Quant 17GB πŸ’‘ high 🏠64.4β‰ˆ 24 GB
4Qwen3.6 27B Q8_0 πŸ’‘ on* 🏠62.6β‰ˆ 32 GB
5Qwen3.8 27B Q8_0 πŸ’‘ xhigh** 🏠61.2β‰ˆ 48 GB
6Qwen3.8 27B Q6_K πŸ’‘ xhigh** 🏠58.7β‰ˆ 32 GB
7Gemma 4 26B A4B Q6_K πŸ’‘ on* 🏠57.1β‰ˆ 32 GB
8Qwen3.6 27B Q4_K_M πŸ’‘ on* 🏠55.9β‰ˆ 24 GB
9Gemma 4 31B Q6_K πŸ’‘ on* 🏠55.8β‰ˆ 32 GB
10Gemma 4 31B Q8_0 πŸ’‘ on* 🏠55.5β‰ˆ 48 GB

Some software can keep part of the model in system memory. This reduces the VRAM required, but usually makes generation slower.

Local AI β†’

92models compared
63,038published AI verdicts

The ranking evaluates writing quality alone. Prompts are written directly in the language being tested, and model names remain hidden until the verdict so reputation cannot influence the evaluation. This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.

Explore the results

Does AI write better in English with or without reasoning?

Same file, same prompts, one switch. Across 21 local models, reasoning adds 33.6 points on average for the Qwen family and 10.6 for the larger Gemma 4 models; the small Gemma E2B and E4B builds gain almost nothing and sometimes lose.

Q4, Q6 or Q8: what quantization does to writing quality

Sometimes nothing, sometimes 20 points. On English prompts Qwen3.8 27B writes as well at Q4_K_M as at Q8_0 with a file half the size; Granite 4.2 30B loses 20 points between Q8 and Q4. The rule of thumb is: there is no rule of thumb.

Do AI judges prefer their own model?

Two of our five judges give their own model a clear bonus: GPT-5.6 Sol awards its own texts up to 21 points more than the other judges do, Claude Opus 5 up to 11. The two Fable judges do the opposite and mark their own model down.

How much do the AI judges agree with each other?

Claude Opus 5 and GPT-5.6 Sol, who supply most verdicts, point the same way on about nine model pairs out of ten. The disagreements are elsewhere: Sol almost never calls a tie, Claude Fable 5.1 does so up to one time in six, and Claude Opus 4.8 sides with the others only two times in three.

One model, four languages, four different ranks

Ninety-one models are ranked in all four languages. The top three barely move: Claude Opus 5, Claude Fable 5 and Claude Fable 5.1 are in the top four everywhere. Below them, GLM 5.3 Flash is 10th in English and 53rd in German, Grok 4.6 25th in English and 61st in French.