The AI Writing Benchmark measures how well leading AI models write in English. Responses to a shared corpus of English writing prompts are compared blind, without model names, and the verdicts are combined into a ranking for English.
Leading across all categories
The 10 best models
The ten best models for writing in English
#
Model
Score
1
Kimi K3 π‘ enabled
93.8
2
Claude Opus 5 π‘ adaptive high
93.7
3
Claude Fable 5 π‘ adaptive high
93.2
4
Claude Fable 5.1 π‘ adaptive high
92.2
5
GPT-6 Astra π‘ high
91.4
6
GLM 5.3 π‘ max
90.3
7
Claude Opus 4.8 π‘ adaptive high
89.5
8
Claude Sonnet 5 π‘ adaptive high
86.5
9
GPT-5.6 Sol π‘ high
85.5
10
GLM 5.3 Flash π‘ high
83.1
This score is not a percentage grade: 50 means an estimated one-in-two chance of beating the average model in English.
The ranking evaluates writing quality alone. Prompts are written directly in the language being tested, and model names remain hidden until the verdict so reputation cannot influence the evaluation. This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.
Same file, same prompts, one switch. Across 21 local models, reasoning adds 33.6 points on average for the Qwen family and 10.6 for the larger Gemma 4 models; the small Gemma E2B and E4B builds gain almost nothing and sometimes lose.
Sometimes nothing, sometimes 20 points. On English prompts Qwen3.8 27B writes as well at Q4_K_M as at Q8_0 with a file half the size; Granite 4.2 30B loses 20 points between Q8 and Q4. The rule of thumb is: there is no rule of thumb.
Updated with 1,500 ranked duels per build and language. The K-Quant Dynamic build is the best local model for English writing at 77.6 points; in French, German and Spanish the smaller 17GB build is the one to download.
Two of our five judges give their own model a clear bonus: GPT-5.6 Sol awards its own texts up to 21 points more than the other judges do, Claude Opus 5 up to 11. The two Fable judges do the opposite and mark their own model down.
Claude Opus 5 and GPT-5.6 Sol, who supply most verdicts, point the same way on about nine model pairs out of ten. The disagreements are elsewhere: Sol almost never calls a tie, Claude Fable 5.1 does so up to one time in six, and Claude Opus 4.8 sides with the others only two times in three.
Ninety-one models are ranked in all four languages. The top three barely move: Claude Opus 5, Claude Fable 5 and Claude Fable 5.1 are in the top four everywhere. Below them, GLM 5.3 Flash is 10th in English and 53rd in German, Grok 4.6 25th in English and 61st in French.