Does AI write better in English with or without reasoning?
Same file, same prompts, one switch. Across 21 local models, reasoning adds 33.6 points on average for the Qwen family and 10.6 for the larger Gemma 4 models; the small Gemma E2B and E4B builds gain almost nothing and sometimes lose.
In July we asked this question with 34 English judgments and answered carefully: reasoning-enabled responses earned 66% of points, but the interval included parity. Two months and 63,000 verdicts later, the careful answer has become a precise one. Every local model on this benchmark was run twice on the same 240 English prompts, once with its reasoning phase enabled and once with it disabled, and both variants took part in the same blind duels. The scores below are their positions on the main English ranking, where 50 means an even chance of beating the average model.
Twenty-one models, one switch
| Local model | With reasoning | Without | Gap |
|---|---|---|---|
| Qwen3.8 27B Q4_K_M | 65.1 | 18.6 | +46.5 |
| Qwen3.8 27B Q8_0 | 61.2 | 20.5 | +40.7 |
| Qwen3.6 27B Q8_0 | 62.6 | 30.0 | +32.6 |
| Qwen3.8 27B Q6_K | 58.7 | 27.8 | +30.9 |
| Qwen3.6 35B A3B Q8_0 | 54.5 | 24.0 | +30.5 |
| Gemma 4 12B Q6_K | 48.9 | 25.0 | +23.9 |
| Qwen3.6 35B A3B Q4_K_M | 45.3 | 24.7 | +20.6 |
| Gemma 4 12B QAT-Q4_0 | 46.4 | 30.8 | +15.6 |
| Gemma 4 26B A4B Q8_0 | 46.0 | 31.1 | +14.9 |
| Gemma 4 31B Q6_K | 55.8 | 44.8 | +11.0 |
| Gemma 4 31B Q8_0 | 55.5 | 44.6 | +10.9 |
| Gemma 4 26B A4B Q6_K | 57.1 | 47.9 | +9.2 |
| Gemma 4 E2B Q8_0 | 21.3 | 12.4 | +8.9 |
| Gemma 4 E2B Q6_K | 13.9 | 6.1 | +7.8 |
| Gemma 4 12B Q8_0 | 44.3 | 40.1 | +4.2 |
| Gemma 4 E2B Q4_K_M | 12.9 | 8.8 | +4.1 |
| Gemma 4 31B QAT-Q4_0 | 54.7 | 51.4 | +3.3 |
| Gemma 4 E4B Q6_K | 18.7 | 16.5 | +2.2 |
| Gemma 4 26B A4B QAT-Q4_0 | 44.1 | 42.0 | +2.1 |
| Gemma 4 E4B Q8_0 | 14.6 | 15.3 | -0.7 |
| Gemma 4 E4B Q4_K_M | 11.8 | 18.7 | -6.9 |
Qwen needs its reasoning, Gemma much less
Every Qwen variant gains between 20.6 and 46.5 points with reasoning enabled. Qwen3.8 27B at Q4_K_M goes from 18.6 to 65.1: without reasoning it is one of the weakest models in the ranking, with it the second-best local model. The Qwen models seem to draft in the reasoning phase and only polish in the visible answer; take the drafting away and the answer is thin.
The larger Gemma 4 models (12B, 26B A4B, 31B) gain between 2.1 and 23.9 points, 10.6 on average. Reasoning helps, but Gemma writes a decent text without it, which matters when you want the answer fast. The QAT builds gain the least: Gemma 4 31B QAT-Q4_0 moves by three points, Gemma 4 26B A4B QAT-Q4_0 by two.
The small Gemma 4 E2B and E4B builds average 2.6 points and 2 of their 6 pairs are negative: for these models the reasoning phase sometimes produces a worse text than no reasoning at all.
What changed since July
The July article had four model variants and 34 English judgments; it found 66.2% of points for reasoning with an interval from 48.6% to 80.6%. The direction was right and the caution was justified. Today each pair in the table rests on 1,000 to 1,800 ranked duels per variant, the intervals are two to three points wide, and the split between families is the finding: reasoning is not a general setting that helps writing, it is a property of each model.
The same experiment in three other languages
The French, German and Spanish editions run the same 21 pairs on their own prompts. Qwen gains everywhere. Gemma’s gains are smaller and, in French and German, several Gemma builds write better with reasoning off. The setting to test is per model and per language, which is exactly what the four leaderboards let you do with the reasoning filter.