In July we asked this question with 34 English judgments and answered carefully: reasoning-enabled responses earned 66% of points, but the interval included parity. Two months and 63,000 verdicts later, the careful answer has become a precise one. Every local model on this benchmark was run twice on the same 240 English prompts, once with its reasoning phase enabled and once with it disabled, and both variants took part in the same blind duels. The scores below are their positions on the main English ranking, where 50 means an even chance of beating the average model.

Twenty-one models, one switch

Local model With reasoning Without Gap
Qwen3.8 27B Q4_K_M 65.1 18.6 +46.5
Qwen3.8 27B Q8_0 61.2 20.5 +40.7
Qwen3.6 27B Q8_0 62.6 30.0 +32.6
Qwen3.8 27B Q6_K 58.7 27.8 +30.9
Qwen3.6 35B A3B Q8_0 54.5 24.0 +30.5
Gemma 4 12B Q6_K 48.9 25.0 +23.9
Qwen3.6 35B A3B Q4_K_M 45.3 24.7 +20.6
Gemma 4 12B QAT-Q4_0 46.4 30.8 +15.6
Gemma 4 26B A4B Q8_0 46.0 31.1 +14.9
Gemma 4 31B Q6_K 55.8 44.8 +11.0
Gemma 4 31B Q8_0 55.5 44.6 +10.9
Gemma 4 26B A4B Q6_K 57.1 47.9 +9.2
Gemma 4 E2B Q8_0 21.3 12.4 +8.9
Gemma 4 E2B Q6_K 13.9 6.1 +7.8
Gemma 4 12B Q8_0 44.3 40.1 +4.2
Gemma 4 E2B Q4_K_M 12.9 8.8 +4.1
Gemma 4 31B QAT-Q4_0 54.7 51.4 +3.3
Gemma 4 E4B Q6_K 18.7 16.5 +2.2
Gemma 4 26B A4B QAT-Q4_0 44.1 42.0 +2.1
Gemma 4 E4B Q8_0 14.6 15.3 -0.7
Gemma 4 E4B Q4_K_M 11.8 18.7 -6.9

Qwen needs its reasoning, Gemma much less

Every Qwen variant gains between 20.6 and 46.5 points with reasoning enabled. Qwen3.8 27B at Q4_K_M goes from 18.6 to 65.1: without reasoning it is one of the weakest models in the ranking, with it the second-best local model. The Qwen models seem to draft in the reasoning phase and only polish in the visible answer; take the drafting away and the answer is thin.

The larger Gemma 4 models (12B, 26B A4B, 31B) gain between 2.1 and 23.9 points, 10.6 on average. Reasoning helps, but Gemma writes a decent text without it, which matters when you want the answer fast. The QAT builds gain the least: Gemma 4 31B QAT-Q4_0 moves by three points, Gemma 4 26B A4B QAT-Q4_0 by two.

The small Gemma 4 E2B and E4B builds average 2.6 points and 2 of their 6 pairs are negative: for these models the reasoning phase sometimes produces a worse text than no reasoning at all.

What changed since July

The July article had four model variants and 34 English judgments; it found 66.2% of points for reasoning with an interval from 48.6% to 80.6%. The direction was right and the caution was justified. Today each pair in the table rests on 1,000 to 1,800 ranked duels per variant, the intervals are two to three points wide, and the split between families is the finding: reasoning is not a general setting that helps writing, it is a property of each model.

The same experiment in three other languages

The French, German and Spanish editions run the same 21 pairs on their own prompts. Qwen gains everywhere. Gemma’s gains are smaller and, in French and German, several Gemma builds write better with reasoning off. The setting to test is per model and per language, which is exactly what the four leaderboards let you do with the reasoning filter.