Q4, Q6 or Q8: what quantization does to writing quality
Sometimes nothing, sometimes 20 points. On English prompts Qwen3.8 27B writes as well at Q4_K_M as at Q8_0 with a file half the size; Granite 4.2 30B loses 20 points between Q8 and Q4. The rule of thumb is: there is no rule of thumb.
Quantization shrinks a model by storing its weights with fewer bits. A 27-billion-parameter model is 27 GB at Q8_0 and 16 GB at Q4_K_M, which is the difference between fitting a 32 GB graphics card or not. Every guide says the smaller file writes a little worse. We measured how much, on English writing, with the same blind duels that build the main ranking. Scores are the main-leaderboard scale for the English edition: 50 means an even chance of beating the average model. Sizes are the GGUF files; speeds are medians measured on our Strix Halo machine, reasoning tokens included.
Same model, three files, three scores
| Family | Quantization | Score | File (GB) | Speed (tok/s) | Duels |
|---|---|---|---|---|---|
| Qwen3.8 27B | Q4_K_M | 65.1 | 15.7 | 8.9 | 1158 |
| Q8_0 | 61.2 | 27.1 | 8.4 | 1175 | |
| Q6_K | 58.7 | 20.9 | 7.0 | 1163 | |
| Qwen3.6 27B | Q8_0 | 62.6 | 26.6 | 7.5 | 1019 |
| Q4_K_M | 55.9 | 15.4 | 5.2 | 1157 | |
| Qwen3.6 35B A3B | Q8_0 | 54.5 | 34.4 | 22.1 | 1776 |
| Q4_K_M | 45.3 | 19.7 | 26.6 | 1717 | |
| Gemma 4 31B | Q6_K | 55.8 | 23.5 | 5.8 | 1170 |
| Q8_0 | 55.5 | 30.4 | 4.8 | 1846 | |
| QAT-Q4_0 | 54.7 | 16.4 | 10.0 | 1777 | |
| Gemma 4 26B A4B | Q6_K | 57.1 | 21.1 | 20.2 | 1179 |
| Q8_0 | 46.0 | 25.0 | 39.8 | 1655 | |
| QAT-Q4_0 | 44.1 | 13.4 | 27.9 | 1787 | |
| Gemma 4 12B | Q6_K | 48.9 | 9.1 | 9.0 | 1096 |
| QAT-Q4_0 | 46.4 | 6.5 | 11.0 | 1652 | |
| Q8_0 | 44.3 | 11.8 | 14.6 | 1631 | |
| Muse Glimmer 30B | K-Quant Dynamic (≈ Q5) | 77.6 | 18.3 | 8.0 | 1536 |
| K-Quant 17GB (≈ Q4) | 64.4 | 15.6 | 8.8 | 1536 | |
| Granite 4.2 30B | Q8_0 | 34.1 | 29.0 | 4.9 | 1080 |
| Q6_K | 26.8 | 22.4 | 6.6 | 975 | |
| Q4_K_M | 14.1 | 16.5 | 7.6 | 1065 | |
| Nemotron 3 Nano Omni 30B A3B | Q8_0 | 23.1 | 31.3 | 21.2 | 1045 |
| Q4_K_M | 14.4 | 22.8 | 25.1 | 1464 |
Where quantization is free
Gemma 4 31B is the model that does not care: QAT-Q4_0, Q6_K and Q8_0 land within a point of each other, and the 16 GB file runs twice as fast as the 30 GB one. Google trained the QAT build specifically to survive 4-bit storage, and here it shows. Qwen3.8 27B goes further: its Q4_K_M build is the best of the three in English, four points above Q8_0, on more than 1,100 duels each. Two builds of Muse Glimmer 30B differ in the other direction: the larger K-Quant Dynamic (18 GB) beats the 17GB build by 13 points and is the best local model in English.
Where it is expensive
Granite 4.2 30B loses seven points from Q8 to Q6 and another thirteen to Q4; the Q4 file is not worth its 16 GB. Nemotron 3 Nano loses nine points to Q4. Qwen3.6 27B and the 35B A3B mixture both lose seven to nine points at Q4. Gemma 4 26B A4B is the odd one: its Q6_K build is eleven points above both Q8_0 and the QAT build, on a thousand duels, and we do not have an explanation yet.
Q6 is rarely the answer
Across nine families, Q6_K is the best of its family only twice (Gemma 4 26B A4B and Gemma 4 12B) and often the worst. It saves less memory than Q4 and rarely keeps the quality of Q8. If you must choose blind, choose Q8 when it fits and a QAT or K-Quant Q4 when it does not, and skip Q6.
The same table in another language says something else
This is English. The French, German and Spanish editions of this article use the same files and reach different verdicts: Qwen3.8 27B at Q4_K_M loses 21 points against Q8_0 in German and 27 in Spanish. A quantization is not good or bad in itself; it is good or bad for a model in a language, which is why we measure it per edition.