Quantization shrinks a model by storing its weights with fewer bits. A 27-billion-parameter model is 27 GB at Q8_0 and 16 GB at Q4_K_M, which is the difference between fitting a 32 GB graphics card or not. Every guide says the smaller file writes a little worse. We measured how much, on English writing, with the same blind duels that build the main ranking. Scores are the main-leaderboard scale for the English edition: 50 means an even chance of beating the average model. Sizes are the GGUF files; speeds are medians measured on our Strix Halo machine, reasoning tokens included.

Same model, three files, three scores

Family Quantization Score File (GB) Speed (tok/s) Duels
Qwen3.8 27B Q4_K_M 65.1 15.7 8.9 1158
Q8_0 61.2 27.1 8.4 1175
Q6_K 58.7 20.9 7.0 1163
Qwen3.6 27B Q8_0 62.6 26.6 7.5 1019
Q4_K_M 55.9 15.4 5.2 1157
Qwen3.6 35B A3B Q8_0 54.5 34.4 22.1 1776
Q4_K_M 45.3 19.7 26.6 1717
Gemma 4 31B Q6_K 55.8 23.5 5.8 1170
Q8_0 55.5 30.4 4.8 1846
QAT-Q4_0 54.7 16.4 10.0 1777
Gemma 4 26B A4B Q6_K 57.1 21.1 20.2 1179
Q8_0 46.0 25.0 39.8 1655
QAT-Q4_0 44.1 13.4 27.9 1787
Gemma 4 12B Q6_K 48.9 9.1 9.0 1096
QAT-Q4_0 46.4 6.5 11.0 1652
Q8_0 44.3 11.8 14.6 1631
Muse Glimmer 30B K-Quant Dynamic (≈ Q5) 77.6 18.3 8.0 1536
K-Quant 17GB (≈ Q4) 64.4 15.6 8.8 1536
Granite 4.2 30B Q8_0 34.1 29.0 4.9 1080
Q6_K 26.8 22.4 6.6 975
Q4_K_M 14.1 16.5 7.6 1065
Nemotron 3 Nano Omni 30B A3B Q8_0 23.1 31.3 21.2 1045
Q4_K_M 14.4 22.8 25.1 1464

Where quantization is free

Gemma 4 31B is the model that does not care: QAT-Q4_0, Q6_K and Q8_0 land within a point of each other, and the 16 GB file runs twice as fast as the 30 GB one. Google trained the QAT build specifically to survive 4-bit storage, and here it shows. Qwen3.8 27B goes further: its Q4_K_M build is the best of the three in English, four points above Q8_0, on more than 1,100 duels each. Two builds of Muse Glimmer 30B differ in the other direction: the larger K-Quant Dynamic (18 GB) beats the 17GB build by 13 points and is the best local model in English.

Where it is expensive

Granite 4.2 30B loses seven points from Q8 to Q6 and another thirteen to Q4; the Q4 file is not worth its 16 GB. Nemotron 3 Nano loses nine points to Q4. Qwen3.6 27B and the 35B A3B mixture both lose seven to nine points at Q4. Gemma 4 26B A4B is the odd one: its Q6_K build is eleven points above both Q8_0 and the QAT build, on a thousand duels, and we do not have an explanation yet.

Q6 is rarely the answer

Across nine families, Q6_K is the best of its family only twice (Gemma 4 26B A4B and Gemma 4 12B) and often the worst. It saves less memory than Q4 and rarely keeps the quality of Q8. If you must choose blind, choose Q8 when it fits and a QAT or K-Quant Q4 when it does not, and skip Q6.

The same table in another language says something else

This is English. The French, German and Spanish editions of this article use the same files and reach different verdicts: Qwen3.8 27B at Q4_K_M loses 21 points against Q8_0 in German and 27 in Spanish. A quantization is not good or bad in itself; it is good or bad for a model in a language, which is why we measure it per edition.