My name is Johannes Eva, and I have been building websites for 27 years. The AI Writing Benchmark began with a gap: no existing leaderboard measured what interested me, the writing quality of AI models in English and other languages.

Public benchmarks mainly focus on code, mathematics and reasoning. The small part devoted to writing is usually evaluated only in English. Here, only the quality of the generated texts counts. The prompts are written directly in each language, and the answers are compared blindly.

Models that can run at home have a special place here. Gemma 4, Qwen 3.6 and 3.8, Ling 3.0 Flash, Muse Glimmer and Granite 4.2, among others, are marked with 🏠. They were run locally, in several quantisations, and their reasoning on and off.

This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.

The machines used for local models

Framework Desktop used to run local language models for the AI Writing Benchmark
Computer
Framework Desktop, AMD Ryzen AI Max 300 Series, revision A6
Processor
AMD Ryzen AI Max+ 395, 16 cores and 32 threads, up to 5.19 GHz
Graphics
AMD Radeon 8060S integrated graphics
Memory
128 GB of unified memory, shared by the processor and graphics unit

Local models run on a Framework Desktop with an AMD Ryzen AI Max+ 395 and 128 GB of unified memory available to both the processor and graphics unit. This makes it possible to load models larger than most consumer graphics cards can hold. A second machine with an NVIDIA GeForce RTX 5070 Ti contributes the smaller models.

Observed local generation speed (Strix Halo)

Median tokens generated per second on the Strix Halo machine (Framework Desktop, Ryzen AI Max+ 395) across valid benchmark responses generated there, including reasoning tokens. Speed varies with context length and software. Models without enough Strix Halo measurements are not listed.

ModelQuantisationMedian speedResponses
Gemma 4 26B A4BQ8_039.8 tokens/s465
Gemma 4 E2BQ4_K_M39.4 tokens/s224
Gemma 4 E4BQ8_035.9 tokens/s542
Gemma 4 E2BQ6_K33.2 tokens/s216
Ling 3.0 FlashQ4_K_M31.6 tokens/s277
Ling 3.0 FlashQ4_K_M30.5 tokens/s240
Gemma 4 E2BQ8_029.2 tokens/s224
Gemma 4 26B A4B QATQAT-Q4_027.9 tokens/s118
Qwen3.6 35B A3BQ4_K_M26.6 tokens/s347
Nemotron 3 Nano OmniQ4_K_M25.1 tokens/s960
Gemma 4 E4BQ4_K_M22.4 tokens/s221
Qwen3.6 35B A3BQ8_022.1 tokens/s312
Nemotron 3 Nano OmniQ8_021.2 tokens/s391
Nemotron 3 SuperQ4_K_M21.1 tokens/s413
Gemma 4 26B A4BQ6_K20.2 tokens/s539
Gemma 4 E4BQ6_K18.6 tokens/s218
Granite 4.2 8BQ4_K_M16.9 tokens/s36
Gemma 4 12BQ8_014.6 tokens/s124
Gemma 4 12B QATQAT-Q4_011 tokens/s752
Gemma 4 31B QATQAT-Q4_010 tokens/s548
Gemma 4 12BQ6_K9 tokens/s216
Qwen3.8 27BQ4_K_M8.9 tokens/s546
Muse Glimmer 30B Kquant 17gbK-Quant 17GB8.8 tokens/s394
Qwen3.8 27BQ8_08.4 tokens/s498
Muse Glimmer 30B Kquant DynamicK-Quant Dynamic8 tokens/s397
Granite 4.2 30BQ4_K_M7.6 tokens/s423
Qwen3.6 27BQ8_07.5 tokens/s190
Qwen3.8 27BQ6_K7 tokens/s504
Granite 4.2 30BQ6_K6.6 tokens/s397
Gemma 4 31BQ6_K5.8 tokens/s527
Qwen3.6 27BQ4_K_M5.2 tokens/s588
Granite 4.2 30BQ8_04.9 tokens/s447
Gemma 4 31BQ8_04.8 tokens/s316

Comparable published benchmarks for Ryzen AI Max+ 395