The four editions of this benchmark share the same models, the same judges and the same protocol. Only the prompts change: 240 per language, written natively. So when a model ranks 10th in English and 53rd in German, the difference is not the method. It is the model.

Ninety-one model variants have a published score in all four languages (data as of 19 September 2026). Here is how far they travel.

The top does not move

Model English French German Spanish
Claude Fable 5 3 3 3 3
Claude Opus 5 2 1 1 2
Claude Fable 5.1 4 2 2 1
GPT-6 Astra 5 6 4 7
Kimi K3 1 5 5 4
Claude Opus 4.8 7 4 8 5

Six models never leave the top eight. Kimi K3 is the only open-weight model among them, and it is first in English but fourth or fifth elsewhere. The scores tell the same story with more nuance: Claude Opus 5 scores 93.7 in English and 99.5 in French, meaning the French pool is easier for it to dominate, not that it writes better French than English. Scores are only comparable within a language.

The middle is where the languages disagree

Model English French German Spanish Spread
GLM 5.3 Flash 10 27 53 44 43
Qwen3.8 27B Q4_K_M 🏠 27 17 58 54 41
Ling 3.0 Flash Q4_K_M 🏠 66 35 76 66 41
Grok 4.6 25 61 53 21 40
Hy3 14 43 24 49 35
Seed 2.1 Turbo ↯ 57 28 61 43 33
Laguna S 2.1 UD Q4_K_M 🏠 58 87 91 87 33

GLM 5.3 Flash is a top-ten model in English and a bottom-half model in German. Grok 4.6 is strong in English and Spanish, weak in French and German. Hy3 is good in English, mediocre everywhere else. These are not sampling accidents: each model has between 1,000 and 1,800 ranked duels per language, and the confidence intervals are a few points wide.

Local models show the same pattern with an extra twist. Qwen3.8 27B at Q4_K_M is the best local model in French (17th overall) and one of the weakest in German (58th). The same file, on the same machine, with the same reasoning setting. Whatever Qwen learned about French, it did not learn about German.

The languages themselves are not equally hard

The gap between the best cloud model and the best local model is the widest in German and Spanish:

Edition Best model Best local model Local rank
English Kimi K3, 93.8 Muse Glimmer 30B K-Quant Dynamic, 77.6 18
French Claude Opus 5, 99.5 Qwen3.8 27B Q4_K_M, 76.5 17
German Claude Opus 5, 98.2 Gemma 4 31B Q8_0, 71.8 26
Spanish Claude Fable 5.1, 98.5 Qwen3.8 27B Q8_0, 70.7 23

And the best local model is a different model in every language. Muse Glimmer wins English, Qwen3.8 wins French and Spanish, Gemma 4 31B wins German. Nobody who reads only one leaderboard would guess that.

What to take from it

Reading an English leaderboard to choose a model for German is a mistake for about a third of the field. The frontier closed models are safe in every language we test. Below them, check the edition in your language, and if you run models locally, check the specific quantization: the file that wins in one language can lose in the next.