One model, four languages, four different ranks
Ninety-one models are ranked in all four languages. The top three barely move: Claude Opus 5, Claude Fable 5 and Claude Fable 5.1 are in the top four everywhere. Below them, GLM 5.3 Flash is 10th in English and 53rd in German, Grok 4.6 25th in English and 61st in French.
The four editions of this benchmark share the same models, the same judges and the same protocol. Only the prompts change: 240 per language, written natively. So when a model ranks 10th in English and 53rd in German, the difference is not the method. It is the model.
Ninety-one model variants have a published score in all four languages (data as of 19 September 2026). Here is how far they travel.
The top does not move
| Model | English | French | German | Spanish |
|---|---|---|---|---|
| Claude Fable 5 | 3 | 3 | 3 | 3 |
| Claude Opus 5 | 2 | 1 | 1 | 2 |
| Claude Fable 5.1 | 4 | 2 | 2 | 1 |
| GPT-6 Astra | 5 | 6 | 4 | 7 |
| Kimi K3 | 1 | 5 | 5 | 4 |
| Claude Opus 4.8 | 7 | 4 | 8 | 5 |
Six models never leave the top eight. Kimi K3 is the only open-weight model among them, and it is first in English but fourth or fifth elsewhere. The scores tell the same story with more nuance: Claude Opus 5 scores 93.7 in English and 99.5 in French, meaning the French pool is easier for it to dominate, not that it writes better French than English. Scores are only comparable within a language.
The middle is where the languages disagree
| Model | English | French | German | Spanish | Spread |
|---|---|---|---|---|---|
| GLM 5.3 Flash | 10 | 27 | 53 | 44 | 43 |
| Qwen3.8 27B Q4_K_M 🏠 | 27 | 17 | 58 | 54 | 41 |
| Ling 3.0 Flash Q4_K_M 🏠 | 66 | 35 | 76 | 66 | 41 |
| Grok 4.6 | 25 | 61 | 53 | 21 | 40 |
| Hy3 | 14 | 43 | 24 | 49 | 35 |
| Seed 2.1 Turbo ↯ | 57 | 28 | 61 | 43 | 33 |
| Laguna S 2.1 UD Q4_K_M 🏠 | 58 | 87 | 91 | 87 | 33 |
GLM 5.3 Flash is a top-ten model in English and a bottom-half model in German. Grok 4.6 is strong in English and Spanish, weak in French and German. Hy3 is good in English, mediocre everywhere else. These are not sampling accidents: each model has between 1,000 and 1,800 ranked duels per language, and the confidence intervals are a few points wide.
Local models show the same pattern with an extra twist. Qwen3.8 27B at Q4_K_M is the best local model in French (17th overall) and one of the weakest in German (58th). The same file, on the same machine, with the same reasoning setting. Whatever Qwen learned about French, it did not learn about German.
The languages themselves are not equally hard
The gap between the best cloud model and the best local model is the widest in German and Spanish:
| Edition | Best model | Best local model | Local rank |
|---|---|---|---|
| English | Kimi K3, 93.8 | Muse Glimmer 30B K-Quant Dynamic, 77.6 | 18 |
| French | Claude Opus 5, 99.5 | Qwen3.8 27B Q4_K_M, 76.5 | 17 |
| German | Claude Opus 5, 98.2 | Gemma 4 31B Q8_0, 71.8 | 26 |
| Spanish | Claude Fable 5.1, 98.5 | Qwen3.8 27B Q8_0, 70.7 | 23 |
And the best local model is a different model in every language. Muse Glimmer wins English, Qwen3.8 wins French and Spanish, Gemma 4 31B wins German. Nobody who reads only one leaderboard would guess that.
What to take from it
Reading an English leaderboard to choose a model for German is a mistake for about a third of the field. The frontier closed models are safe in every language we test. Below them, check the edition in your language, and if you run models locally, check the specific quantization: the file that wins in one language can lose in the next.