Clearer rank gaps and more accurate model terminology

AI-judge result pages now explain visible gaps in their ranks, and the English local-model guide no longer promises a parameter column that the ranking does not contain.

  • Small: Each edition now explains that AI-judge ranks are assigned before model filters hide rows, which is why gaps can remain in the visible numbering.
  • Small: The English local-AI guide now says where total and active parameter counts actually appear instead of referring to a nonexistent standalone column.

Checks: Checked with focused page tests, the complete Django test suite, and mobile browser checks of the affected public pages.

Clearer filtered ranks and leaderboard search snippets

The model leaderboard now explains visible rank gaps after filtering and describes its practical purpose more precisely in search results.

  • Small: The leaderboard now explains that selected writing categories determine the ranking first. Publisher, size, access, reasoning and quantization filters then hide models without renumbering the remaining rows, so gaps can stay visible.
  • Small: Search descriptions now identify the full leaderboard as a model-comparison page with practical filters instead of repeating the broader home-page description.

Checks: Checked with focused Django tests, the complete test suite, and mobile browser checks across all four editions.

Clearer local variants, ranks and memory guidance

Local rankings now distinguish exact variants, explain gaps left by filters, and provide verified memory guidance for three ranked Gemma 4 Q6 variants.

  • Medium: Gemma 4 12B, 26B A4B and 31B in Q6_K now show approximately 16, 32 and 32 GB of memory instead of an unknown value. Each guide includes the exact downloaded GGUF and its companion projection file before applying the existing safety reserve.
  • Small: The local leaderboard now explains that ranks are assigned across every eligible local variant before display filters are applied, so hidden variants can leave gaps in the visible rank numbers.
  • Small: The home-page local top ten is now labelled as a ranking of variants, because quantization and reasoning mode can create several tested entries from the same model checkpoint.

Checks: Checked with focused Django tests, the complete test suite, and mobile browser checks across all four editions.

Faster AI results and clearer ranking scores

AI-results pages now fetch category filters efficiently, explain wins-first ranks, and define the AI score used in Model Face-off.

  • Medium: The AI-results category query now returns each category once instead of transferring tens of thousands of duplicate rows, without changing available filters or published judgments.
  • Small: Each AI-judge table now states that total wins determine rank first and asks readers to consider wins, losses, ties, and comparisons together when coverage differs.
  • Small: Model Face-off now explains that its AI score is not a percentage grade and that 50 represents estimated even odds against the average model in the selected cohort.

Checks: Checked with focused Django tests, the complete test suite, and mobile browser checks across all four editions.

A clearer local-model score and a keyboard-ready human ranking

The local-model ranking now explains its score scale, and wide human-ranking tables can be reached and panned with a keyboard.

  • Small: The local-model page now states that its score uses the main leaderboard scale, is not a percentage grade, and explains what 50 means for the active category selection.
  • Small: A populated human-ranking table now has keyboard focus and a localized accessible region name for horizontal navigation.

Checks: Checked with focused Django tests, the complete test suite, and mobile browser checks across all four editions.

Clearer scores and more transparent AI judging

All four editions now explain leaderboard scores on the home page, disclose possible self-model judging, and make wide results tables reachable by keyboard.

  • Medium: Every horizontally scrollable ranking, matrix and judge table now has keyboard focus and a distinct accessible region name.
  • Small: The home leaderboard now says that its score is not a percentage grade and explains what a score of 50 means.
  • Small: The AI-results introduction now notes that a judge may evaluate a pair containing its own model’s response while answer identities remain hidden.

Checks: Checked with focused Django tests, the complete test suite, and mobile browser checks across all four language editions.

Muse Glimmer now compared with local rivals

The four native editions now show the available head-to-head evidence against Gemma, Qwen and Nemotron instead of hiding it behind the stricter main-duel threshold.

  • Medium: The Muse Glimmer article now includes provisional, sample-labelled comparisons for both quantizations against four local models, with counts and language-specific results.
  • Small: All four editions received new native titles and revised copy; the French article was rewritten more extensively.
  • Small: A new accessible chart replaces the judge-disagreement view and displays eight local-rival matchups without mobile overflow.

Checks: Checked against the production snapshot, editorial and SEO tests, the full Django suite, and mobile browser rendering in all four languages.

Muse Glimmer 30B tested for local writing

A new data-led feature compares Muse Glimmer’s two official quantizations and explains what the early writing evidence can, and cannot, support.

  • Medium: The new Muse Glimmer analysis combines 100 blind AI judgments across English, French, German and Spanish with official specifications and early Arena evidence.
  • Small: Five accessible charts cover the direct quantization duel, judge agreement, file sizes, official DFlash speeds and Arena uncertainty.
  • Small: The article gives practical memory, privacy, licensing and factual-verification guidance without ranking Muse against peers that have not reached the same evidence floor.

Checks: Checked with source-backed data validation, focused editorial tests, the complete 972-test Django suite and local browser rendering of all five charts.

Faster results with clearer model details

The benchmark now answers two practical reader questions more directly: which judge produced a ranking, and how much graphics memory a local model needs.

  • Medium: AI-results no longer assembles internal run provenance that the page never displays. Rankings, scores, verdict counts and freshness dates are unchanged.
  • Small: The historical evaluator is now labelled Claude Opus 4.8, with its decimal point restored.
  • Small: The German home page now reports local-model graphics memory in GB, matching its detailed local-AI page, instead of using the French abbreviation Go.

Checks: Checked with focused page tests, the complete Django suite and public verification across all four editions.

A safer eight-hour audit cycle

The benchmark can now improve on a regular schedule without turning maintenance into an unchecked automation loop.

  • Medium: Added a bounded systemd audit every eight hours. GPT-5.6 Sol may implement only changes approved independently by Claude Opus 5.
  • Small: Added release gates that require restoration of the previous version after failed tests, public checks or a rejected post-publication review.
  • Small: Published this timestamped changelog in four independently written editions.

Checks: Full Django test suite, four-site public checks and independent Opus 5 review required.