Quality

Where each model is right, wrong, or cut off

Every chart on this page has one message, stated above it. Bars are counts of questions, split into correct, wrong, and "not finished" (the model hit the token limit while still reasoning, which counts as wrong). Hover for the numbers; the explorer shows the questions themselves.

Budget changes the ranking, not the models

At 12k tokens Muse leads MMLU-Pro; at 32k the Qwens catch up, by finishing more.

The Qwens are not worse reasoners; they are slower to finish

GPQA Diamond: same score (80.3 vs 79.8%), opposite error profiles.

198 graduate-level science questions, 32k budget. Muse finishes nearly everything and is wrong on 38; Qwen3.8 is right on 95% of what it finishes but does not finish 31. Per question: 14 only Muse, 13 only Qwen3.8, a tie. Qwen3.6 did not run this task (left out by decision after tying Qwen3.8 on the others).

On competitive programming the same mechanism separates them clearly

LiveCodeBench v6: Muse 79%, Qwen3.8 60%. Qwen3.8 does not finish 38% of the problems at 32k.

175 problems from contests of Jan–Apr 2025, graded by running every test (judge environment with numpy/numba available, as AtCoder's). Per problem: 34 only Muse, 1 only Qwen3.8, 104 both. Of Qwen3.8's 66 unfinished problems, Muse solved 34. By difficulty: easy 43/43 for both; hard 49/80 vs 21/80.
Two scores for LiveCodeBench: judge environment vs strict

With only the standard library available (strict), Muse scores 66.9%: 23 of its solutions import numba, a JIT compiler the real AtCoder judge provides. With those libraries available, 21 of them pass (all hard problems). Qwen3.8 used no such libraries, so its score is 60.0% either way. Both are in the appendix.

The same model gives the same scores on any machine

Qwen3.8 run three times (A6000 without MTP, A6000 with MTP, Spark): within ±1.5 points.

Identical checkpoint (unsloth/Qwen3.8-27B-NVFP4) and recipe, 12k budget. This spread is the run-to-run noise (temperature 1.0), and the scale to read every other difference against. It also means everything measured on the A6000 applies to the Spark.

Quantization does not change the outcome

Muse-Glimmer, 12k budgetQ4 GGUF (llama.cpp)NVFP4 (vLLM)

Per question on MMLU-Pro: 5 only Q4, 3 only NVFP4, 159 both. NVFP4 here is NVIDIA's AutoQuant checkpoint (mixed NVFP4/FP8/BF16); Q4 is unsloth's dynamic GGUF. Muse vs Qwen3.8 both in NVFP4 on the same vLLM shows the same MMLU-Pro gap as the GGUF runs (17 vs 4 per question).

GSM8K and HumanEval do not separate these models

12k budgetMuseQwen3.6Qwen3.8

Everyone scores 96–99%. They stay in the study as sanity checks with known expected scores (they validated the harness and showed that quantization, speculative decoding and hardware do not change answers), but they cannot rank the models. That is why the harder tasks above were added.