Where each model is right, wrong, or cut off
Every chart on this page has one message, stated above it. Bars are counts of questions, split into correct, wrong, and "not finished" (the model hit the token limit while still reasoning, which counts as wrong). Hover for the numbers; the explorer shows the questions themselves.
Budget changes the ranking, not the models
At 12k tokens Muse leads MMLU-Pro; at 32k the Qwens catch up, by finishing more.
The Qwens are not worse reasoners; they are slower to finish
GPQA Diamond: same score (80.3 vs 79.8%), opposite error profiles.
On competitive programming the same mechanism separates them clearly
LiveCodeBench v6: Muse 79%, Qwen3.8 60%. Qwen3.8 does not finish 38% of the problems at 32k.
Two scores for LiveCodeBench: judge environment vs strict
With only the standard library available (strict), Muse scores 66.9%: 23 of its solutions import numba, a JIT compiler the real AtCoder judge provides. With those libraries available, 21 of them pass (all hard problems). Qwen3.8 used no such libraries, so its score is 60.0% either way. Both are in the appendix.
The same model gives the same scores on any machine
Qwen3.8 run three times (A6000 without MTP, A6000 with MTP, Spark): within ±1.5 points.
Quantization does not change the outcome
| Muse-Glimmer, 12k budget | Q4 GGUF (llama.cpp) | NVFP4 (vLLM) |
|---|
Per question on MMLU-Pro: 5 only Q4, 3 only NVFP4, 159 both. NVFP4 here is NVIDIA's AutoQuant checkpoint (mixed NVFP4/FP8/BF16); Q4 is unsloth's dynamic GGUF. Muse vs Qwen3.8 both in NVFP4 on the same vLLM shows the same MMLU-Pro gap as the GGUF runs (17 vs 4 per question).
GSM8K and HumanEval do not separate these models
| 12k budget | Muse | Qwen3.6 | Qwen3.8 |
|---|
Everyone scores 96–99%. They stay in the study as sanity checks with known expected scores (they validated the harness and showed that quantization, speculative decoding and hardware do not change answers), but they cannot rank the models. That is why the harder tasks above were added.