Same quality if you can afford the wait; Muse if you can't
Muse-Glimmer-30B, Qwen3.6-27B and Qwen3.8-27B, run locally on the same questions, prompts, sampling and token budgets. Every answer was saved and every wrong answer inspected. The three findings below carry the decision; each links to the page that holds the evidence.
1. Given room to finish thinking, the three models are equivalent
MMLU-Pro with a 32k-token budget: 82 / 82 / 80%.
2. Muse gets there with less than half the tokens
Tokens per answer on the same questions: 1,592 vs 4,083 vs 4,935.
3. With a tight budget, Muse wins because it finishes
MMLU-Pro at 12k tokens: Muse 82%, Qwen3.6 73%, Qwen3.8 77%. The gap is answers that were never finished.
What decides it
The token budget is the whole story, and it is a product decision, not a benchmark one. A 32k-token answer takes about 11 minutes on the A6000 with speculative decoding and over 25 minutes on the Spark. The question still open on the team's side: how long may an answer take in the intended use? That fixes the budget, and the budget fixes which model is ahead.
This held at every scale we tested: MMLU-Pro and HumanEval at 12k vs 32k, GPQA Diamond (a tie at 80%), and LiveCodeBench, where Qwen3.8 does not finish 38% of the problems even at 32k (79% vs 60%). See Quality.
Also checked, so the comparison holds
- Quantization does not change the result. Muse in NVFP4 and in Q4 GGUF score the same (82.0 vs 81.0 on MMLU-Pro); Muse vs Qwen3.8 both in NVFP4 on the same runtime shows the same gap.
- Same checkpoint, same scores on the Spark and on the A6000. Three Qwen3.8 runs land within ±1.5 points, which is also the noise level to read every table with.
- Speculative decoding (MTP, DFlash) changes speed, not answers. Qwen3.8 31.8 → 50.8 tok/s; Muse 38.5 → 56.5 tok/s. See Speed.
- The A6000 decodes 2.5× faster than the Spark on the same model and recipe; prefill is about equal. The Spark's config is well tuned; it is memory-bandwidth bound.
Not measured: agentic / SWE-bench-style work. Vendor-published SWE-bench numbers are in the repository (swe-bench-cards.md). See Caveats.