Summary

Same quality if you can afford the wait; Muse if you can't

Muse-Glimmer-30B, Qwen3.6-27B and Qwen3.8-27B, run locally on the same questions, prompts, sampling and token budgets. Every answer was saved and every wrong answer inspected. The three findings below carry the decision; each links to the page that holds the evidence.

1. Given room to finish thinking, the three models are equivalent

MMLU-Pro with a 32k-token budget: 82 / 82 / 80%.

200 MMLU-Pro questions, 32,768-token completion budget. Muse ran at 12k (it hit the limit on only 3 questions). Margin of error ±5–6 points; differences this size are noise. Quality page has every task.

2. Muse gets there with less than half the tokens

Tokens per answer on the same questions: 1,592 vs 4,083 vs 4,935.

Mean completion tokens (reasoning + answer) on the run above. Time per question follows the same ratio: 50 s, 176 s, 98 s with 3 requests in flight.

3. With a tight budget, Muse wins because it finishes

MMLU-Pro at 12k tokens: Muse 82%, Qwen3.6 73%, Qwen3.8 77%. The gap is answers that were never finished.

Each bar is 200 questions. "Not finished" = the model was still reasoning when it hit the 12,288-token limit; it counts as wrong. Among the questions each model did finish, accuracy is 83% / 81% / 86%.

What decides it

The token budget is the whole story, and it is a product decision, not a benchmark one. A 32k-token answer takes about 11 minutes on the A6000 with speculative decoding and over 25 minutes on the Spark. The question still open on the team's side: how long may an answer take in the intended use? That fixes the budget, and the budget fixes which model is ahead.

This held at every scale we tested: MMLU-Pro and HumanEval at 12k vs 32k, GPQA Diamond (a tie at 80%), and LiveCodeBench, where Qwen3.8 does not finish 38% of the problems even at 32k (79% vs 60%). See Quality.

Also checked, so the comparison holds

Not measured: agentic / SWE-bench-style work. Vendor-published SWE-bench numbers are in the repository (swe-bench-cards.md). See Caveats.