Caveats

What these numbers can and cannot tell you

Read differences against the noise

Samples of 200 questions carry a margin of error of ±5–6 points on MMLU-Pro and GPQA, ±2–4 on HumanEval. Measured directly: the same model run three times spread by ±1.5 points. A difference of 2 points between models means nothing; 9 points with a paired count of 22 vs 4 does.

Our prompts, our parsers

Scores compare these models with each other under identical conditions. They are not comparable to public leaderboards, which use other prompt formats, few-shot examples and larger budgets. Where vendor numbers exist for the same task at full precision, ours land close (GPQA Diamond: Muse 80.3 here vs 83.5 published) once the budget is accounted for.

4-bit weights, dynamic quantization

Everything ran quantized to fit 48 GB. The formats are not bit-for-bit comparable across models (unsloth's Q4 is dynamic per layer; NVIDIA's NVFP4 for Muse is a searched mix; the Qwen3.8 checkpoint is NVFP4 only in its MLPs, FP8 elsewhere). The check that matters was made: Muse scores the same in both formats, and the Muse–Qwen3.8 gap is the same with both in NVFP4.

Not measured

Muse's own numbers vs the vendor's table

Meta's model card compares Muse with Qwen3.6-27B on its own agentic benchmarks, where it reports Muse ahead on most. None of those were run here; this study used independent public datasets on purpose. Both sets of numbers are self-consistent; they measure different things.