What these numbers can and cannot tell you
Read differences against the noise
Samples of 200 questions carry a margin of error of ±5–6 points on MMLU-Pro and GPQA, ±2–4 on HumanEval. Measured directly: the same model run three times spread by ±1.5 points. A difference of 2 points between models means nothing; 9 points with a paired count of 22 vs 4 does.
Our prompts, our parsers
Scores compare these models with each other under identical conditions. They are not comparable to public leaderboards, which use other prompt formats, few-shot examples and larger budgets. Where vendor numbers exist for the same task at full precision, ours land close (GPQA Diamond: Muse 80.3 here vs 83.5 published) once the budget is accounted for.
4-bit weights, dynamic quantization
Everything ran quantized to fit 48 GB. The formats are not bit-for-bit comparable across models (unsloth's Q4 is dynamic per layer; NVIDIA's NVFP4 for Muse is a searched mix; the Qwen3.8 checkpoint is NVFP4 only in its MLPs, FP8 elsewhere). The check that matters was made: Muse scores the same in both formats, and the Muse–Qwen3.8 gap is the same with both in NVFP4.
Not measured
- Agentic work (SWE-bench-style: read a repository, run tools, iterate). LiveCodeBench is single-shot competitive programming, a different skill. Vendor-published SWE-bench numbers are in
swe-bench-cards.md; the two models are within a few points there, by the vendors' own measurement. - Budgets above 32k. Qwen3.8 finishes 96% of what it finishes correctly; a larger cap would raise its LiveCodeBench score at a proportional cost in time. The right cap depends on the tolerated answer latency, which is the open question.
- Muse on the Spark. It is not installed there and the team chose not to install it. Everything measured on the A6000 transfers for quality (same checkpoint gives the same scores on both machines), not for speed.
- Multi-user throughput. Speed numbers are for one request at a time; the harness can measure concurrency but this was not asked.
Muse's own numbers vs the vendor's table
Meta's model card compares Muse with Qwen3.6-27B on its own agentic benchmarks, where it reports Muse ahead on most. None of those were run here; this study used independent public datasets on purpose. Both sets of numbers are self-consistent; they measure different things.