Decode speed is tokens generated per second for one request on an otherwise idle server; prefill is how fast the prompt is read before the first token. Both measured by the same script on every server, with unique prompts so caches cannot inflate them.
Speculative decoding speeds up without changing answers
A6000, single request. A small "drafter" proposes tokens and the model verifies them in one pass; the output distribution is unchanged, and the scores confirm it: Qwen3.8 with vs without MTP scored 97.5 vs 98.0 (GSM8K), 77.0 vs 74.5 (MMLU-Pro), 97.0 vs 94.5 (HumanEval), all within noise. MTP is Qwen's built-in head; DFlash is Meta's separate drafter for Muse. On vLLM this A6000 could not make DFlash pay off for Muse (25% draft acceptance); the number above is llama.cpp.
The A6000 decodes 2.5× faster than the Spark; prefill is about equal
Same checkpoint, same vLLM recipe, MTP on both: 50.8 vs 20.3 tok/s.
Decode, A6000
tok/s, Qwen3.8 NVFP4 + MTP
Decode, DGX Spark
tok/s, same model and recipe
Time to first token, ~10k-token system prompt
both machines, roughly
Prefill throughput by prompt length. The Spark is faster on short prompts and slower on long ones; for a 40,000-character system prompt (~10k tokens) both start answering in 6–9 s. The Spark's config was verified: native FP4 kernels, MTP accepting 53% of drafted tokens (~2.6 tokens per pass). Its limit is memory bandwidth: only the MLPs of this checkpoint are NVFP4; all attention and the lm_head are FP8, so each step reads ~25 GB.
The prefix cache has no effect on the Spark's vLLM; it works on 0.29
Repeating a 20k-token prompt: 16.7 s again on the Spark; 1.7 s on vLLM 0.29 with MTP on.
Time of the second request that shares its prefix with the first (bars), with the first request's time as reference (hollow marks). Same prompt twice, max_tokens=1, idle server. On the Spark (vLLM 0.26.1) no scenario hits, from 6k up to a 129k-token repeat (164.7 s vs 165.2 s). On vLLM 0.29 the cache works with MTP on (2× faster repeats; 1,600-token cache blocks) and better with it off (4×; 224-token blocks). A vLLM upgrade should restore caching on the Spark with MTP kept on.
All decode and prefill numbers
Configuration
Decode tok/s
TTFT
Prefill ~1k
~4k
~12k
Compare speed only within the same runtime (llama.cpp vs vLLM differ in prefill), or the same model and runtime across machines. Muse NVFP4 on vLLM carries two Ampere workarounds that cost it speed.