MTP speculative decoding
Local LLM stackNote
What MTP is
M1 / Propose → verify → keep
One check.
Several useful tokens.
The cat sat
Proposes the next tokens
Replaces “rug” with “mat”.
1 → 3
pass → tokens in this example
1 → 1
without MTP
--parallel 1. Drafts are checked, not blindly accepted. In the final state, “mat” replaces the rejected “rug”.Multi-token prediction (MTP) is a small extra head trained with the model. It guesses the next few tokens. The main model checks all guesses in one pass and keeps the ones it agrees with. When the guesses are right, you get several tokens for the price of one pass.
Qwen3.8-27B has the head inside the GGUF. You do not need a second draft model.
Measured gain
| Server | Without MTP | With MTP | Gain |
|---|---|---|---|
| llama-swap (upstream-based build) | 28.8 t/s | 51+ t/s | +78 % |
| Unsloth Studio, 4k deep | 27.2 t/s | 35.0 t/s | about 1.3x |
| Unsloth Studio, 32k deep | 24.4 t/s | 24.3 t/s | none |
MTP costs about 1.3 GiB of VRAM. It was never harmful, so it stays on.
The first single-GPU run saw 56 to 87 % of drafts accepted.
The flag name trap
An early note said "MTP does not work on this build" because llama-server --help had no mtp flag. It was wrong. The flag is --spec-type draft-mtp, part of the --spec-draft-* family. That mistake cost a full investigation. Search the help text for the family, not the acronym.
Settings
| Flag | Value | Why |
|---|---|---|
--spec-type | draft-mtp | Use the built-in head |
--spec-draft-n-max | 3 main, 2 uncensored twin, 2 on Studio | Upstream calls 2 the sweet spot for 16 to 24 GB cards and 3 to 4 for bigger ones. This box sits between |
--spec-draft-p-min | 0 | See below |
--parallel | 1 | Required. More slots are unsupported with MTP, and the gain is gone by 4 |
Why p-min 0 beats the common advice
Many guides say to gate drafts at p-min 0.75. On a smaller MTP model on this box (900-token generations, median of 3):
| n-max | p-min | Decode | Acceptance |
|---|---|---|---|
| 3 | 0.75 | 116.5 t/s | 93 % |
| 3 | 0 | 144.2 t/s | 58 % |
| 6 | 0.75 | 116.3 t/s | 93 % |
| 6 | 0 | 129.1 t/s | 39 % |
High acceptance is not the goal. The gate skips drafting when the head is unsure, and throws away a free token. When the model is fully on the GPU, a rejected draft costs almost nothing. Upstream suggests 0.60 to 0.75 for bandwidth-starved rigs and reports it hurts fast cards. For Qwen3.8 on this box, 0.60 to 0.75 is not tested.
MoE models want shorter drafts
On an offloaded MoE model, n-max 3 was slower than 2 (111.9 against 116.5 t/s). Every drafted token can need different experts, and those come over from system RAM. This does not apply to Qwen3.8-27B, which is dense in its weights, but matters if you run MoE models on the same server. Doctor-Shotgun's MoE offload guide explains the offload side.
Lossless, but not bit-identical
On a short greedy prompt, output with and without MTP was byte-identical. On long greedy runs they diverged. MTP checks several tokens per pass, which changes the GPU's reduction order and the last bits of the logits. On a near tie, one token flips and the rest of the text forks. The math is lossless, the output is not deterministic.
Measure MTP with long outputs
A benchmark capped at 160 output tokens made MTP look worthless. The crossover was around 900 tokens on the smaller model. Use long generations.
Not everything keeps the head
Some quants strip the MTP tensors. On a Qwen3.6-35B UD-Q4_K_M quant, --spec-type draft-mtp failed with failed to create MTP context at any context size. A quant of the same model that kept MTP ran 36 % faster. If MTP fails, check the GGUF before you blame VRAM. On Unsloth Studio, check the fit settings first (see Unsloth Studio).
Beyond MTP: DFlash2 (not tested here)
Not testedllama.cpp PR #27342 adds DFlash2, a separate draft model with a candidate selector:
cd ~
./build/bin/llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
-hfd incoai/Qwen3.8-27B-DFlash2-GGUF:Q4_K_M --spec-type draft-dflash --spec-draft-n-max 7
Claims from a YouTube video and its comments, not verified:
- DFlash2 alone: 68 to 154 t/s on one RTX PRO 6000 Blackwell across 100 LiveCodeBench problems. With n-gram drafting added: 305 t/s. The same server gave 112 or 268 t/s depending on the prompt.
- One commenter stopped the draft chain early when the top two logits were close. They reported +31 % on code and +61 % on prose.
Next step: Benchmarking honestly
- Leads to
- Benchmarking honestly