Ask what those extra predictions represent. A multi-token prediction head can propose tokens beyond the immediate next one, but that does not make each proposal an exact sample from the base model's next-token distribution after the earlier proposed tokens. The base model may choose a different first token. Then the later draft tokens were made under a continuation that never happened. DeepSeek-V3's technical report describes a multi-token prediction objective, while vLLM's MTP serving documentation treats the head's tokens as drafts verified by the target model.

For a simple example, the draft proposes A B C D. If the target rejects B, the server cannot continue by emitting C D as though the sequence were still A B C D. It must follow the target's accepted or corrected continuation and rebuild subsequent speculation from that state. Even if all four are accepted under the verification algorithm, the precise sampling rule matters. Greedy matching and distribution-preserving speculative sampling have different acceptance machinery. Do not promise exact target-model output by merely taking the argmax of four heads.

I would test target-only decoding against MTP-assisted decoding using fixed weights, tokenization, sampling settings and prompts. For greedy mode, compare output tokens across varied rejection positions. For sampling mode, check the implemented acceptance and correction rule against the intended target distribution, then measure statistical behavior and quality. Inspect whether rejected drafts polluted KV state or token positions, and measure accepted tokens per verification pass against the extra head's latency.

The interviewer might say the MTP head is part of the same checkpoint, so it is "the same model." That tells us how it was trained, not that its proposal distribution equals the base decoder's conditional distribution at every later position. MTP can improve the proposal quality and reduce external draft-model cost. Verification is still what makes this a correctness-preserving serving path under the documented algorithm. A speculative decoder gets faster by sampling from the wrong model covers using the wrong target distribution in speculative sampling. This question asks why a future-token training objective does not itself authorize four committed output tokens.