Model and Inference Engineering · Principal
The verifier rejected a draft token. Why does the next decode still remember it?
The question
Interview question
A draft model proposes four tokens. The target accepts the first two and rejects the third. The server emits only the accepted prefix and the corrected next token. Yet subsequent logits differ from ordinary target-model decoding. The output token list looks clean. What else might still contain the rejected continuation?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The draft and target may have computed KV for speculative positions that never became committed tokens. Cache allocation is not the same as acceptance. The next target step must attend to the committed prefix, with positional indices and cache lengths aligned to that prefix. It must not see K/V from rejected positions. The original speculative decoding paper establishes sampling from the target distribution under its acceptance procedure. That mathematical result assumes the implementation carries forward only the appropriate state. vLLM's KV cache manager discusses lookahead allocation for speculative tokens. This incident is a hypothetical bug, not a claim that vLLM retains rejected KV.
I would record, per verification step, the proposed tokens, accepted count, replacement or bonus token, committed token count, target computed-token count, draft computed-token count and allocated block ownership. Stop at the first rejection and inspect both models' subsequent state. Some engines can retain a physical block while marking only part of it valid. That is fine if attention cannot read the invalid slots and a later append overwrites them correctly. A length counter advanced too far, a reused partial block, or a block hash published before commit can turn speculative work into apparent history.
The clean test is a forced early rejection on a short fixed prefix. Run one target decode without speculation and one with it. Compare the target's next-step logits when both have the same committed token sequence, model revision and sampling configuration. For stochastic sampling, compare distributions or seeded conditional behavior rather than demand byte-for-byte text identity from independent random draws. Repeat with rejection at each draft position, at a block boundary, after preemption and under prefix-cache reuse. A pass on all-accepted drafts says almost nothing about rollback.
The interviewer may say it is cheaper to keep the extra KV and mask it. That can work if the validity boundary and later append semantics are proven for every layer and worker. Otherwise trim or rebuild the affected tail. The optimization must preserve the target-model conditioning state. A speculative decoder gets faster by sampling from the wrong model is about sampling from the wrong probability distribution even when cache state is correct. This is a state-consistency failure after the sampling decision itself was made correctly.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →