Model and Inference Engineering · Staff
Speculative decoding accepts most draft tokens. Why is latency unchanged?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Acceptance is useful only if each verification round commits enough tokens to pay for drafting and verification. In speculative decoding, a smaller model proposes several tokens, then the target model scores them in a parallel verification pass. That can replace several sequential target decode steps. It also adds serial draft steps, a target verification pass, cache management and scheduling work. The original speculative-decoding paper explains the opportunity. It does not guarantee a speedup for every target, draft and serving batch.
Say the draft proposes four tokens and the dashboard says 80 percent of proposed tokens are accepted. That single percentage is not enough. How many tokens were committed per round after the first rejection? How much wall time went to drafting, target verification, sampling, KV updates and waiting for a batch slot? A draft model that takes nearly as long as the target can erase the saved target steps. Verification of four positions need not cost the same as one ordinary decode step under the current kernel and batch shape.
I would compare no-speculation and speculation on the same prompts, output lengths, batch sizes and quality contract. Report committed tokens per target pass, target passes per generated token, decode latency per completed request, throughput and p95. Separate time-to-first-token because the first speculative block can delay the first visible token even when later decode speeds up. The cost picture also changes when many concurrent users make target verification compute-heavy. A production-engine study measures how batch size and verification overhead limit gains, rather than inferring speed from acceptance alone.
If the interviewer says to use a longer draft, I would measure the marginal accepted tokens against the extra serial draft cost. Rejection at position two can make positions three through eight wasted work in a chain-based scheme. Tune draft length and turn speculation off for workload slices where it loses. Speculative decoding wins a benchmark but raises cost per successful task asks whether a benchmark win lowers cost per successful task. This is the inner serving mechanism explaining why a seemingly strong acceptance rate may not even reduce latency.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →