Model and Inference Engineering · Staff
GPU utilization improved. Why is one long answer repeatedly starting over?
The question
Interview question
After raising concurrent requests, the GPU stays busy and aggregate tokens per second rises. A long response sometimes takes three times as long. The serving engine logs repeated preemptions of its sequence when KV space runs out. Is this just the unavoidable price of higher utilization?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The cache contains keys and values for tokens already processed. If a scheduler preempts a sequence and discards its cache, it can later recompute that prefix from the prompt and generated tokens before continuing. The user does not necessarily see the model repeat text, but the GPU repeats work. vLLM's optimization documentation describes recompute preemption under KV pressure and its latency cost. High utilization by itself cannot tell productive decoding from replaying the same context.
I would trace a single request: prompt tokens, generated tokens, KV blocks reserved, preemption count, recomputed tokens, queue wait between resumes and final completion time. Then compare the sum of new output tokens with total tokens processed by the GPU. A request with a long accumulated prefix can make each successive restart more expensive. Measure the tail over request cohorts, not just successful short requests. Include cancellations and timeouts because they may disappear from a completed-request latency chart.
Raising the concurrency limit can admit more sequences than the cache can keep resident. Reducing maximum active sequences, reserving KV capacity for already running long jobs, or using a different preemption or swap policy may help, but each costs something. Swapping cache to host can avoid recomputation and still lose on transfer latency and bandwidth. A larger GPU or lower-precision KV may free space, with cost or quality consequences. I would run a traffic replay with the real prompt and output-length distributions, and set an admission policy based on predicted memory with headroom for unknown continuation length.
What if throughput still wins and only a few long jobs suffer? That may be a product decision, but make the affected class and its SLO explicit. KV cache is nearly full and p99 doubled. What would you investigate? investigates KV saturation and p99 generally. Swap KV to host memory or recompute it? compares swapping with recomputation. Here the incident is a repeated-preemption feedback loop hidden by busy GPU charts, and the interviewer should ask how much of that busy time served new user tokens.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →