Model and Inference Engineering · Principal
QPS is flat. Why did the inference autoscaler run out of GPUs?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
QPS counts requests, not the work inside them. One short prompt with a two-line answer and one long context with a thousand-token answer both count as one. Prefill processes the input, decode emits output tokens one step at a time, and active sequences occupy KV cache. NVIDIA's inference metrics separate time to first token, inter-token latency, output token throughput and requests per second for this reason.
For a quick sanity check, 100 requests per second producing 100 output tokens each imply 10,000 output tokens per second in steady state. At the same request rate, 1,000 output tokens each imply 100,000. That is ten times the output-token demand, before we even count a shift to longer prompts. It does not mean you need exactly ten times as many GPUs. Batching, memory bandwidth, attention at longer contexts and the prefill/decode mix change actual capacity. Longer completions also stay resident longer, so concurrent KV occupancy rises even when arrivals per second stay flat.
I would slice traffic by input tokens, generated tokens, context length, model, tenant and priority. Put queue wait, time to first token, inter-token latency, active sequences, KV occupancy, prefill and decode throughput next to the QPS chart. vLLM's metrics documentation gives useful examples of serving-side observability. Check whether the autoscaler reacts to completed QPS, which can even fall under saturation, and whether GPU provisioning takes longer than the workload spike. Kubernetes HPA is a periodic feedback controller, so metric lag and warm-up matter too.
Capacity planning should use measured throughput and latency curves for representative prompt and output distributions. Scale ahead of predictable rises, reserve headroom for warm GPUs, and set admission and per-tenant budgets for requests that can consume most of the KV pool. Do not promise a precise number of GPUs from token arithmetic alone. The important correction is that request rate is one dimension of demand, while serving cost and residency vary by request.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →