Model and Inference Engineering · Staff
Decode uses little GPU compute. Why doesn't that mean the GPU is underloaded?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
At batch size one, each decode step produces only one new token. Its matrix multiplies have much less work to amortize the reading of model weights than a long prefill does. The GPU can wait on memory movement or small-kernel launch and synchronization while its arithmetic units are mostly idle. So low compute utilization alone does not tell me spare throughput exists for this workload. The POD-Attention paper describes the prefill and decode phases as having different compute and memory behavior. NVIDIA's inference sizing guidance makes the same basic distinction.
As a rough bound, an 8-billion-parameter model with two-byte weights has about 16 GB of weights. If each single-token step had to stream all of them from HBM and the device sustained 2 TB/s, the weight-read portion alone would take about 8 ms. That is a simplified lower bound, not a measured tokens-per-second prediction. Some weights can be reused from cache, kernels overlap work, KV reads grow with context, and actual sustained bandwidth differs from the headline number. Communication and launch overhead can also dominate. The purpose of the arithmetic is to ask what resource might cap progress before trying to raise SM occupancy blindly.
I would profile prefill and decode separately. Look at achieved memory bandwidth, cache behavior, kernel launch gaps, active batch size, GEMM shapes, KV reads, synchronization and any tensor-parallel all-reduce. If weight movement is dominant, continuous batching can reuse a loaded weight tile across more sequences, improving aggregate tokens per second. That may raise per-request latency or memory pressure. If bandwidth is not saturated, investigate serial dependencies, small kernels or scheduling overhead before calling it a bandwidth ceiling.
The interviewer might ask whether a GPU at 20 percent compute utilization can take five times the traffic. Only a load experiment with the real sequence-length mix and latency SLO can answer that. Saturating another resource can make p99 worse long before arithmetic reaches 100 percent. A tokens per second benchmark looks great. What did it measure? challenges a tokens-per-second benchmark's measurement. This page asks which hardware resource limits single-token decode and which intervention can change that limit.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →