Model and Inference Engineering · Principal
Tensor parallelism doubled, but KV cache per GPU did not halve. Why?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Start with the number of KV heads, not the number of query heads or GPUs. In grouped-query attention, several query heads read one KV head. A serving implementation can partition KV heads across tensor-parallel ranks while there are enough heads to give each rank a whole one. After that, some implementations replicate a KV head across ranks serving different query heads.
Say the model has 8 KV heads, the query-head count permits tensor parallelism of 16, and the server supports this configuration by replication. At TP=8, each rank holds one KV head's cache. At TP=16, each rank still needs one. Relative to the unsharded eight-head cache, a rank uses about one eighth for KV, not one sixteenth. Across all 16 ranks, the aggregate KV storage is twice the logical unsharded KV data, before allocator overhead and other features. This vLLM model implementation explicitly partitions KV heads when their count is at least TP size and replicates them when it is smaller.
It is easy to miss this if a capacity spreadsheet divides the whole cache by TP size. Query projections and output projection work may still be sharded, and per-rank weight memory can fall, so the total GPU memory picture is mixed. But the KV term stops shrinking once this per-rank head floor is reached. For a large active-token budget, that floor can decide concurrency.
I would verify the actual serving kernel's head mapping, cache dtype, block size, number of layers and whether prefix sharing or sliding-window attention changes allocation. Then measure reserved and live cache blocks per rank with a fixed set of requests at TP=8 and TP=16. Compare request throughput and interconnect cost too. TP=16 may be justified by compute or model fit even when it does not buy KV capacity per rank.
The important qualification is that this is an implementation choice, not a universal promise that any GQA model runs at any TP degree. Some runtimes reject a configuration they cannot partition or replicate correctly. If the interviewer asks whether an extra GPU always increases the token budget, the answer is no. Work through the physical head placement first.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →