The cache estimate is a good first step. At each attention layer, one token stores a key and a value for each KV head. Roughly, KV bytes are 2 × layers × live tokens × KV heads × head dimension × bytes per element. Holding the other terms fixed, eight KV heads instead of 32 gives one quarter as many cached KV elements. It does not make weights, activations, temporary buffers or the whole GPU footprint one quarter as large. For a hypothetical 32 layer model with head dimension 128 and two byte KV elements, eight KV heads cost 128 KiB per live token. The same calculation with 32 KV heads is 512 KiB. That is a cache comparison, not a complete latency model.

The 32 query heads did not disappear. In grouped query attention, four query heads share one set of key and value heads. Each query head still forms its own attention scores and weighted output. For a long prompt, prefill computes many query versus prior token interactions across those query heads. It also runs the feed forward layers and writes the KV that decode will later reuse. An optimized attention kernel need not materialize the full quadratic score matrix in memory, but there is still substantial attention arithmetic as the prompt gets longer. The original GQA paper motivates the trade-off between fewer KV heads and model quality. NVIDIA's attention API shows the separate query and KV head dimensions.

Why is decode different? When generating one new token, the engine repeatedly reads a growing history of KV. Reducing that history can relieve memory capacity and bandwidth, which may allow larger batches or longer active sequences. Prefill processes a whole prompt in parallel. Its bottleneck may instead be attention computation, matrix multiplies outside attention, a shape dependent kernel fallback, scheduling delay, or copying the prompt through the stack. GQA can help parts of prefill too, including smaller K and V projections and writes. It simply cannot justify a universal four times speed claim. The actual benefit depends on sequence length, batch size, hardware and kernels.

I would separate time waiting in the queue, prefill GPU time, any prefix reuse, and time to first token. Then profile representative short and long prompts at the same quality target, batch mix and decoding length. Compare attention kernel time, feed forward time, achieved bandwidth, compute utilization, KV allocation, active batch size and tokens per second after the first token. If prefix caching is on, measure computed prompt tokens rather than reported prompt length. A cache hit can hide prefill work in one test and make the comparison misleading.

Suppose the interviewer changes the workload to thousands of short prompts with long generations. Now I would expect the KV reduction to matter more for capacity and repeated decode reads, although the scheduler may turn that headroom into concurrency rather than one request becoming four times faster. If the long prompt is 64,000 tokens and first token time dominates, I would investigate chunked prefill, parallelism and the actual kernel before buying more GPUs on the basis of cache bytes alone.