Model and Inference Engineering · Staff
We disabled half the attention heads. Why didn't inference become faster?
The question
Interview question
A team masks half the attention heads, but inference latency does not improve. What work does the engine still perform, and what would structural pruning change?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
What does “disabled” mean in the implementation? If we multiply half the head outputs by zero after attention, the GPU may still project Q, K and V for those heads, run their attention kernel, write their KV cache and run the same output projection. The result changes but the dense tensor shapes do not. A dense kernel generally computes zeroed weights too unless a sparse or structurally smaller operation is actually selected. PyTorch's pruning tutorial makes the same basic point for masked dense weights: zeros alone do not remove dense matrix-multiply work.
Structural pruning is different. Remove the chosen head slices from Q/K/V projections and the corresponding input columns of the output projection, update attention head counts and layouts, and use kernels that support the resulting shapes. For grouped-query attention, a query head may share a KV head with others, so dropping query heads does not automatically shrink KV cache in the same proportion. If the number of KV heads stays fixed, measure that separately. A model may also need fine-tuning after pruning to recover quality. Don't assume all heads are interchangeable because one eval set did not notice their removal.
I would profile prefill and decode separately before promising a speedup. Compare actual GEMM dimensions, attention kernel shapes, allocated KV bytes, memory bandwidth, launch count and tokens per second at relevant batches. Halving head count does not halve the whole Transformer, which still has FFN layers, norms, embedding and output work. Smaller matrices can even use the hardware less efficiently. NVIDIA's fused attention documentation shows that attention kernels distinguish query and KV head dimensions and supported layouts.
The clean test is a truly resized checkpoint and engine, alongside the masked baseline, with quality and latency measured at the same load. If they have identical kernel shapes, we have not actually bought a structural serving saving. A zero in the computation graph is not the same thing as absent computation.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →