KV is the state we keep for future tokens. It is not the entire peak memory of a forward pass. Prefill also needs temporary tensors and kernel workspaces while it processes prompt tokens. Weight storage, cached KV, runtime buffers and allocator reservations all share the device. A spreadsheet that subtracts weights and KV from total VRAM and calls the remainder "free" may undercount a shape-specific temporary peak. vLLM's worker memory profiling measures peak model memory to determine what can be allocated to KV, precisely because KV cannot take the whole remaining device.

Suppose short prompts fit under a KV pool sized for 32K tokens, but a 32K prefill OOMs when several prompts run together. I would measure allocated and reserved memory before, during and after prefill for that batch shape. Use PyTorch's peak-allocation and memory snapshot tools where the stack uses them, and account separately for allocations invisible to that allocator. Profile per worker and rank. The peak might come from attention or MLP intermediates, image encoders, a fallback kernel, graph capture or another concurrent request. Do not announce the cause from nvidia-smi idle memory alone.

Chunked prefill can reduce the number of prompt positions processed simultaneously, trading extra scheduling steps for a smaller temporary peak. Limiting concurrent prefills and reserving a measured workspace budget can help too. But if a kernel switches implementation at a shape threshold, the right fix may be to repair or route around that fallback. Re-run quality checks after a chunking change because positional or mask bugs can make the same prompt produce different answers, as The prompt is identical. Why does chunked prefill change the model's answer? explores.

An interviewer might say FlashAttention has linear memory, so prefill cannot OOM. It avoids materializing the quadratic attention matrix in the usual implementation, but it does not remove all activations or workspace. Also distinguish an allocation failure from actual live memory exhaustion and from a request rejected by an admission limit. The capacity contract needs peak memory for real batch shapes, not only steady-state KV slots. KV memory is allocated in blocks. Why can short requests waste so much of it? addresses wasted slots inside KV blocks. This one addresses memory outside the KV budget during computation.