Paged KV storage avoids reserving a maximum context window for every request. It allocates fixed-size blocks as tokens arrive. That solves a large part of the external fragmentation problem, but the last block of each active sequence may be only partly filled. The PagedAttention paper motivates block-based KV management to reduce waste and enable sharing. It does not make the final partially used block free.

Take a simplified allocator with 16-token blocks. A request with 17 live tokens needs two blocks, with only one token used in the second. If thousands of concurrent requests have lengths just over a block boundary, their tail blocks consume capacity that a spreadsheet counting only actual tokens misses. If a block stores KV for many layers and heads, those unused token slots are not merely a few bytes. The exact allocation depends on the engine's layer grouping and block implementation, so measure real reserved blocks rather than extrapolating from a toy formula.

I would report three numbers separately: logical live KV tokens, allocated block slots and total reserved KV pool. The ratio of slots to live tokens reveals internal waste, while the gap from allocated slots to pool size shows free blocks held for future work. Prefix sharing, beam branches and eviction can complicate ownership, so break the view down by request length and cache state. A high GPU-memory reading alone does not say whether the blocks are live, reusable or stranded.

Could we make blocks smaller? That reduces worst-case tail waste per sequence but raises metadata, block-table and scheduling overhead, and it can change kernel efficiency. Larger blocks can improve some access patterns while hurting workloads dominated by short sessions. Replay the real length distribution at several block sizes, checking throughput, p99 latency and admitted concurrency. Do not size for the average prompt alone, since output length and number of active sequences determine how many partially filled tails coexist.

The useful distinction is external versus internal fragmentation. Paging keeps blocks physically noncontiguous, so a long sequence does not need one contiguous allocation. It can still leave unused slots inside the blocks it owns. The capacity estimate should count physical slots, not just semantic tokens.