A captured CUDA graph replays operations against stable memory addresses. Its input and output buffers and the allocations tied to graph capture cannot simply be returned after each request like ordinary temporary tensors. PyTorch's CUDA graph documentation explains the long-lived addresses and private graph pools. vLLM's memory guidance also notes that CUDA graphs take extra GPU memory. So idle requests do not imply idle allocations. The process is keeping captured work ready for a future decode step.

The amount depends on the engine, model, graph mode, captured batch or sequence shapes, and whether captures share a memory pool. A serving engine may prepare several shapes so common decode batches avoid Python and kernel-launch overhead. More graph variants can mean more persistent memory. I would record allocated and reserved memory before model load, after weights, after graph capture, and after representative traffic. Break out KV cache, graph pools and any compiled-kernel workspace. GPU process memory or nvidia-smi alone does not identify which owner is responsible.

Suppose an operator sizes admission using free memory before graph capture. The first request later triggers capture and the engine has less room for KV than the plan assumed. Reverse the order: capture required shapes during startup, then derive a safe KV budget from the post-capture steady state. If graph memory is too large, reduce captured shape variants, use a more selective graph mode, share pools where the implementation supports it, or fall back to eager execution for rare shapes. Each change trades some latency for capacity. Test under actual batch composition because a graph configuration that helps batch one may not help all traffic.

The harder follow-up is why disabling graphs appears to “fix” OOM but worsens p99. Both results can be true. The correct comparison is completed throughput and first/inter-token latency at a fixed memory and load budget, with startup capture included in rollout capacity. CUDA graph replay is faster. Why does it return the wrong tokens for some batches? is about a graph replay producing wrong tokens for some batches. This page is about persistent memory ownership and how that changes admission even when replay is correct.