First, distinguish memory reserved by an allocator from KV blocks still assigned to active or orphaned requests. An allocator may hold freed GPU pages for reuse, so nvidia-smi remaining high does not by itself prove a leak. PyTorch's CUDA memory documentation distinguishes allocated tensor memory from cached reserved memory. A serving engine may use its own KV pool on top of that, so I also want engine-level counts: free blocks, blocks held by active requests, blocks held as reusable prefixes, and blocks awaiting release after in-flight GPU work.

A gateway timeout is a decision by the gateway to stop waiting. It may not have sent a cancellation to the worker, or the worker may still be generating into a closed stream. A request can also be removed from the scheduler but leave shared prefix references or pending decode work. I would trace one request ID through admission, block allocation, gateway deadline, worker cancellation receipt, final GPU completion and block release. The reference count and ownership transitions matter more than a fleet-wide utilization average. vLLM's prefix-cache design describes blocks with request references and eviction after references reach zero. That is a useful model, not evidence that this hypothetical engine has that exact bug.

The contract should make cancellation an explicit state transition. The gateway sends a stable request ID and cancel message, retries cancellation if its acknowledgement is lost, and does not interpret a transport disconnect as proof of release. The worker marks the sequence non-runnable, stops scheduling new tokens, then releases its unique KV blocks only after kernels that might read them have completed. Shared prefix blocks lose this request's reference but remain valid for other owners or the cache policy. Use a generation or ownership token so a late completion from an old request cannot free a block now assigned to a new one. Log final release once, even if timeout, client disconnect and worker error all race to end the same request.

To reproduce, send a long generation, time it out at the gateway, then compare the worker's request table and KV block accounting before and after the cancellation acknowledgement and stream completion. Repeat with a shared prefix, client close, worker error and a cancellation while a GPU step is running. If active ownership falls to zero but the allocator's reserved bytes stay flat and new requests reuse capacity, there may be no leak at all. If free KV blocks keep shrinking while no requests or intentional cache entries own them, we have an accounting defect. If the worker never hears cancel, the fault is upstream of the allocator.

I would protect capacity during the fix with a worker-side deadline or orphan reap that checks actual request state, not just age. A blind timer freeing blocks while GPU work is in flight can create cross-request corruption. The release gate is stable free-block accounting after every terminal path and no wrong-token or stale-prefix exposure under cancellation races. KV cache is nearly full and p99 doubled. What would you investigate? is the diagnosis of ordinary KV pressure under live load. This is the lifecycle of resources after the caller believes the request ended.