Reusing a prefix does not make the cached blocks private to the request that computed them first. Both live sequences may reference the same immutable prefix blocks. Releasing A can drop A's reference, but the allocator must not overwrite a block while B still uses it. vLLM's prefix-caching design discusses cached KV blocks and reuse. An implementation using reference counts must only make a block available for eviction or reuse when no active request still owns it. The exact lifecycle differs by serving engine, so inspect its allocator rather than projecting a particular data structure onto every engine.

I would build a deterministic two-request reproduction. Start B after A populates the prefix, verify both point at the expected shared block IDs or hashes, cancel A at several moments, then watch block ownership, free-list movement, eviction and B's next logits. Add a third request that forces allocation pressure. Without pressure, a prematurely freed block may remain unchanged and the test passes by luck. Compare B against an uncached reference with the same weights and token IDs. Separate an actual use-after-free from a bad cache key that reused a different prompt or adapter state.

The invariant needs to include more than a count. Shared prefix blocks should be immutable for all active consumers. Cache identity must include whatever changes the computed KV, such as model revision, adapter, tokenization, positions and relevant multimodal inputs. Cancellation and timeout must release each request's ownership once, including paths that race with preemption or normal completion. A stress test should interleave acquire, cancel, evict and allocate while validating allocator invariants and output parity. Refcount underflow and double release deserve explicit assertions.

What if B uses a copy-on-write continuation block? Fine. The shared completed prefix still needs safe lifetime management, and B's private suffix should not mutate it. Requests time out, but GPU memory stays high. Who still owns the KV blocks? asks who owns KV blocks after requests time out, with memory that stays high. This is the opposite failure: a block is freed too early while another live sequence still needs it. Cache hit rate can rise as correctness falls, so measure both.