The two branches are allowed to share the KV blocks for tokens they have already consumed in common. Those blocks are immutable as far as the branches are concerned. The first time they need different KV state, they need independent writable storage. If both block tables still point to the same writable physical block, a write for branch A can alter what branch B reads on its next attention step. The output then depends on the order in which the scheduler ran them. vLLM's PagedAttention design uses reference counts and copy-on-write to share blocks safely for parallel sampling and beam search.

A block holds several tokens, so a fork can happen in the middle of its last block. It is not enough to allocate a new block only after the old one fills. If the next token writes into a shared, partially filled block, make a private copy first. Earlier complete blocks can still be shared. Each logical beam needs its own block table, sequence length, position and beam score. When beams are pruned or reordered, update the owners and references without freeing a block still reachable by another beam. This is exactly where a seemingly harmless scheduler optimization can break correctness.

I would reproduce with a tiny fixed prompt that branches at a known token and forces different continuations. Compare each branch's logits and KV reads to a run that allocates fully separate caches. Then fork at every possible offset within a block, reorder beams, prune one, cancel another, and interleave requests from different customers. Validate block reference counts and copy events under all those transitions. Same outputs in a single fixed scheduling order are weak evidence because the bug may require the other order.

Could we copy the entire prefix for each branch? Yes, and that is a useful correctness baseline. But beam width and long prefixes make that costly. Copy-on-write preserves sharing until a real write needs isolation. Two requests share a cached prefix. Why did cancelling one corrupt the other? is about cancellation corrupting a shared prefix cache across requests. This page is about ownership at the moment a decode path forks within one search request. Both need an explicit answer to who may write each physical KV block.