I would stop routing that class through graph replay while reproducing the mismatch. Wrong tokens across requests are a correctness incident and potentially a data exposure. The tempting explanation is “CUDA graphs are nondeterministic.” I would want evidence before blaming arithmetic. A graph replay runs a captured sequence of work against specific memory locations and the graph's chosen shape. If a request ID, token ID, position, sequence length, KV block table or sampling input is stale in one of those locations, the kernels can run perfectly and still compute for the wrong request.

NVIDIA's CUDA graph guidance describes the static-address requirement and the danger of rebinding a tensor instead of copying new data into persistent graph inputs. In a serving engine, there are more moving parts than a plain tensor batch. The scheduler may reuse slots as requests finish, and the captured decode step may read metadata that identifies which KV pages belong to which sequence. I would inspect every input that crosses the capture boundary, including pointer arrays and device-side scalar tensors, not just the visible input_ids tensor.

My smallest reproduction would use two unmistakably different prompts, deterministic sampling and a fixed batch shape that alternates slot membership. Run eager and graph modes from the same cache state and compare logits before sampling, then compare token IDs. If logits diverge, check the actual input and metadata bytes immediately before replay and after synchronization. If logits agree but sampled tokens differ, inspect sampling seed and per-request RNG state, output slot mapping and any CPU-side postprocessing. Make the failing shape repeat at low concurrency first, then add cancellation, preemption, prefix cache hits and request replacement one at a time.

The order of operations is part of the bug search. The host must copy each new input into the persistent buffers, ensure those copies complete on the relevant stream before replay reads them, and not reuse output storage until the consumer has finished. A graph key should represent every shape and mode that changes captured work or memory interpretation. If the scheduler pads a smaller batch into a captured bucket, masks and valid lengths must say which slots are real. A stale KV block table can mix contexts even with a correct token ID. Look for an alias to a freed or recycled allocation and for a metadata update made after replay was queued.

I would add a canary that compares graph and eager logits on representative transitions, including a batch that loses one request and gains another. It need not run on every production token. Keep per-request traces of graph key, slot generation, KV ownership and output mapping with privacy-safe IDs, so a future mismatch is diagnosable. The release gate is exact behavior for deterministic fixtures, no cross-request content, and acceptable differences only where floating-point execution order reasonably explains small logit variation. A speedup is real only if it preserves the request boundary. A generic eager fallback is a valid temporary outcome.