Look for one Boolean standing in for a vector of request states. A line equivalent to finished = (next_tokens == eos_id).any() tells the loop that someone finished. It must not mean everyone finished. Track termination per sequence using that request's EOS tokens, length cap, stop rules and cancellation state. Transformers' stopping-criteria interface returns a Boolean per sequence in a batch, which is the shape a custom server needs to preserve.

Once a request finishes, stop extending its KV and stop billing or streaming generated tokens for it. Keep the other requests alive. A static-shape kernel may still have a padded slot for the finished sequence, but that slot is a serving implementation detail, not an extra user-visible token. When the scheduler compacts the batch, carry the mapping from batch row to request ID. Otherwise the next token may go to the wrong stream even after the stop bug is fixed.

I would test three requests: one ends immediately, one reaches a per-request length limit, and one keeps generating. Repeat with rows reordered, cancellation between decode steps and a model with several valid EOS IDs if that is part of its tokenizer contract. Check returned text, terminal reason, KV release and usage for each request. A global all() is not automatically the answer either. It keeps the kernel loop running but can still append junk to a request that already ended if the per-row mask is missing.

This is why a single-prompt parity test does not validate a batching rewrite. The failure lives in the relationship between a row of logits and the request lifecycle that row belongs to.