Model and Inference Engineering · Staff
One sequence emitted EOS. Why did every request in the batch stop?
The question
Interview question
A custom decode loop serves several requests in one GPU batch. A short answer emits the end-of-sequence token, and the server ends all the requests. The model weights are fine when each prompt runs alone. Where would you inspect the batch code?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Look for one Boolean standing in for a vector of request states. A line equivalent to finished = (next_tokens == eos_id).any() tells the loop that someone finished. It must not mean everyone finished. Track termination per sequence using that request's EOS tokens, length cap, stop rules and cancellation state. Transformers' stopping-criteria interface returns a Boolean per sequence in a batch, which is the shape a custom server needs to preserve.
Once a request finishes, stop extending its KV and stop billing or streaming generated tokens for it. Keep the other requests alive. A static-shape kernel may still have a padded slot for the finished sequence, but that slot is a serving implementation detail, not an extra user-visible token. When the scheduler compacts the batch, carry the mapping from batch row to request ID. Otherwise the next token may go to the wrong stream even after the stop bug is fixed.
I would test three requests: one ends immediately, one reaches a per-request length limit, and one keeps generating. Repeat with rows reordered, cancellation between decode steps and a model with several valid EOS IDs if that is part of its tokenizer contract. Check returned text, terminal reason, KV release and usage for each request. A global all() is not automatically the answer either. It keeps the kernel loop running but can still append junk to a request that already ended if the per-row mask is missing.
This is why a single-prompt parity test does not validate a batching rewrite. The failure lives in the relationship between a row of logits and the request lifecycle that row belongs to.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →