Model and Inference Engineering · Staff
The same prompt answers differently alone and in a padded batch. Which position did its first token get?
The question
Interview question
In isolation a decoder sees four prompt tokens. In a batch with a longer request, the engine left-pads it by six slots. The attention mask hides the padding, but the first real token now receives position 6 instead of position 0. Why might its output change even with greedy decoding?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The attention mask answers which tokens may be attended to. Position IDs answer where the real tokens are placed in the model's positional scheme. Those are not the same thing. For a model with learned absolute positional embeddings, shifting a four-token prompt from positions 0–3 to 6–9 changes the vectors presented to the model, even though the padded keys are masked. The next generated token also needs position 4, not the padded tensor's physical index 10. For rotary or relative position schemes, a uniform shift may cancel in some self-attention interactions, but implementations, caching, multimodal positions and other positional components can make blanket invariance unsafe. Name the architecture before predicting exact behavior.
I would capture the token IDs, attention mask, position IDs and past-key-value positions for the same request alone and in mixed-length batches. Compare logits before sampling. If the earliest divergence occurs in prefill, inspect how the model's input preparation derived positions. If prefill agrees but later decode differs, inspect cache length, per-request sequence length and position advancement across batches. Hugging Face's generation code has explicit position-ID preparation, and its GPT-2 model documentation distinguishes attention masks from positional indices. Those are implementation references, not a promise that every model uses the same scheme.
The fix is to map real tokens to the logical positions the checkpoint expects, and to keep that mapping consistent when requests join, leave or resume a continuous batch. With a mask [0,0,1,1,1,1], the nonpadding tokens should generally correspond to logical positions [0,1,2,3] for an absolute-position model trained that way. Test with left and right padding where supported, several pad lengths, and transitions from prefill to decode. Compare logits at a tight tolerance for deterministic kernels before looking at sampled text. Some numerical variation is expected across kernels, but a repeatable pad-length-dependent shift in logits is a strong signal of a position bug.
Would I change padding side to make the symptom disappear? Only after checking generation semantics. Right padding in a decoder batch can make a generic next-token path read logits from a pad slot. A correct server tracks each sequence's last real token and positions explicitly. The interview point is the separation of masking, logical position and physical batch layout.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →