The attention mask answers which tokens may be attended to. Position IDs answer where the real tokens are placed in the model's positional scheme. Those are not the same thing. For a model with learned absolute positional embeddings, shifting a four-token prompt from positions 0–3 to 6–9 changes the vectors presented to the model, even though the padded keys are masked. The next generated token also needs position 4, not the padded tensor's physical index 10. For rotary or relative position schemes, a uniform shift may cancel in some self-attention interactions, but implementations, caching, multimodal positions and other positional components can make blanket invariance unsafe. Name the architecture before predicting exact behavior.

I would capture the token IDs, attention mask, position IDs and past-key-value positions for the same request alone and in mixed-length batches. Compare logits before sampling. If the earliest divergence occurs in prefill, inspect how the model's input preparation derived positions. If prefill agrees but later decode differs, inspect cache length, per-request sequence length and position advancement across batches. Hugging Face's generation code has explicit position-ID preparation, and its GPT-2 model documentation distinguishes attention masks from positional indices. Those are implementation references, not a promise that every model uses the same scheme.

The fix is to map real tokens to the logical positions the checkpoint expects, and to keep that mapping consistent when requests join, leave or resume a continuous batch. With a mask [0,0,1,1,1,1], the nonpadding tokens should generally correspond to logical positions [0,1,2,3] for an absolute-position model trained that way. Test with left and right padding where supported, several pad lengths, and transitions from prefill to decode. Compare logits at a tight tolerance for deterministic kernels before looking at sampled text. Some numerical variation is expected across kernels, but a repeatable pad-length-dependent shift in logits is a strong signal of a position bug.

Would I change padding side to make the symptom disappear? Only after checking generation semantics. Right padding in a decoder batch can make a generic next-token path read logits from a pad slot. A correct server tracks each sequence's last real token and positions explicitly. The interview point is the separation of masking, logical position and physical batch layout.