Model and Inference Engineering · Staff
The same weights and greedy decoding give different tokens on eight GPUs. Is that a bug?
The question
Interview question
A model generates one answer on a single GPU and another with tensor parallelism across eight GPUs. Same checkpoint, prompt and greedy decoding. Engineers say greedy output is deterministic, so one deployment must be broken. How would you distinguish a correctness error from numerical sensitivity?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Greedy decoding removes sampling randomness. It does not make floating-point arithmetic bitwise identical across execution layouts. Parallel reductions may add terms in another order, fused kernels may use different accumulation precision, and matrix libraries may choose different algorithms. Addition in finite precision is not associative. PyTorch's numerical accuracy note explicitly cautions that mathematically equivalent batched and unbatched computations may not be bitwise identical. If the top two next-token logits are nearly tied, a small logit difference can reverse their order. From then on the model conditions on a different token, so whole responses can diverge sharply.
That explanation is not a free pass. First hold the input contract fixed: token IDs, position IDs, mask, chat template, weights, quantization, cache and stop conditions. Compare prefill logits for the same prompt before any generated token diverges. Record top logit gap, maximum absolute and relative error, and the layer where hidden states first depart. If the gap is large and the winning token flips, suspect a sharding, all-reduce, packing or cache bug. If only near-ties flip within an expected tolerance and most logits track closely, numerical layout sensitivity is plausible. Repeat over representative prompts, lengths and batch sizes.
For a suspected tensor-parallel bug, check how weight matrices are partitioned, where partial results are summed, and whether a bias or normalization is applied once or once per shard. Test a tiny deterministic layer against a high-precision reference. A kernel can be internally deterministic and still implement the wrong operation. Conversely, a bitwise mismatch can be harmless when task outcomes and probabilities are stable. A production regression needs both mechanism and impact, not a blanket equality requirement or a blanket tolerance.
What guarantee should users get? If an API promises reproducible byte-for-byte output, fix the hardware and software path or document the narrow conditions under which that promise holds. For most interactive serving, evaluate task-level stability, rare high-impact flips and quality across layouts. The same prompt answers differently alone and in a padded batch. Which position did its first token get? covers a padded batch changing logical positions, which is a semantic input mismatch. This question assumes the input is identical and asks whether numerical differences in the computation can alter a greedy choice.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →