Greedy decoding removes sampling randomness. It does not make floating-point arithmetic bitwise identical across execution layouts. Parallel reductions may add terms in another order, fused kernels may use different accumulation precision, and matrix libraries may choose different algorithms. Addition in finite precision is not associative. PyTorch's numerical accuracy note explicitly cautions that mathematically equivalent batched and unbatched computations may not be bitwise identical. If the top two next-token logits are nearly tied, a small logit difference can reverse their order. From then on the model conditions on a different token, so whole responses can diverge sharply.

That explanation is not a free pass. First hold the input contract fixed: token IDs, position IDs, mask, chat template, weights, quantization, cache and stop conditions. Compare prefill logits for the same prompt before any generated token diverges. Record top logit gap, maximum absolute and relative error, and the layer where hidden states first depart. If the gap is large and the winning token flips, suspect a sharding, all-reduce, packing or cache bug. If only near-ties flip within an expected tolerance and most logits track closely, numerical layout sensitivity is plausible. Repeat over representative prompts, lengths and batch sizes.

For a suspected tensor-parallel bug, check how weight matrices are partitioned, where partial results are summed, and whether a bias or normalization is applied once or once per shard. Test a tiny deterministic layer against a high-precision reference. A kernel can be internally deterministic and still implement the wrong operation. Conversely, a bitwise mismatch can be harmless when task outcomes and probabilities are stable. A production regression needs both mechanism and impact, not a blanket equality requirement or a blanket tolerance.

What guarantee should users get? If an API promises reproducible byte-for-byte output, fix the hardware and software path or document the narrow conditions under which that promise holds. For most interactive serving, evaluate task-level stability, rare high-impact flips and quality across layouts. The same prompt answers differently alone and in a padded batch. Which position did its first token get? covers a padded batch changing logical positions, which is a semantic input mismatch. This question assumes the input is identical and asks whether numerical differences in the computation can alter a greedy choice.