Top-p is defined on one distribution over the whole vocabulary. After temperature and other applicable logit transforms, sort tokens by global probability and keep the smallest prefix whose cumulative probability reaches p. A GPU's local softmax divides by only its shard's mass, so its local p is a different threshold. For example, imagine the global probabilities are 0.35 and 0.25 on shard A, and 0.30 and 0.10 on shard B. At p = 0.60, the global nucleus contains 0.35 and 0.30. Local top-p on A includes both 0.35 and 0.25 because 0.35 is only 0.583 of that shard's mass. The combined candidates now contain a token excluded by the intended sampler.

There is a straightforward correctness baseline: gather the relevant global logits, apply the same penalties, temperature, masks and top-p rule once, then sample one token and broadcast its global ID. vLLM's vocabulary-parallel logits path documents gathering logits across model-parallel ranks before processing. That is an implementation reference, not a claim that every correct sampler must gather the full vocabulary. A distributed algorithm can be exact if it computes the global normalization and nucleus and draws from the correct combined mass. For bounded top-k, exchanging local top-k candidates can be enough. For top-p, a fixed small local candidate count is not generally sufficient without a bound or a fallback, because the nucleus size varies with the distribution.

I would compare next-token distributions with an unsharded reference on tiny crafted logits, not just compare one seeded sample. Include a token outside the nominal tokenizer vocabulary if the sharded head has padding rows. Then check temperature, repetition penalties and allowed-token masks in the actual order used by the model. If ranks use independent random draws, there is a second bug even when their probabilities are correct: one logical decode step must choose one global token.

Greedy decoding can pass every regression test here while top-p remains wrong. Compare distributions as well as sampled strings before calling the optimization equivalent.