Model and Inference Engineering · Principal
Each GPU sampled from its own vocabulary shard. Why did top-p change the model?
The question
Interview question
A serving team shards the output vocabulary across tensor-parallel GPUs. To avoid gathering logits on every decode step, each GPU applies top-p locally, and the system combines its candidates. The answer distribution changes. What did it normalize over?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Top-p is defined on one distribution over the whole vocabulary. After temperature and other applicable logit transforms, sort tokens by global probability and keep the smallest prefix whose cumulative probability reaches p. A GPU's local softmax divides by only its shard's mass, so its local p is a different threshold. For example, imagine the global probabilities are 0.35 and 0.25 on shard A, and 0.30 and 0.10 on shard B. At p = 0.60, the global nucleus contains 0.35 and 0.30. Local top-p on A includes both 0.35 and 0.25 because 0.35 is only 0.583 of that shard's mass. The combined candidates now contain a token excluded by the intended sampler.
There is a straightforward correctness baseline: gather the relevant global logits, apply the same penalties, temperature, masks and top-p rule once, then sample one token and broadcast its global ID. vLLM's vocabulary-parallel logits path documents gathering logits across model-parallel ranks before processing. That is an implementation reference, not a claim that every correct sampler must gather the full vocabulary. A distributed algorithm can be exact if it computes the global normalization and nucleus and draws from the correct combined mass. For bounded top-k, exchanging local top-k candidates can be enough. For top-p, a fixed small local candidate count is not generally sufficient without a bound or a fallback, because the nucleus size varies with the distribution.
I would compare next-token distributions with an unsharded reference on tiny crafted logits, not just compare one seeded sample. Include a token outside the nominal tokenizer vocabulary if the sharded head has padding rows. Then check temperature, repetition penalties and allowed-token masks in the actual order used by the model. If ranks use independent random draws, there is a second bug even when their probabilities are correct: one logical decode step must choose one global token.
Greedy decoding can pass every regression test here while top-p remains wrong. Compare distributions as well as sampled strings before calling the optimization equivalent.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →