A seed initializes a particular random-number path. It is not a cross-engine contract for which random numbers are drawn, how they are assigned to requests, or how candidate tokens are filtered. Different samplers may consume draws in another order, use different RNG algorithms or apply temperature and top-p with slightly different numerical kernels. A batched server may also schedule requests differently. vLLM's sampling parameters define a seed alongside temperature and filtering options, but that does not promise byte-identical output to another engine. More importantly, if logits differ before sampling, blaming RNG misses the real fault.

I would separate the stages. Compare exact token IDs, chat template, position IDs, masks and model revision. Run a fixed prompt at greedy temperature to compare prefill logits and the first next-token choice, recognizing that even greedy behavior can flip on a near tie under different numerical layouts. Then capture a few steps of processed logits, top-k/top-p candidate sets, renormalized probabilities and sampled token for the same requests. Confirm how each engine interprets zero temperature, minimum probabilities, stop tokens and any hidden default. An API surface with the same field names can still implement different effective distributions.

If the product requires reproducible sampling, define the supported scope explicitly: engine build, kernels, device type, batch policy and request ordering. A per-request RNG state can reduce interference from other requests, but it cannot by itself align different algorithms. Persisting the seed without the sampler version is weak provenance. If the goal is quality rather than replaying exact bytes, compare distributional outcomes on a large task set, including tool correctness and refusal behavior. A text diff of two sampled answers is not a meaningful regression test by itself.

One useful diagnostic is to feed exactly the same logit vector to both sampling implementations. If the candidate set or chosen-token distribution differs, isolate the sampler. If they agree on that vector but actual serving differs, investigate the model or prompt path. The same weights and greedy decoding give different tokens on eight GPUs. Is that a bug? asks why greedy outputs differ across GPU layouts because of floating-point computation. This question holds a seed in the story and asks what the seed does and does not make deterministic in a new sampler.