Model and Inference Engineering · Staff
Every training rank has a different shard. Why are its random augmentations identical?
The question
Interview question
Eight data-parallel ranks each read different examples. The team set one global seed for reproducibility. Logs show that every rank applies the same crop pattern and corruption sequence to corresponding positions in its local stream. Is sharing a seed always wrong?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
No. There are several random streams with different jobs. Ranks may need the same initial model weights, and a distributed sampler commonly uses a shared shuffle seed plus rank-aware partitioning so ranks choose nonoverlapping indices. But augmentation RNG copied unchanged into every rank and worker can correlate the transformations applied to different data. In some setups that correlation reduces useful variation, and in others it creates exact duplicate views when examples repeat. The seed alone does not tell us whether examples overlap. We need to inspect sampled IDs and transformed outputs. PyTorch's DistributedSampler documents the shared sampler seed and rank partition, while its data-loading guidance explains worker seeds. Those mechanisms should not be conflated with augmentation randomness.
I would log, for a small deterministic run, sample ID, epoch, rank, worker ID, augmentation parameters and a content hash before and after transformation. Check whether the framework seeds workers, whether NumPy or a custom image library was also seeded, and whether persistent workers advance their random state across epochs. Check calls to the sampler's epoch setter so the shuffle order actually changes. If random augmentation is run in the model process rather than the loader, inspect that generator separately. A single manual_seed(42) copied into every subprocess is a common route to synchronized random choices, but it is not a universal outcome for every framework.
For reproducible independent choices, derive augmentation randomness from a stable run seed and identifiers such as sample, epoch and augmentation view, with rank or worker only where it is part of the intended sampling policy. Counter-based or stateless transforms can make replay robust to worker scheduling changes. If the same sample should get the same view regardless of which rank owns it after a restart, do not bake rank into that particular seed. If two intentional views of one sample must differ, include a view ID. Write down which reproducibility property is needed before choosing the formula.
The interviewer may say training loss looks normal. It could. The question is whether the data distribution and diversity match the recipe. Compare augmentation coverage, repeated-view rate and downstream quality, not just one loss curve. The checkpoint resumes at the right step. Why does training repeat old samples? covers replaying old samples after checkpoint resume. Here examples may be distinct and the bug is in how each one is transformed.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →