Model and Inference Engineering · Staff
Every training rank got equal work. Why did the epoch repeat some examples?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Distributed training often requires each rank to take the same number of steps. Suppose a dataset has ten examples and four ranks. A sampler that does not drop the tail may pad the index list to twelve and give three indices to each rank. Two examples are seen twice in that epoch. PyTorch's DistributedSampler documentation says drop_last=False adds extra indices to make the dataset evenly divisible, while drop_last=True drops the tail. Equal work does not imply one exposure per example.
For a huge dataset and a few ranks, two extra indices may be irrelevant. For a small fine-tuning or evaluation set with many ranks, the fraction can matter. In this toy epoch it is two extra exposures out of twelve draws. It is not a claim that the same two examples are always duplicated after shuffling. If the training loop calls set_epoch(epoch) as documented, the shuffle can change which indices receive the extra exposure across epochs. Without it, the shuffle order can repeat in a way you did not intend. With drop_last=True, the opposite tradeoff appears: in this example two of ten examples are absent from that epoch.
I would log dataset IDs per rank for a small test epoch and compare total draws, unique IDs and per-ID counts. Check how the trainer defines an epoch, especially if it uses a streaming dataset or custom sampler rather than PyTorch's map-style DistributedSampler. For validation, aggregate by unique example ID or use an evaluation sampler that avoids padded duplicates. Averaging per-rank accuracy over duplicated validation examples can bias a small metric. In training, decide whether unequal last batches, tail dropping, or controlled repetition better matches the objective.
Does repetition automatically make the model worse? No. Sampling with replacement and repeated training epochs are normal choices. The error is to report “one full pass over each example” or a unique-example validation score when the sampler did something else. The index padding changes the exposure count even with a perfectly functioning distributed optimizer.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →