Model and Inference Engineering · Principal
We enabled shuffle. Why did training still see one data source for hours?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
A streaming shuffle with a finite buffer does not make an entire multi-terabyte corpus uniformly random. If the upstream stream delivers ten million code examples before any support dialogues and the buffer holds ten thousand, it can shuffle locally within the code run but cannot choose a dialogue that has not arrived yet. Some dataset libraries also shuffle shard order, which helps if sources are spread across many reasonably balanced shards, but a few giant source-homogeneous files can still produce long blocks. Hugging Face Datasets' streaming guide describes buffer-based approximate shuffling and shard-order shuffling.
I would log source, language, license tier and token count over training steps for each rank. A global mixture target of 20% code is not enough. Plot rolling windows of tokens and inspect each rank's shard assignment. If one rank receives a huge code shard and another a dialogue shard, they can have different losses and compute time even while the global end-of-epoch totals match. Also inspect what the optimizer sees at step boundaries. Long runs from one source can shift gradient direction and optimizer state before the missing sources arrive. The model may later recover, but “the epoch had the right proportions” does not establish the intended training trajectory.
Fix the source stream first. Use smaller, interleaved shards or separate source streams with a controlled token-level mixture sampler. Choose buffer size relative to the longest homogeneous run and available memory. Seed and checkpoint the sampler state so a restart does not replay one source or silently change the sequence. Test the actual first million consumed token IDs and source labels, not only the configuration. Full random shuffle of a massive remote corpus may be too expensive, so state the approximation and its measured window-level error.
If the interviewer asks whether shuffling more is always better, no. A deliberate curriculum can be useful. But it is a curriculum only when we chose its order, measured it and compared quality. The data mixture says 20% code. Why did code dominate the training tokens? asks why a claimed 20% code mixture dominated tokens. The data loader got faster. Why did the model see an easier training set? asks how faster loading changes which samples get consumed under a step budget. This one is about source autocorrelation that remains after a local buffer shuffle.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →