The step number names a model update. It does not uniquely name a position in a shuffling and prefetching data pipeline. A loader may have built a shuffled index, assigned disjoint positions to ranks, queued batches ahead in worker processes and committed fewer batches to the optimizer than it already read. At checkpoint time, “read from disk,” “put on a queue,” “used in backward” and “included in a committed optimizer step” are different positions. Resuming from the wrong one can replay or skip samples while the model state restores perfectly.

I would define the authoritative cursor as the set of samples consumed by committed updates. For a simple deterministic global stream, a dataset generation, shuffle seed, global consumed-sample count and exact batch construction recipe may reconstruct the next sample. That only works if mapping from global position to per-rank examples is stable across restore. In a general loader, save the sampler's permutation or sufficient seed and epoch state, per-rank cursor, any partial accumulation state, and the treatment of prefetched batches. Megatron's data loading guide describes document, sample and shuffle indices used in its pipeline. PyTorch's DistributedSampler documentation requires epoch handling for shuffling. Neither document means every custom trainer automatically saves a recoverable cursor.

To find this bug, take a narrow window around the checkpoint. Log sample IDs and token ranges for the last few committed steps before failure and the first few after resume, grouped by rank and accumulation slot. Check whether batches in the loader queue were counted as consumed when the checkpoint was written. Verify that the resume code restores the same dataset manifest, index files and shuffle state before it constructs the iterator. An identical integer seed can still produce a different order if rank count, dataset order, filtering, packing or library behavior changed. If workers are persistent, inspect their states rather than only the main process RNG.

Now suppose the interviewer says the job also restarts on 32 ranks. Exact per-rank cursor restoration is no longer a simple copy. We need a global definition of the unconsumed stream and a new partitioning of it, or accept a documented replay window. A fixed global sample sequence with stable IDs makes this manageable, although sequence packing and variable batch sizes require extra care. If the training contract permits a small amount of replay, measure it and ensure no eval examples or excluded documents enter. If it promises each sample once within a generation, do not claim that merely restoring the step number satisfies it.

For recovery I would restart from the last checkpoint that includes coherent model and data-position state, then run a short deterministic comparison against an uninterrupted control job. Check exact sample-ID continuity, not only loss. Rank 7 dies while the training checkpoint is being written. Which checkpoint can you resume? asks whether a distributed checkpoint is complete enough to load. This page asks whether the loaded run continues the intended data stream. A successful optimizer.step() at 80,001 proves neither.