Model and Inference Engineering · Staff
The Adam moments restored. Why is the next update wrong?
The question
Interview question
A long training run resumes from a checkpoint. The weights, first moments and second moments all load. The learning-rate scheduler is at the expected point. But the first optimizer update after resume differs from an uninterrupted run. Someone restored the moment tensors and reset each parameter's Adam step count to zero. Why does that matter?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Adam does not update weights from just m and v. It corrects the moving averages using the step number: roughly m_hat = m / (1 - β1^t) and v_hat = v / (1 - β2^t), followed by an update proportional to m_hat / (sqrt(v_hat) + eps). The correction is large at the beginning because moments initialized at zero are biased toward zero. A mature moment tensor paired with t = 1 is a state that never existed in the uninterrupted run. PyTorch's Adam implementation carries step state alongside the moment tensors for this reason.
Take a simple steady gradient g. After many steps, m ≈ g and v ≈ g². If the step count is wrongly reset, the next step uses bias-correction factors close to 0.1 and 0.001 for the common β1 = 0.9, β2 = 0.999. Ignoring epsilon, that gives an update around 0.316 times the expected magnitude, even though the learning rate is unchanged. Real gradients and mixed precision may make the direction and size of the discrepancy less neat. The example just proves that restoring two tensors is not enough.
I would take one saved training batch and branch at the checkpoint. Compare the uninterrupted step with the restored step on parameter gradients, optimizer step, exp_avg, exp_avg_sq, effective update and scheduler state. Check per-parameter step counts in sharded optimizers, not only a global iteration number. Also check whether the loader attached each moment to the right parameter. That is a different bug with a similar symptom.
If the original optimizer state cannot be recovered, resetting all Adam state and treating this as a new optimization phase may be safer than combining mature moments with a fresh counter. It still changes the training recipe, so test the continuation and describe it honestly. A checkpoint that loads and reports the right model step has not proved optimizer continuity.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →