Model and Inference Engineering · Principal
The optimizer checkpoint loads. Did its moments attach to the right parameters?
The question
Interview question
A model refactor changes the order in which parameters are registered. All tensor shapes still match. The weights load by name and a training resume succeeds without an exception, but loss behavior changes. Someone says the optimizer checkpoint loaded, so its momentum must be correct. What would you check?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
An Adam-style optimizer has state associated with each parameter, such as first and second moment estimates. Those values are not interchangeable just because two matrices have the same dimensions. PyTorch's optimizer state documentation explains that saved parameter IDs are matched with current parameter groups in order and that the loader does not verify the semantic identity of each matched parameter. Depending on how the optimizer was rebuilt, a refactor can bind old moments to a different weight without a shape error. Sharded training adds another mapping layer, but the underlying invariant is the same.
I would compare checkpoint and runtime parameter names, group assignment, shape, dtype and a stable identity for each logical weight. Inspect optimizer state for representative layers and confirm the intended moment tensors follow them after save and load. Then run a controlled one-step continuation from the old code and the refactored code on the same microbatch. Compare weights, gradients and optimizer state, accounting for expected numeric differences in distributed kernels. A raw loss chart over the first hundred resumed steps is a late signal, not proof of correct mapping.
The safest migration is an explicit mapping from old logical parameters to new ones, including aliases and changed sharding, with a declared policy for any new or removed parameter. If a layer changed meaning while preserving shape, reject an automatic state transfer for it. A fresh optimizer is an option, but that changes the training recipe and must be evaluated as such. Preserve a migration report alongside the checkpoint so a future rollback can reproduce the transition.
Could loading by parameter names solve everything? It helps if the names are stable and the implementation actually uses them to remap state. Names can also change or be reused after an architecture change. A training job restarts with half as many GPUs. Can its optimizer state be loaded safely? asks whether optimizer state can survive a GPU-count change. The checkpoint loads and loss starts normally. Why do tied embeddings drift apart after training resumes? asks whether formerly tied embeddings became separate. This question isolates silent semantic misbinding after parameter registration order changes, even when every tensor fits.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →