Model and Inference Engineering · Staff
The optimizer update is valid. Why are normalization scales shrinking throughout training?
The question
Interview question
A refactor builds one AdamW parameter group with weight decay for every trainable tensor. A comparison run used separate groups and excluded normalization scales and biases. Both runs have finite gradients and a similar early loss. Over time, the new run's RMSNorm scales trend down. Is AdamW broken?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
AdamW applies a decoupled decay to parameters assigned a nonzero weight-decay value, independent of whether the current task gradient wants a scale or bias to shrink. PyTorch's AdamW documentation describes the decoupled update and parameter-group settings. Excluding normalization and bias from decay is a common recipe choice, not a universal theorem. What matters here is that a supposedly identical training run changed its optimization recipe by moving those parameters into a decay group.
I would enumerate all trainable parameters by logical name and identity, the optimizer group they enter, learning rate and decay, including tied weights and any newly introduced modules. Check that each parameter appears exactly once and that the intended group policy matches the saved training configuration. On a single step, compute the expected contribution of decay for a selected normalization scale, separate from the adaptive gradient update. Compare with the previous run using the same weights and microbatch. A norm scale may also decline because of gradients, so the graph alone does not prove decay caused it.
Then choose the recipe deliberately. Typical transformer training excludes norm scales and biases, while applying decay to many weight matrices, but embedding and output-head policy can differ. If weights are tied, they cannot meaningfully live in two conflicting optimizer groups. A layer name filter is easy to get wrong after a refactor, so validate by parameter type and explicit exceptions, then emit a group inventory with the checkpoint. Evaluate training and downstream quality over enough steps to distinguish a transient from a persistent effect. If the new policy was intentional, record it as a changed experiment rather than calling the runs comparable.
The interviewer may ask why this is separate from The optimizer checkpoint loads. Did its moments attach to the right parameters?, where moments attach to the wrong parameters after loading. Here the optimizer can be freshly created, every state correctly attached and every update numerically valid. The failure is the policy of which parameters receive decay. The new LoRA adapter has gradients. Why do its weights never change? is a missing LoRA parameter. This parameter is present, but it gets an update the old recipe did not authorize.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →