Model and Inference Engineering · Staff
The new LoRA adapter has gradients. Why do its weights never change?
The question
Interview question
A fine-tuning job builds its optimizer, then attaches a new LoRA adapter and marks its weights trainable. Backward produces nonzero gradients for the adapter. The loss moves because other trainable parameters still update, but the new adapter's values remain unchanged. What did the optimizer know when it was constructed?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
An optimizer updates the parameters in its parameter groups, not every tensor that happens to have a gradient. If the adapter was attached after the optimizer captured its parameter list, backward can compute .grad while optimizer.step() never visits those tensors. PyTorch's add_param_group documentation explicitly covers adding newly trainable parameters. The same symptom can appear if the adapter exists but requires_grad was set after a filtered optimizer was built, or if an adapter module was replaced with new parameter objects after optimizer creation.
I would compare object identity and names, not just counts. List model parameters that require gradients, parameters owned by each optimizer group, duplicates, and the difference between them. Save one adapter value, gradient and optimizer state before and after a known step. Check whether AMP skipped that step, whether gradient clipping set it to zero, and whether the adapter's learning-rate group is zero. A nonzero gradient alone is not proof that the intended optimizer updated the intended parameter. Confirm the base weights are frozen if that is the recipe, and verify only the expected adapter matrices and any chosen biases enter the optimizer.
The repair depends on when the adapter is created. Prefer constructing all trainable modules before building the optimizer and scheduler. If unfreezing or adding parameters during training is intentional, add a group with explicit learning rate and weight decay, initialize its optimizer state, and decide how its schedule relates to existing groups. On resume, inspect the saved parameter-group layout and mapping. Optimizer state may be associated with group order and parameter IDs rather than semantic names, so a changed layout can restore moments to the wrong place or fail to load. Rebuild and validate the mapping instead of ignoring a warning.
For a small test, run one deterministic batch and assert a particular adapter tensor changes by a nonzero amount after a committed step while a frozen base tensor does not. Then save and resume and make the same assertion again. A training job restarts with half as many GPUs. Can its optimizer state be loaded safely? asks whether optimizer state can load after changing GPU count. This question is about a trainable parameter that never belonged to the optimizer at all. Fluent outputs and a decreasing loss can mask it for a surprisingly long time.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →