Model and Inference Engineering · Staff
The gradient norm is clipped on every microbatch. Why is the accumulated norm still too large?
The question
Interview question
The training loop accumulates eight microbatches before one optimizer step. An engineer adds `clip_grad_norm_` after every backward call, with a maximum norm of one, and expects the gradient passed to the optimizer to have norm at most one. Is that true?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
No. Clipping is a nonlinear operation, and it matters what vector you clip. For a simple one-dimensional example, suppose each microbatch contributes a gradient of 0.5 after the intended loss scaling. Each individual norm is under one, so clipping each one does nothing. The accumulated gradient is four. Even if the code clips the running partial sum after each microbatch, it is no longer equivalent to clipping the final sum once. Early gradients can be shrunk before later gradients cancel them. For instance, contributions of +2 and -2 should sum to zero. Clipping the partial +2 to +1 before adding -2 leaves -1. This is a different optimization algorithm.
State the intended update first. If the effective batch is eight microbatches of equal sample count, scale each microbatch loss appropriately, accumulate their gradients, then clip the final accumulated gradient immediately before the optimizer step. Unequal microbatch sizes need sample or token weighting consistent with the target objective. In distributed training, make sure the norm is computed on the global logical gradient, not independently on each shard or rank with a threshold of one. The required reduction depends on whether gradients and parameters are replicated or partitioned. Check the exact implementation, especially when optimizer states and parameter shards are distributed.
Mixed precision adds another ordering constraint. If a GradScaler scaled the loss, the gradients are still scaled during backward. Unscale them before measuring or clipping their norm, and do so once the effective batch is complete. PyTorch's AMP examples explicitly show both rules: accumulate with a consistent scale across the effective batch, then unscale and clip before the step. A threshold applied to scaled gradients has no stable relationship to the intended threshold. Check nonfinite gradients and synchronized skip behavior before claiming a step committed everywhere.
How would I prove the bug? With a tiny deterministic model, save per-microbatch gradients and manually compute the weighted sum, its global norm and the clipped result. Compare that with the actual optimizer input, before momentum or Adam state changes it. Clipping a gradient does not itself bound the parameter delta under an adaptive optimizer. Add a cancellation case, not only gradients pointing in one direction. Then check real training traces for preclip and postclip norms, skipped steps and loss spikes. We doubled gradient accumulation after losing GPUs. Why did the training recipe change? asks whether changing accumulation changed the effective training recipe. Here the recipe may be unchanged on paper while the clip location quietly changes the optimizer input.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →