Model and Inference Engineering · Staff
A 12-layer Transformer trains. Why does the 48-layer version fail after moving LayerNorm?
The question
Interview question
The team scales a decoder from 12 to 48 layers and, during a refactor, changes each block from normalization before the sublayer to normalization after the residual addition. The loss spikes early. Someone says both versions have a residual connection and the same number of parameters, so normalization placement cannot be the reason. How would you reason about it and test it?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would write the two updates before touching a learning rate. Ignoring the two sublayers inside a block, pre-normalization looks like x_next = x + F(LN(x)). Post-normalization looks like x_next = LN(x + F(x)). In the first expression, there is a direct identity path from x to x_next, so a gradient has a path backward through the stack that does not have to pass through each sublayer or its normalization. In the second, that residual addition is itself followed by normalization. It is still a residual architecture, but the gradient does not get the same untouched path across blocks.
That is a mechanism, not a claim that post-normalization is broken. The original analysis of normalization placement found different initialization gradient behavior and showed why post-normalization often needed careful warm-up. Pre-normalization can have its own depth and representation trade-offs. Many viable architectures add residual scaling or other changes. We should compare the actual 12 and 48 layer recipes rather than make a rule out of one paper.
First establish whether the refactor is really the only change. Run a small deterministic forward and backward test of one block with identical weights where the two formulas are implemented exactly. Check the residual branch, attention and MLP order, norm epsilon, initialization and dtype. Then build a controlled matrix: 12 and 48 layers, each with pre and post placement, same token batch, optimizer, schedule and initialization policy where meaningful. Log per-layer activation RMS, gradient norms, update-to-weight ratios and any clipped or skipped steps during the first few hundred updates. A global loss trace hides a gradient that grows only near the output or dies near the input.
If the 48-layer post variant alone blows up, try an appropriate warm-up and initialization or residual scaling as separate changes, with a budgeted pilot before a full run. Do not simply turn on a huge clip value and declare the architecture stable because the chart looks finite. Compare validation and task quality once it trains, because a stable pre-normalized model is not automatically the better model for the target. Also watch output logit scale and final normalization. A “fix” that changes the function at every layer can change more than gradient flow.
One useful interviewer push is, “The refactor passed unit tests.” I would expect it to. Both formulas produce tensors of the same shape, and a 12-layer smoke run can pass. The failing property is how the update behaves repeatedly through 48 layers under the optimizer. The test that matters here includes gradients across depth and an actual short training trajectory, not just shape and a finite forward pass.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →