I would write the two updates before touching a learning rate. Ignoring the two sublayers inside a block, pre-normalization looks like x_next = x + F(LN(x)). Post-normalization looks like x_next = LN(x + F(x)). In the first expression, there is a direct identity path from x to x_next, so a gradient has a path backward through the stack that does not have to pass through each sublayer or its normalization. In the second, that residual addition is itself followed by normalization. It is still a residual architecture, but the gradient does not get the same untouched path across blocks.

That is a mechanism, not a claim that post-normalization is broken. The original analysis of normalization placement found different initialization gradient behavior and showed why post-normalization often needed careful warm-up. Pre-normalization can have its own depth and representation trade-offs. Many viable architectures add residual scaling or other changes. We should compare the actual 12 and 48 layer recipes rather than make a rule out of one paper.

First establish whether the refactor is really the only change. Run a small deterministic forward and backward test of one block with identical weights where the two formulas are implemented exactly. Check the residual branch, attention and MLP order, norm epsilon, initialization and dtype. Then build a controlled matrix: 12 and 48 layers, each with pre and post placement, same token batch, optimizer, schedule and initialization policy where meaningful. Log per-layer activation RMS, gradient norms, update-to-weight ratios and any clipped or skipped steps during the first few hundred updates. A global loss trace hides a gradient that grows only near the output or dies near the input.

If the 48-layer post variant alone blows up, try an appropriate warm-up and initialization or residual scaling as separate changes, with a budgeted pilot before a full run. Do not simply turn on a huge clip value and declare the architecture stable because the chart looks finite. Compare validation and task quality once it trains, because a stable pre-normalized model is not automatically the better model for the target. Also watch output logit scale and final normalization. A “fix” that changes the function at every layer can change more than gradient flow.

One useful interviewer push is, “The refactor passed unit tests.” I would expect it to. Both formulas produce tensors of the same shape, and a 12-layer smoke run can pass. The failing property is how the update behaves repeatedly through 48 layers under the optimizer. The test that matters here includes gradients across depth and an actual short training trajectory, not just shape and a finite forward pass.