skip to content

In pre-norm versus post-norm blocks, where does normalization sit relative to the residual add?

level: middleimportance: must knowfreq 62%

answer

  1. same layer, different side of the plus
  2. one placement leaves a path untouched
  3. expand the stack as a running sum
  4. Norm(x + F(x)) versus x + F(Norm(x))

basics

~20 s

Post-norm normalizes after the residual addition: out = Norm(x + F(x)). Pre-norm normalizes the branch input instead: out = x + F(Norm(x)). Pre-norm leaves an un-normalized identity path running the whole depth of the stack, which is why deep stacks train far more easily.

solid answer

~50 s

Both use the same branch and the same normalization layer; only the position relative to the `+` differs. Post-norm computes `x_next = Norm(x + F(x))`, so every block's output is re-standardized before it moves on. Pre-norm computes `x_next = x + F(Norm(x))`, so the branch sees a standardized input while the addition itself is left alone. Expand a pre-norm stack and you get `x_L = x_0 + F_1(Norm(x_0)) + F_2(Norm(x_1)) + ...` — a literal un-normalized path from input to output. Backward, that path hands block 1 an undistorted copy of the output gradient no matter what the branches do. In post-norm no such path exists: anything reaching an early block has been rescaled by every intervening block's activation statistics, and those statistics are still moving in the opening steps. That is why post-norm gets fragile as you add depth and pre-norm does not.

go deeper

for a junior

Be able to state the two formulas without hesitating: post-norm normalizes after the addition, pre-norm normalizes the branch input before it. Knowing which one is the modern default for deep stacks is enough at this level.

for a middle

Expect to expand a stack of blocks on the whiteboard and point at the un-normalized identity path pre-norm creates. Explain why that path matters for the gradient reaching the first block, and name the cost pre-norm pays for it.

for a senior

Be ready to recognise the symptom in a real run — a deep post-norm stack diverging in the opening steps at a learning rate a shallow one survives — and to say which levers you would pull, including that placement is a retrain, not a hot fix.

for a principal

Own the framing that placement is an architectural commitment made before the first large run: it sets the depth you can reach, the learning rates you can use, and the tuning you inherit. Argue it against the depth you actually plan to scale to.

## The two placements A residual block computes some branch function `F` — an attention sublayer, a feed-forward sublayer, a small convolution stack, whatever the block is made of — and adds its output back to the block's input `x`. A normalization layer rescales a vector of activations using statistics gathered over some chosen axis, then applies a learned scale and shift. The pre-norm/post-norm question is *only* about where that normalization layer sits relative to the addition: - **Post-norm**: `x_next = Norm(x + F(x))` - **Pre-norm**: `x_next = x + F(Norm(x))` Same branch, same normalization layer, same parameter count. One symbol of placement separates them, and it changes how the whole stack behaves at depth. ## The identity path Stack `L` pre-norm blocks and expand the recursion: `x_L = x_0 + F_1(Norm(x_0)) + F_2(Norm(x_1)) + ... + F_L(Norm(x_{L-1}))` The running sum is usually called the **residual stream**. Notice what the expansion says: `x_0` appears in the output *as itself*, with nothing applied to it. Forward, the input reaches the output undistorted. Backward, the gradient with respect to `x_0` is the output gradient times `(I + <sum of branch terms>)`; the `I` is a clean copy of the output gradient that arrives at block 1 without having passed through a single normalization layer. That term is present at initialization, before any branch has learned anything, and it does not depend on the activation statistics of any intermediate block. Post-norm has no such expansion. Every block's output passes through `Norm` before it participates in the next addition, so the route from block 1 to block `L` crosses a normalization layer at every step. A normalization layer is not the identity map: it divides by the standard deviation of its own input and removes the component that would change that statistic. So what reaches an early block has been rescaled by the activation statistics of every block above it — statistics that are themselves changing fastest in the opening steps of training, exactly when the model is least stable. ## What this looks like in practice The canonical demonstration: a 24-layer post-norm encoder stack blows up inside the first few hundred steps at a learning rate a 6-layer version of the same model handles without complaint. The loss spikes, activations grow, and the run never recovers. Move the normalization inside the branch — the same layers, the same width, the same optimizer, one placement change — and the 24-layer stack trains cleanly from step one. The usual rescue for post-norm is a learning-rate warmup, and post-norm's reputation as "warmup-sensitive" is precisely this effect: the placement makes the survival of the first few hundred steps depend on a schedule detail. Pre-norm tolerates larger learning rates and is far less sensitive to how the opening steps are scheduled, which is the practical reason it became the default for stacks past a few dozen blocks. ## What pre-norm costs Pre-norm is not free. Because each block adds a branch output whose magnitude is set by its own normalized input and not by the stream, the residual stream's scale grows as you go deeper, and the last block's output was never normalized at all — so a pre-norm stack needs one final normalization layer after the last block, before the output head. There is also a reported quality wrinkle: at moderate depth, where both placements train, post-norm has been reported to reach slightly better final quality, while pre-norm is the one that still trains at extreme depth. ## Confusions to avoid - **Pre-norm is not "normalize before the activation".** That is an ordering question *inside* the branch. Pre-norm/post-norm is about the position relative to the residual addition. - **Pre-norm does not remove normalization.** The layer is still there, once per branch; it moved, it did not disappear. - **Post-norm is not untrainable.** It trains well at moderate depth and with careful early-step handling; it is depth plus an aggressive learning rate that breaks it. - **The two are different functions.** You cannot reinterpret a trained post-norm checkpoint as pre-norm; the weights mean something different. ## How to answer in an interview Write the two expressions, expand the pre-norm stack to show the un-normalized sum, and say what the identity term guarantees for the gradient reaching block 1. Then name the tradeoff in one sentence: pre-norm buys trainability at depth and pays with a growing residual stream and a required final normalization; post-norm keeps every block's input standardized and pays with fragility in the opening steps as depth grows.

  • Trace the gradient reaching block 1 of a 48-block stack under each placement — what actually differs?
    Under pre-norm, expanding the stack shows the input appearing in the output as itself, so block 1 receives an undistorted copy of the output gradient alongside the branch contributions — an identity term present at initialization. Under post-norm there is no such term: everything reaching block 1 has been rescaled by the activation statistics of all 47 blocks above it, and those statistics are still moving early in training.
  • Does moving the normalization inside the branch change the model's parameter count?
    No. It is the same normalization layer with the same learned scale and shift, applied to a different tensor. Parameter count, width and depth are identical. What changes is the function the stack computes and, with it, the optimization behaviour — which is why the two are not interchangeable checkpoints.
  • What does pre-norm buy you in terms of the learning rate the stack tolerates?
    Noticeably more headroom. Because the identity path is untouched, early updates cannot be amplified by a chain of normalization rescalings, so pre-norm stacks generally train at larger learning rates and are much less sensitive to how the opening steps are handled. Post-norm at the same depth typically needs a gentler start to survive the first few hundred steps.

saying these in an interview costs you the question

  • Says the two placements compute the same function
  • Claims pre-norm removes normalization from the block
  • Thinks the difference is which axis the statistics come from
  • Says post-norm simply cannot be trained
  • Confuses it with ordering normalization before the activation inside the branch

context