skip to content

In a pre-norm stack, how does the residual stream's scale change with depth?

level: seniorimportance: should knowfreq 38%

answer

  1. the branch normalizes its own input first
  2. each block adds a fixed-size update
  3. uncorrelated variances add up
  4. what does the output head see at the end?

basics

~20 s

It grows. Each pre-norm block adds a branch output whose size is set by its own normalized input, not by the stream, so contributions accumulate and the stream's standard deviation rises roughly with the square root of depth. That is why a final normalization must sit before the output head.

solid answer

~50 s

In pre-norm, a block computes `x_next = x + F(Norm(x))`. The branch normalizes its own input first, so the size of what it adds does not shrink as the stream grows — each block contributes a roughly fixed-magnitude update. If those updates are similar in size and roughly uncorrelated, variance accumulates about linearly with depth, so the stream's standard deviation grows about like `sqrt(L)`. Two consequences matter in practice. First, the output of the last block was never normalized, so the head would see activations whose scale depends on depth and drifts during training — pre-norm stacks therefore add one final normalization after the last block. Second, a fixed-size update against a growing stream is a relatively smaller perturbation, so later blocks change the representation proportionally less. Post-norm has neither issue: it re-standardizes after every addition, and pays instead on the gradient side.

go deeper

for a junior

Know that a pre-norm network ends with one extra normalization layer after the last block, before the output head, and that this is part of the design rather than an optional extra.

for a middle

Explain the mechanism: the branch normalizes its own input, so what it adds does not shrink as the stream grows, and roughly independent contributions accumulate. Be able to state the square-root-of-depth trend with its assumption.

for a senior

Show you would diagnose this from logs — per-block stream norms and branch-to-stream ratios — and that you can tell a design property apart from a bug, rather than hunting for an imaginary numerical failure.

for a principal

Frame it as something depth scaling must budget for: the head's input scale and each block's share of influence change as you add depth, so comparisons across model sizes need this accounted for before they mean anything.

## Where the growth comes from A pre-norm block is `x_next = x + F(Norm(x))`. The important detail is that `F` never sees the raw stream — it sees `Norm(x)`, a standardized version. So the magnitude of the branch output is governed by the branch's own weights and its learned scale, and is essentially decoupled from how large `x` has become. Every block therefore adds an update of roughly its own characteristic size, no matter how far down the stack it sits. Now accumulate. Writing the stack out, `x_L = x_0 + sum over blocks of F_i(Norm(x_{i-1}))`. If those per-block updates are of comparable magnitude and are roughly uncorrelated with each other, their variances add: the variance of the residual stream grows roughly linearly in the number of blocks `L`, so its standard deviation grows roughly like `sqrt(L)`. The exact factor depends on how correlated the branch outputs are and on the learned scales, so treat `sqrt(L)` as the shape of the trend, not a formula to quote. Empirically the per-block norm of the residual stream in a trained deep pre-norm model is monotonically increasing with depth, which is the same statement. Post-norm does not do this. `x_next = Norm(x + F(x))` re-standardizes the sum at every block, so the stream is reset to a fixed scale at every depth. Its problem lives on the other side of the ledger: nothing reaches an early block without being rescaled by all the intervening statistics. ## Consequence 1: you need a final normalization Because the addition is what ends a pre-norm block, the very last thing the stack produces is an un-normalized sum. Feed that straight into an output head and the head's inputs have a scale that (a) depends on how many blocks you stacked and (b) drifts upward during training as the branches learn. Both make the head's effective learning rate a moving target, and both make the same head hyperparameters behave differently across model depths. The fix is standard and cheap: one normalization layer after the last block, before the head. This is not decoration — it is the layer that closes the pre-norm design, and forgetting it is a real bug that shows up as an unstable or badly calibrated head rather than as an obvious crash. ## Consequence 2: later blocks matter relatively less A fixed-magnitude update added to a stream that has grown is, in relative terms, a smaller perturbation than the same update added near the input. So the fraction of the representation that a block can change tends to fall with depth. This is a well-known characterisation of pre-norm stacks: the deeper blocks behave more like refinements and less like transformations. It is the usual explanation offered for the reported observation that, at moderate depth where both placements train, post-norm can reach slightly better quality per layer — its blocks are not competing against an ever-larger stream. It also has a diagnostic flavour. If you are debugging a deep pre-norm model that seems to waste its upper blocks, this is the mechanism to reason about, not a bug to hunt. ## What to monitor - The **per-block residual-stream norm**, logged as a curve over depth. It should increase smoothly; a jump at one block usually means a branch whose learned scale has run away. - The **ratio of branch-output norm to stream norm** at each depth. That ratio directly expresses how much a block can still change, and a collapse toward zero in the top blocks is the effect above taken to its extreme. - The **scale of the head's input** across depth variants when you are comparing model sizes. If it moves, check that the final normalization is actually there. ## Bounds and non-problems This growth is not divergence. Nothing here says activations blow up: `sqrt(L)` growth over even a hundred blocks is a modest factor, and every branch still sees a standardized input regardless of stream scale, so no branch is driven into saturation by it. The failure mode is not numerical overflow — it is a depth-dependent input to the head and a depth-dependent share of influence per block. Treat it as a design property of pre-norm you must accommodate, not a defect to eliminate. ## How to answer in an interview Say why the growth happens (branch magnitude is decoupled from the stream because the branch normalizes its own input), give the trend (`sqrt(L)` under the independence assumption, hedged), then name the two consequences: a required final normalization before the head, and the shrinking relative contribution of the deepest blocks. Finish by contrasting with post-norm, which resets the scale every block and pays elsewhere.

  • What would you log during training to see this effect happening?
    Two curves over depth: the residual-stream norm at each block, and the ratio of that block's branch-output norm to the stream norm. The first should rise smoothly with depth; a discontinuity points at a branch whose learned scale is running away. The second shows how much influence each block still has, and a collapse toward zero in the top blocks is this effect at its extreme.
  • Does a post-norm stack suffer the same growth?
    No. Post-norm normalizes after the addition, so the stream is re-standardized to a fixed scale at every block and nothing accumulates — it also needs no extra final normalization, because its last operation already is one. Post-norm's price is paid on the gradient side instead, where nothing reaches an early block without being rescaled by every block above it.
  • Why doesn't the growing stream push the branches into saturation?
    Because each branch normalizes its input before doing anything else, so it always sees standardized activations regardless of how large the stream has become. The growth affects what the output head sees and how much relative influence each block has — not the operating range of the branch functions themselves.

saying these in an interview costs you the question

  • Thinks pre-norm keeps activations standardized at every depth
  • Calls the final normalization before the head optional
  • Claims the growth means activations overflow or diverge
  • Believes deeper pre-norm blocks contribute relatively more
  • Attributes the growth to the skip connection rather than the placement

context