How do you choose between pre-norm and post-norm for a model you intend to scale deeper?
answer
- decide by target depth, not prototype depth
- trainability gates quality, not the reverse
- one shared config biases the comparison
- the swap is a retrain and a retune
basics
~20 sDecide by the depth you plan to reach, not by a shallow benchmark. Post-norm has been reported slightly better at around six blocks; pre-norm is the placement that still trains at a hundred. Trainability at the target depth dominates, and switching later invalidates your tuned hyperparameters.
solid answer
~50 sTreat it as an architectural commitment made before the first large run. The evidence people quote is genuinely two-sided: at moderate depth, where both train, post-norm has been reported to reach slightly better final quality, while pre-norm is the placement that still trains reliably at extreme depth. If your roadmap is six blocks forever, that shallow result is decision-relevant; if you intend to scale, trainability wins outright, because a placement that diverges at your target depth has no quality number at all. Two second-order points matter. Placement is not hyperparameter-neutral: it changes the learning rates you can use, how much early-step care you need, and how much explicit regularization the model wants — so an A/B run at one fixed configuration measures whichever placement the configuration was tuned for. And a placement change is a retrain: the two compute different functions, so no checkpoint carries over.
go deeper
Know the default: deep stacks are normally built pre-norm, and the placement is chosen up front rather than adjusted during training. Do not expect to defend the tradeoff yet.
Be able to give both halves of the evidence — post-norm reported slightly better where both train, pre-norm the one that trains when the stack gets deep — instead of repeating a one-sided rule of thumb.
Show that you would test the choice fairly: tune each arm before comparing, and recognise that placement moves the learning rate, the early-step handling and the effective regularization along with it.
Own the call as an architectural commitment tied to the depth roadmap, with the retrain cost stated, the comparison methodology specified, and the rationale written down so the team does not re-litigate it at the next scale step.
## What the decision actually is Post-norm puts the normalization after the residual addition, `x_next = Norm(x + F(x))`; pre-norm puts it on the branch input, `x_next = x + F(Norm(x))`. Choosing between them is not a tuning knob you can revisit at step 50,000 — it is a commitment made before the first serious run, because the two compute different functions and no trained weights transfer between them. ## The evidence is genuinely two-sided The honest summary of the published picture is a depth-versus-quality framing: - At **moderate depth** — a handful of blocks — both placements train, and post-norm has been reported to reach slightly better final quality. - At **large depth** — on the order of a hundred blocks — plain post-norm becomes very hard to train at all, while pre-norm trains routinely. A candidate who reports only half of this is not being careful. The reason the field defaulted to pre-norm is not that it wins everywhere; it is that trainability is a precondition and quality is an optimization on top of it. A configuration that diverges in the opening steps of your target-depth run does not have a quality number to compare. So the decision rule is: **pick by the depth you intend to reach, not the depth you are prototyping at.** If a six-block prototype is a stepping stone to a much deeper model, tuning the prototype's placement on its own quality is optimizing the wrong objective — you will pay for it by having to redo the tuning after the switch. ## The comparison trap The most common analysis mistake here is an A/B run at a single fixed hyperparameter configuration. Placement changes what the model wants: how large a learning rate it tolerates, how much care the opening steps need, and how much explicit regularization is appropriate. Whichever placement your configuration was tuned for will look better, and the margin you measure is mostly the tuning, not the architecture. If you want a comparison that means something, tune each arm separately — at minimum sweep the learning rate for each — and report the best of each. Otherwise report the result as what it is: a comparison at the incumbent's operating point. ## The incidental regularizing effect Placement carries a regularization side effect that is easy to miss. Post-norm re-standardizes the main path at every block, which continuously constrains how far the representation's scale can drift; pre-norm deliberately leaves that path alone. So the two arrive at the same nominal weight decay and dropout settings with different effective regularization, and a team switching from post-norm to pre-norm often finds it needs a little more explicit regularization than before to land in the same place. This is another reason a placement swap is a retune, not a substitution: the tuned recipe you trusted was co-adapted to the placement it was tuned with. ## Middle grounds exist The choice is not strictly binary. Published variants keep the post-norm structure but modify the addition so that very deep stacks remain trainable — DeepNorm, for example, scales up the identity branch before the post-norm addition and pairs that with a matched initialization scale, and is reported to train stacks far deeper than plain post-norm manages. There are also designs that normalize both the branch input and the branch output before adding. These are worth knowing about, but they carry their own tuning obligations, and for most teams the interesting question is whether the extra machinery buys anything over pre-norm plus a final normalization. ## How to actually make the call 1. **State the target depth**, including where you expect to be two scale steps from now. This is the input that dominates. 2. **If the target is deep**, take pre-norm, and design in the final normalization before the head from the start. 3. **If the model is genuinely shallow and will stay shallow**, post-norm is defensible, and the reported quality edge is real — but budget for careful early-step handling. 4. **Decide once.** Put the rationale in the model card or design note so it is not re-litigated by whoever inherits the run. 5. **If you must compare**, tune both arms and say so; never ship a placement conclusion drawn from a single shared configuration. ## What an interviewer is listening for That you refuse the false choice between "pre-norm is better" and "post-norm is better", that you name trainability as a gate rather than a metric, that you know the swap invalidates tuning and the checkpoint, and that you would write the decision down rather than leave it as folklore.
- Can you convert a trained post-norm checkpoint to pre-norm to buy trainability at depth?No. The two placements compute different functions, so the learned weights do not mean the same thing under the other arrangement — the model would not merely lose a little quality, it would produce nonsense. A placement change is a retrain from scratch, which is exactly why the decision belongs before the first large run rather than after it.
- Your A/B at one fixed configuration says post-norm wins; do you trust it?Only if both arms were tuned. Placement changes the usable learning rate, the early-step handling and the effective regularization, so a shared configuration favours whichever placement it was tuned for — usually the incumbent. Rerun with at least a learning-rate sweep per arm and compare the best of each, or report the result explicitly as a comparison at the incumbent's operating point.
- What changes in your regularization budget when you move from post-norm to pre-norm?Usually you need a bit more explicit regularization. Post-norm re-standardizes the main path at every block, which constrains representation drift as a side effect; pre-norm deliberately leaves that path untouched and loses the side effect. So the same weight decay and dropout settings land at a different effective strength, and the recipe has to be re-tuned rather than carried over.
saying these in an interview costs you the question
- Treats it as a pure quality benchmark and ignores trainability
- Assumes tuned hyperparameters transfer across the switch
- Says pre-norm is strictly better in every regime
- Proposes swapping placement partway through a training run
- Claims a post-norm checkpoint can be reinterpreted as pre-norm