How does weight decay affect a layer whose output feeds straight into a normalization layer?
answer
- the next layer undoes any rescaling
- magnitude changes gradients, not outputs
- the gradient is orthogonal to the weight
- effective step size goes like one over norm squared
basics
~20 sFor a layer whose output is immediately normalized, rescaling its weights leaves the function unchanged. Weight decay there is not capacity control: it keeps the weight norm from growing, which stabilizes the effective step size those weights receive.
solid answer
~50 sNormalization standardizes its input, so if you multiply the preceding layer's weights by any positive constant, both the mean and the standard deviation of that layer's output scale by the same constant and the normalized result is identical. Those weights are scale-invariant: their magnitude does not change the function the network computes at all, so shrinking them cannot be constraining capacity. What magnitude does change is gradients. For a scale-invariant parameter the gradient is orthogonal to the weight vector and its size scales like one over the norm, so the angular step the weights actually take per update goes like `lr / ||w||^2`. Without decay each orthogonal step strictly increases the squared norm, so the norm climbs through training and the effective step size quietly decays. Weight decay supplies the only shrinking force, and the norm settles at an equilibrium where growth and shrinkage balance — which is why on normalized networks decay behaves much more like a learning-rate control than like a regulariser.
go deeper
Know the core surprise: if a layer's output is immediately standardized, multiplying that layer's weights by a constant changes nothing about what the network computes.
Be able to derive the invariance in one line — mean and standard deviation both scale with the weights, so the ratio is unchanged — and say why that undercuts the usual capacity story for decay.
Demonstrate the operational consequence: log per-layer weight norms, read a monotonically climbing norm as an unplanned learning-rate decay, and re-tune decay whenever the learning-rate schedule changes.
Own the argument that on a normalized architecture decay and learning rate parameterise one underlying effective step size, so a tuning protocol that sweeps them independently is asking a badly posed question.
## The invariance Suppose a linear or convolutional layer produces pre-activations `z = W x`, and those pre-activations go straight into a normalization layer that subtracts a mean and divides by a standard deviation computed over the normalized axis. Replace `W` with `cW` for any positive constant `c`. Every entry of `z` scales by `c`, so the mean scales by `c` and the standard deviation scales by `c`, and the ratio `(cz - c*mean) / (c*std)` is exactly `(z - mean) / std`. The normalized output is unchanged. The learned scale and shift applied afterwards are unchanged too, because they act on the already-normalized values. So `W` is scale-invariant: only its direction matters to the function the network computes; its length is a free parameter that the loss cannot see. This is true of essentially every weight matrix or kernel that sits immediately before a normalization layer, which in a modern normalized architecture is most of them. ## The immediate consequence for weight decay The standard story for weight decay is that a smaller norm means a less flexible function. That story assumes the norm affects the function. Here it does not. Shrinking a scale-invariant `W` cannot make the network less expressive, because a scaled-down `W` computes precisely the same thing. Whatever weight decay is doing on these layers, it is not constraining capacity in the usual sense. ## What the norm does control: gradients Two facts follow from scale invariance, and both are provable from it directly. First, the gradient is orthogonal to the weight. If a function satisfies `f(cW) = f(W)` for all `c > 0`, differentiate with respect to `c` at `c = 1` and you get `W . grad = 0`. The gradient never has a component along the weight vector; it only ever rotates the direction. Second, the gradient magnitude scales inversely with the norm. Differentiating `f(cW) = f(W)` with respect to `W` gives `grad at cW = (1/c) * grad at W`. Double the weights and every gradient entry halves. Put those together. A plain gradient step adds a vector orthogonal to `W`, so by Pythagoras the squared norm strictly increases: `||W - lr*g||^2 = ||W||^2 + lr^2 * ||g||^2`. The norm can only grow, never shrink, under gradient steps alone. And the quantity that actually matters — how far the direction of `W` rotates per update — is roughly the step length divided by the norm, which is `lr * ||g|| / ||W||`. Since `||g||` itself scales like `1 / ||W||`, the angular progress per step goes like `lr / ||W||^2`. This is the punchline. In a normalized network without decay, the weight norm climbs monotonically, and because effective progress scales like the inverse squared norm, you get an unplanned learning-rate decay schedule that nobody wrote down. ## What decay restores Add shrinkage and the norm update becomes a competition: multiplicative shrinkage pulls the squared norm down by a factor each step while the orthogonal gradient component pushes it up by the squared step length. The norm settles at an equilibrium where the two cancel, and once there, the effective angular step size is stable rather than steadily shrinking. Raising the decay coefficient lowers the equilibrium norm and therefore raises the effective step size; lowering it does the reverse. On a normalized architecture, the decay coefficient and the learning rate are not independent knobs at all — they jointly determine the effective step size, which is why tuning one without re-examining the other so often gives confusing sweeps. ## Practical symptoms and what to do The symptom of setting decay to zero on a heavily normalized network is a run that starts fine and then stalls earlier than the schedule intends, with per-layer weight norms visibly climbing across epochs. The symptom of a very large coefficient is not the usual gentle underfitting but instability, because the equilibrium norm is small and the effective step size correspondingly large. The operational advice is: log per-layer weight norms, not just the loss; treat a monotonically climbing norm as a signal rather than a curiosity; and re-tune the decay coefficient whenever you change the learning-rate schedule, because on these layers they are two views of the same quantity. Note also that the layers where decay retains its ordinary capacity-control meaning are precisely the ones not followed by normalization — commonly the final output layer, which is worth thinking about separately. ## The honest caveat Scale invariance is exact only for the parameters directly preceding the normalization, and only for the multiplicative part. Real networks mix normalized and unnormalized paths, and momentum complicates the clean orthogonal-step argument. The mechanism is nonetheless real and is the standard explanation for why decay remains essential on architectures where, by the naive capacity argument, it ought to do nothing at all.
- If those weights are scale-invariant, why does their norm grow at all without decay?Because the gradient of a scale-invariant function is orthogonal to the weight vector, so each plain step adds a perpendicular component and the squared norm increases by the squared step length every update. There is no force pointing back toward the origin. Decay supplies the only shrinking term, and the norm then settles where growth and shrinkage balance.
- Which layers in a normalized network still get ordinary capacity control from weight decay?The ones whose output is not immediately normalized — most often the final classification or regression layer, and any branch that skips normalization. There the weight magnitude genuinely changes the function, so shrinking it limits how sharply the model can respond. It is a good reason to think about the output layer's decay separately from the trunk's.
- Why does this make the decay coefficient and the learning rate dependent knobs?On a scale-invariant layer the effective angular step per update goes roughly as the learning rate over the squared weight norm, and the decay coefficient is what sets that equilibrium norm. Raising decay shrinks the norm and so raises effective step size. Changing either one alone moves the same underlying quantity, which is why sweeps that vary only one often look inconsistent.
saying these in an interview costs you the question
- Thinks shrinking a pre-normalization weight shrinks that layer's output
- Treats decay on normalized layers as pure capacity control
- Assumes weight norms stay roughly constant during training
- Says learning rate and decay are always independent knobs
- Claims scale invariance means the weights stop mattering entirely