skip to content

Why are biases and normalization scale and shift parameters usually exempt from weight decay?

level: middleimportance: should knowfreq 48%

answer

  1. count the parameters in each group
  2. a bias shifts, it does not amplify
  3. the scale is the layer's output gain
  4. decay matrices and kernels only

basics

~20 s

Because shrinking them regularizes almost nothing and can actively hurt. Biases and normalization scale and shift are few in number, do not control how sharply the function bends, and pulling a normalization scale toward zero throttles that layer's output.

solid answer

~50 s

Two arguments, and they point the same way. First, counting: a weight matrix has fan-in times fan-out parameters while its bias has one per unit, and a normalization layer has one scale and one shift per channel, so these groups contribute a negligible share of the total norm — decaying them buys essentially no capacity control. Second, effect: a bias only shifts a unit's output up or down, it cannot create the sharp input-dependent responses that memorise examples, so shrinking it constrains nothing useful. Worse, a normalization layer multiplies its standardized output by the learned scale, so pulling that scale toward zero uniformly attenuates the signal into the next layer and pushes the shift toward forcing zero-mean output. Downstream weights then grow to compensate and the two forces fight. The convention is therefore to decay weight matrices and convolution kernels only, and put everything else in an undecayed group.

go deeper

for a junior

Know the convention itself: decay is applied to weight matrices and convolution kernels, while biases and a normalization layer's learned scale and shift are placed in a group with decay switched off.

for a middle

Give both reasons — these groups are a negligible share of the parameter norm, and shrinking an additive offset or an output gain does not limit how sharply the network can bend.

for a senior

Show you check parameter grouping before trusting a decay sweep, and can describe the symptom of getting it wrong: suppressed normalization scales with downstream weight norms drifting up to compensate.

for a principal

Frame this as removing a confound from a tuning study. Bad grouping makes the decay coefficient mean something slightly different in every run, which quietly corrupts comparisons across a team's experiments.

## What the parameter groups actually are A trained network's parameters fall into a few structurally different kinds. There are the multiplicative weights — the entries of dense weight matrices and convolution kernels — which multiply incoming activations. There are bias terms, one per output unit, which add a constant offset. And where normalization layers are used there is a learned scale and a learned shift, typically one of each per channel or per feature, applied after the layer has standardized its input. Weight decay is conventionally applied to the first group only. The other parameters are placed in a group with the decay coefficient set to zero. Understanding why is a good test of whether a candidate knows what decay is actually for. ## Argument one: they barely exist Decay is a budget on the squared L2 norm of the parameters. A dense layer mapping 512 inputs to 512 outputs has 262,144 weights and 512 biases — the biases are two-tenths of a percent of the parameters in that layer, and their contribution to the norm penalty is correspondingly negligible. The same holds for normalization scales and shifts, which are per-channel and therefore vanishingly few next to the kernels feeding them. Whatever regularisation pressure you hoped to apply, applying it to these groups contributes almost none of it. The capacity of the model lives in the multiplicative weights, so that is where a capacity constraint belongs. ## Argument two: shrinking them does not constrain the right thing The reason large weights are dangerous is that they let a layer amplify small differences in its input into large differences in its output — that amplification is what makes it possible to fit noise. A bias does not amplify anything. It translates a unit's pre-activation by a constant, shifting where a nonlinearity's threshold sits relative to the data. Shrinking every bias toward zero does not reduce how sharply the network can bend; it just forces every unit's operating point toward a particular location, which is a systematic constraint on where the model can sit rather than on how complex it can be. On small networks this can measurably hurt, because the bias is often exactly the degree of freedom the layer needs. The normalization scale is worse. The layer standardizes its input to roughly zero mean and unit variance, then computes `scale * normalized + shift`. That scale is the layer's output gain. Decaying it pulls the gain toward zero, which uniformly attenuates everything the layer passes on. This does not simplify the function in any useful sense — the next layer's weights simply grow to compensate, and now you have decay fighting to shrink one parameter while the gradient inflates another. Decaying the shift toward zero similarly biases the layer toward zero-mean output whether or not the data supports that. ## How the failure shows up If you drop every parameter into one decayed group and watch the per-group norms across epochs, the picture is diagnostic. The weight norms find their equilibrium as expected. The normalization scales, however, are being pulled down by decay and pushed up by the loss, which needs a working gain; the tug of war shows up as scales that sit systematically lower than in a correctly grouped run, with downstream weight norms drifting upward to compensate. Held-out accuracy is usually a little worse, and the effect gets larger the stronger the decay coefficient — which is exactly the regime where you were relying on decay to do real work. Be honest about the magnitude, though. On large networks with moderate decay, the practical difference between decaying everything and decaying only the weights is often small, and plenty of results have been obtained with sloppy grouping. The reason to get it right is that it costs nothing, it removes a confound from your decay sweep, and the error grows precisely as you turn decay up. ## The rule to state in an interview Decay the multiplicative parameters — dense weight matrices and convolution kernels. Exempt the additive and gain parameters — biases, normalization scales, normalization shifts. The justification is that decay is a magnitude budget for the parameters that control how strongly the network can amplify its input, and the exempt groups neither amplify anything nor contribute meaningfully to the norm, while one of them directly controls a layer's output gain and reacts badly to being shrunk.

  • Does exempting biases risk letting them overfit?
    Very little. A bias vector holds one parameter per unit, a tiny share of the model, and a bias can only translate a unit's output — it cannot produce the sharp input-dependent responses that memorise individual examples. Overfitting capacity lives in the multiplicative weights. If a network is overfitting through its biases, the fix is a smaller network, not decay on the bias group.
  • What concretely breaks if you decay a normalization layer's learned scale hard?
    That scale multiplies the layer's standardized output, so pulling it toward zero attenuates everything the layer passes downstream. The next layer's weights grow to restore a usable signal, so decay is shrinking one parameter while the gradient inflates another. With a strong coefficient you can effectively silence channels, and the run underfits sooner than the same coefficient would suggest.
  • If a network uses no normalization layers at all, does the convention change?
    Only the normalization part of it. You still decay the weight matrices and kernels and still exempt the biases, for the same counting and no-amplification reasons. With no learned scales and shifts there is simply nothing else to place in the exempt group, so the split reduces to weights versus biases.

saying these in an interview costs you the question

  • Says biases are exempt because their gradients are zero
  • Assumes decaying every parameter is always harmless
  • Cannot explain what a normalization layer's learned scale does
  • Thinks the exemption is about numerical stability
  • Believes biases carry meaningful overfitting capacity

context