skip to content

Loss-Term Regularizers

Regularizers that change the objective itself: shrinking every weight toward zero, and softening a hard one-hot target. Interviewers start here because both are one-line changes with large effects.

on this pageshow

explore

questions

7

What does weight decay do to a neural network's weights during training?

level: juniorimportance: must knowfreq 82%

answer

  1. two forces act on every weight
  2. one comes from data, one is constant
  3. multiply by slightly less than one
  4. smaller norm, slightly worse training loss

basics

~20 s

Weight decay pulls every decayed weight slightly toward zero on each update, on top of the gradient step. The result is a smaller-norm solution: the network fits with less extreme weights, which usually improves held-out performance.

solid answer

~50 s

Weight decay is a second force acting on every decayed parameter. Alongside the gradient step, each weight is scaled by a factor just below one, so the update reads `w <- (1 - lr * wd) * w - lr * g`. In the loss-term view, you are minimising the training loss plus a penalty proportional to the squared L2 norm of the weights. The gradient pushes each weight wherever the data wants it; decay pulls it steadily back toward zero, and training settles where the two balance. Weights the loss genuinely depends on stay large; weights the data barely supports end up small. What you buy is a smaller-norm solution — a function that cannot swing as sharply, with fewer extreme weights available to memorise individual examples. Training loss gets slightly worse as a result; the bet is that held-out loss gets better. The decay coefficient sets how hard that pull is.

go deeper

for a junior

Be ready to state the effect in one line: every decayed weight is nudged toward zero on each update, which discourages large weights and usually improves held-out performance at a small cost in training loss.

for a middle

Explain the equilibrium — decay and the loss gradient are opposing forces and each weight rests where they cancel. Be able to say what the decay coefficient scales and why shrinkage never reaches exactly zero.

for a senior

Expect to justify decay against the other regularisers already in the recipe, and to show you judge it by held-out loss rather than training loss whenever you change the coefficient.

for a principal

Own the framing that decay is a capacity budget set deliberately per project, whose right value depends on data volume, run length and what else in the pipeline is already constraining the model.

## The one-line definition Weight decay is a shrinkage rule applied during training: on every update, each parameter in the decayed group is moved a little closer to zero, independently of what the training data is asking for. Written as an update it is ``` w <- (1 - lr * wd) * w - lr * g ``` where `w` is the parameter, `g` is the gradient of the training loss with respect to it, `lr` is the learning rate and `wd` is the decay coefficient. The equivalent loss-function view is that you are no longer minimising the training loss `L(w)` alone but `L(w) + (wd / 2) * ||w||^2`, where `||w||^2` is the sum of the squared weights — the squared L2 norm. ## Two forces, one equilibrium The useful mental model is a tug of war. The gradient term is data-driven: it moves a weight in whatever direction reduces the training loss. The decay term is data-blind: it always points at zero, and its strength is proportional to how large the weight currently is. Each weight comes to rest near the point where these cancel — where the marginal reduction in training loss from making the weight larger exactly matches the penalty for doing so. This is why decay does not simply flatten the network. A weight that carries real signal has a strong, persistent gradient opposing the shrinkage, so it settles at a substantial value. A weight that the loss barely depends on has almost nothing opposing the shrinkage, so it drifts toward zero. Decay therefore acts as a magnitude budget that the optimisation spends on the parameters that pay for themselves. One consequence worth being precise about: the shrinkage is multiplicative, not a threshold. A weight approaches zero asymptotically but does not land on it exactly. Exact zeros are the signature of an L1-style penalty, whose pull toward zero has constant magnitude regardless of how small the weight already is; decay's pull shrinks as the weight shrinks, so it never quite finishes the job. ## What a smaller-norm solution buys A network's sensitivity to its input is bounded by the sizes of the weights it multiplies through. Large weights let a layer amplify small input differences into large output differences, which is exactly what a network needs in order to carve out a decision boundary tight enough to fit noise or individual training examples. Constraining the norm removes that budget. Among all the parameter settings that fit the training data comparably well — and in an over-parameterised network there are many — decay biases the optimiser toward the ones with the least magnitude to spare. That is a bias-variance trade in the classical sense. You are accepting a worse fit to the training set in exchange for a function that varies less across resamples of the data. Empirically the training metric degrades monotonically as you raise decay, while the held-out metric improves, plateaus, and then degrades once the penalty starts suppressing weights the data genuinely supports. ## What it is not Weight decay is not a learning-rate schedule, although in the update above the two appear multiplied together and so are not independent knobs — change the learning rate and the shrinkage applied per step changes with it. It is not applied once at the end of training; it acts on every update, so the total shrinkage a weight experiences also depends on how many steps you run. And it is not a substitute for the rest of the recipe: it composes with data augmentation, dropout, early stopping and simply having more data, all of which pull in the same direction. It is also not applied to every parameter in practice. The convention is to decay weight matrices and convolution kernels, and to exempt bias terms and the learned scale and shift of normalization layers, for reasons that turn on how few of those parameters there are and on what shrinking them actually does to a layer's output. ## What an interviewer is checking The answer they want has three beats: the mechanism (a shrinkage applied on each update, equivalently a squared-norm penalty on the objective), the equilibrium (weights settle where data pull and shrinkage balance, so unsupported weights go small), and the trade (training loss worsens, held-out loss is expected to improve). A candidate who states only that it prevents overfitting has recited a label, not an explanation.

  • If weight decay always shrinks weights, why do they not all collapse to zero?
    Because the training loss pushes back. Each weight settles near the point where the data-driven gradient exactly balances the shrinkage pull, so weights the loss genuinely depends on stay large while weights it barely cares about drift toward zero. Only if the decay coefficient is enormous relative to the loss signal does the whole network collapse toward a trivial constant-output function.
  • Does weight decay improve training loss or worsen it?
    It worsens it, essentially always. You are optimising an objective that is no longer pure training loss, so the fit to the training set is deliberately compromised in exchange for a smaller weight norm. If raising decay makes both training and held-out loss worse, you have crossed from regularising into underfitting and should back off.
  • Why does weight decay not produce exactly zero weights the way an L1 penalty does?
    Its pull toward zero is proportional to the weight's current size, so it weakens as the weight shrinks and only approaches zero asymptotically. An L1 penalty applies a constant-magnitude pull regardless of size, which can overwhelm a small gradient completely and park the weight at exactly zero. Decay shrinks; it does not select.

saying these in an interview costs you the question

  • Says weight decay drives weights to exactly zero like an L1 penalty
  • Claims it improves training loss as well as held-out loss
  • Thinks it is applied once at the end of training
  • Confuses it with decaying the learning rate over time
  • Can only say it prevents overfitting, with no mechanism

context

open as a page

How does label smoothing change the target that a classifier's cross-entropy loss fits?

level: middleimportance: must knowfreq 58%

basics

~10 s

Label smoothing replaces the one-hot target with a blend of one-hot and uniform: the true class is asked for about 1 minus epsilon, every other class for epsilon divided by the number of classes.

open as a page

Why are biases and normalization scale and shift parameters usually exempt from weight decay?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because shrinking them regularizes almost nothing and can actively hurt. Biases and normalization scale and shift are few in number, do not control how sharply the function bends, and pulling a normalization scale toward zero throttles that layer's output.

open as a page

Does label smoothing fix an overconfident classifier's calibration?

level: seniorimportance: should knowfreq 42%

basics

~20 s

It usually reduces miscalibration, because it removes the incentive to drive confidence toward 1. But it is a blunt global cap, not a calibration procedure: too much smoothing turns overconfidence into underconfidence, and it never guarantees confidence tracks accuracy.

open as a page

Sweeping weight decay across four decades on a wearable-accelerometer activity classifier, how do you pick the value?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Read the held-out curve, not the training curve. Across a logarithmic sweep the held-out metric is U-shaped: too little decay changes nothing, too much underfits. Pick from the flat top of that curve, then re-sweep when data volume or schedule changes.

open as a page

How does weight decay affect a layer whose output feeds straight into a normalization layer?

level: seniorimportance: nice to knowfreq 20%

basics

~20 s

For a layer whose output is immediately normalized, rescaling its weights leaves the function unchanged. Weight decay there is not capacity control: it keeps the weight norm from growing, which stabilizes the effective step size those weights receive.

open as a page

Should label smoothing be on for a shared backbone serving both classification and retrieval?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Not automatically. Smoothing tightens each class into an equidistant cluster in the penultimate layer and erases the similarity structure between classes, so it can raise top-1 accuracy while degrading nearest-neighbour retrieval built on the same features.

open as a page