skip to content

Starting the Run Right

Picking a weight scale that keeps the signal alive, finding a first learning rate with a range test, and proving the loop can memorize one batch. Cheap steps that prevent long wasted runs.

on this pageshow

explore

questions

13

Why does initializing every weight in a network's hidden layers to zero break training?

level: juniorimportance: must knowfreq 74%

answer

  1. what can ever make two units differ?
  2. same input, same output, same gradient
  3. the layer collapses to width one
  4. zero also freezes the weight gradients
  5. only the initialization supplies asymmetry

basics

~20 s

All units in a layer then compute the same output and receive the same gradient, so they update identically and stay duplicates forever. The layer has the power of a single unit. Random asymmetric values break that tie.

solid answer

~50 s

Two hidden units in the same layer see the same inputs. If their incoming weight rows are identical, their outputs are identical, the gradient arriving at each is identical, and the update applied to each is identical -- so they remain identical for the entire run. A width-256 layer then has the expressive power of one unit, and no amount of data or training time separates them, because nothing in the update rule ever differs. Random initialization is what breaks the symmetry: it is the only source of asymmetry the optimizer has. Zero specifically is worse than a generic constant, since in a plain stack a zero weight matrix also drives the activations and the gradients feeding the layers below it to zero. Biases are a different story -- zeroing them is standard, because the random weight matrix already makes the units distinct.

go deeper

for a junior

Be ready to say in one sentence that identical weights make identical units with identical gradients, and that random initialization is the only thing that breaks the tie.

for a middle

Expect to trace it through one forward and one backward pass, showing that the gradient arriving at two hidden units is the same number, and to explain why zeroed biases are harmless.

for a senior

Show you can recognise the fingerprint in a live run -- identical columns in a hidden layer's activations and a network that plateaus at near-constant-predictor accuracy -- rather than reciting the rule.

for a principal

Own the framing that initialization is a cheap contract on every layer's forward and backward scale. Be able to say which constants are legitimate, such as a zeroed residual branch scale or an output bias set to the class priors, and which never are.

## The symmetry argument Consider a small fully connected network: 60 input features, a hidden layer, and a softmax output -- say a 3-layer speaker-identification MLP. Two units `a` and `b` in the same hidden layer are fed *exactly the same* input vector. The only thing that can make them compute different values is their own incoming weights. Suppose you set the whole weight matrix to a single constant `c`. Then: 1. **Forward:** unit `a` and unit `b` compute the same pre-activation, so the same activation. Every unit in the layer produces one identical number. 2. **Backward:** the gradient of the loss with respect to unit `a`'s output equals the gradient with respect to unit `b`'s output, because they play interchangeable roles in everything downstream (their outgoing weights are also identical). The gradient with respect to their incoming weights is that shared signal times the shared input, so it is identical too. 3. **Update:** whatever the optimizer does, it applies the *same* number to both rows. After the step they are still equal. By induction this holds forever. The layer is permanently one unit wearing 256 hats. This is what "breaking symmetry" means: the initialization must supply the asymmetry, because gradient descent is a deterministic function of the current parameters and never manufactures a difference between two parameters that start equal and are treated identically. ## Why exactly zero is a worse case than a generic constant With `c` nonzero, the shared gradient is generally nonzero -- the units move, they just move together. With `c = 0` the network is more thoroughly stuck. In a plain stack `y = W x`, the gradient with respect to `W` in one layer is proportional to the activation coming in *and* to the weights above it. If every weight matrix is zero, the hidden activations are zero and the signal flowing back is multiplied by zero matrices, so the weight gradients are zero throughout and only the biases move. The network sits at a saddle point that the update rule cannot leave. ## What symptom this produces in a real run It does not look like a crash. It looks like a network that trains a little (the biases and the output layer's offsets adapt) and then plateaus at the accuracy a constant or near-constant predictor would get. If you inspect the hidden layer, every column of the weight matrix is the same vector, and every hidden activation in a batch is the same number repeated across the width. That is a distinctive fingerprint and worth recognising: it points at the initialization, not at the learning rate or the data. ## What you do instead Draw the weights from a zero-mean random distribution -- Gaussian or uniform -- with a spread chosen from the layer's fan-in (Xavier/Glorot for tanh-like activations, He's variant for ReLU). Two roles are being served at once and it helps to keep them apart: - **Randomness** breaks the symmetry. Almost any random draw does this. - **Scale** keeps the signal alive through depth. This is the part that needs a rule; too small and activations shrink layer after layer, too large and they blow up or saturate. ## The legitimate zeros Some parameters *should* start at zero or a fixed constant, and knowing which ones is the mark of someone who understands the argument rather than the slogan: - **Biases:** zero is the normal default. The weights in the same layer are already random, so the units are already distinct; the bias adds no symmetry problem. - **The output layer's bias in a badly imbalanced classifier:** setting it to the log-odds of the class priors starts the model already predicting the base rate, which is a real convenience. - **The final scale parameter of a residual branch:** deliberately zeroed so a very deep residual stack begins as the identity map. The units inside that branch are still randomly initialized, so no symmetry is created -- only the branch's contribution is switched off at step zero. The common thread: a constant is only dangerous when it makes two *interchangeable* units identical. A constant on a parameter that is not duplicated across a width is harmless, and sometimes useful.

  • Does initializing all biases to zero cause the same problem?
    No, and it is the normal default. The symmetry argument bites when two interchangeable units in a layer become identical, and a random weight matrix already makes them distinct -- adding the same bias offset to distinct units keeps them distinct. The one bias worth setting deliberately is the output layer's, which can be started at the log-odds of the class priors so the model opens by predicting the base rate.
  • Would a nonzero constant, say 0.1 everywhere, fix it?
    No. The units now move, because the shared gradient is nonzero, but they move together: every row of the weight matrix receives the same update at every step and the layer stays effectively one unit wide. Constant initialization fails for symmetry reasons, not for magnitude reasons, so no choice of the constant rescues it. You need randomness across the width, and separately a sensible scale.
  • If you inspected a run suffering from this, what would you look at first?
    The hidden activations for a single batch. Under a constant initialization every unit in the layer emits the same value, so the activation matrix has identical columns and its across-width variance is exactly zero. Checking that the incoming weight matrix has identical rows confirms it in one line. The loss curve alone will not tell you -- it just plateaus early, which has a dozen other causes.

saying these in an interview costs you the question

  • Says gradient noise will eventually separate identical units
  • Claims the problem is only slower convergence, not permanent symmetry
  • Thinks zeroing biases is equally harmful
  • Believes any nonzero constant initialization fixes it
  • Confuses symmetry breaking with choosing the right scale

context

open as a page

How do you run a learning-rate range test to pick a first rate for a new model?

level: middleimportance: must knowfreq 52%

basics

~20 s

Run one short pass, multiplying the learning rate each step from tiny to divergent, and plot smoothed loss against rate on a log axis. Start the real run about a factor of ten below where the loss bottoms out.

open as a page

Why do you try to overfit a single small batch before launching a full training run?

level: middleimportance: must knowfreq 62%

basics

~20 s

Memorizing one fixed batch needs no generalization, so if the loss will not fall to near zero on eight examples the fault sits in the training machinery — gradient flow, step size, loss wiring — not in the dataset.

open as a page

A training run pins its initialization seed but still varies run to run - what else is random?

level: middleimportance: must knowfreq 66%

basics

~20 s

Initialization is only one of three random streams. The epoch shuffle that sets data order and the per-step stochastic operations - augmentation parameters, dropout masks, random masking - draw too, and parallel loading workers hold their own generators.

open as a page

Why does He initialization use a weight variance of 2/fan_in for ReLU layers?

level: middleimportance: must knowfreq 62%

basics

~10 s

ReLU zeros about half of a symmetric, zero-mean pre-activation distribution, halving the signal's second moment at every layer. Doubling the weight variance from 1/fan_in to 2/fan_in cancels that halving, so activation scale survives depth.

open as a page

Five seeds of a ranking model span 0.6 AUC points - how do you judge a claimed +0.3 gain?

level: seniorimportance: must knowfreq 56%

basics

~20 s

A single run cannot resolve an effect half the size of the seed spread. Rerun both the baseline and the candidate over the same list of seeds, look at the per-seed differences, and report that distribution rather than the best run.

open as a page

What cross-entropy loss should a freshly initialized 1000-class classifier report on its first batch?

level: juniorimportance: should knowfreq 48%

basics

~20 s

About 6.9, which is ln(1000). An untrained head spreads probability roughly uniformly, so the correct class receives about 1/1000 and the loss is -ln(1/1000). A much larger reading means the initial outputs are far from zero.

open as a page

Why must you re-run a learning-rate range test after the batch size or initialization changes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The curve describes a configuration, not an architecture. A larger batch averages away gradient noise, so the divergence knee moves up; a pretrained encoder sits near a good solution, so rates that were fine from scratch wreck its features.

open as a page

Why does a one-batch overfit test floor at 0.4 loss when random cropping and colour jitter stay on?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Augmentation means the batch is not fixed: each step shows a differently cropped and colour-shifted image, so the model fits a distribution instead of memorizing points. Switch augmentation off first, then the other noise sources.

open as a page

With a fixed training budget for a bake-off of eight variants, how do you split runs between variants and seeds?

level: principalimportance: should knowfreq 35%

basics

~20 s

Screen broadly, confirm narrowly. Give every variant one short run to eliminate clear losers, then promote two or three survivors to a paired multi-seed confirmation. Never ship on the screening ranking - a one-seed leaderboard reorders when rerun.

open as a page

Two training runs with identical seeds and identical data order diverge by step 4000 - why?

level: seniorimportance: nice to knowfreq 38%

basics

~10 s

Floating-point addition is not associative, and parallel kernels sum partial results in an order that varies between launches. Last-bit gradient differences feed back through thousands of updates until the two runs are visibly apart.

open as a page

In a 200-block residual network, why zero-initialize each block's final scale parameter?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

It makes every residual branch output exactly zero at step zero, so the stack starts as the identity map and the signal keeps the scale it entered with instead of growing block by block. The branches then switch on gradually.

open as a page

When is a learning-rate range test the wrong tool for choosing the rate you ship?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

A range test measures one thing: the largest rate that keeps training loss falling over a couple of hundred steps. It is silent on generalization, on stability thousands of steps later, and on recipes whose good rate is already known.

open as a page