Why does initializing every weight in a network's hidden layers to zero break training?
answer
- what can ever make two units differ?
- same input, same output, same gradient
- the layer collapses to width one
- zero also freezes the weight gradients
- only the initialization supplies asymmetry
basics
~20 sAll units in a layer then compute the same output and receive the same gradient, so they update identically and stay duplicates forever. The layer has the power of a single unit. Random asymmetric values break that tie.
solid answer
~50 sTwo hidden units in the same layer see the same inputs. If their incoming weight rows are identical, their outputs are identical, the gradient arriving at each is identical, and the update applied to each is identical -- so they remain identical for the entire run. A width-256 layer then has the expressive power of one unit, and no amount of data or training time separates them, because nothing in the update rule ever differs. Random initialization is what breaks the symmetry: it is the only source of asymmetry the optimizer has. Zero specifically is worse than a generic constant, since in a plain stack a zero weight matrix also drives the activations and the gradients feeding the layers below it to zero. Biases are a different story -- zeroing them is standard, because the random weight matrix already makes the units distinct.
go deeper
Be ready to say in one sentence that identical weights make identical units with identical gradients, and that random initialization is the only thing that breaks the tie.
Expect to trace it through one forward and one backward pass, showing that the gradient arriving at two hidden units is the same number, and to explain why zeroed biases are harmless.
Show you can recognise the fingerprint in a live run -- identical columns in a hidden layer's activations and a network that plateaus at near-constant-predictor accuracy -- rather than reciting the rule.
Own the framing that initialization is a cheap contract on every layer's forward and backward scale. Be able to say which constants are legitimate, such as a zeroed residual branch scale or an output bias set to the class priors, and which never are.
## The symmetry argument Consider a small fully connected network: 60 input features, a hidden layer, and a softmax output -- say a 3-layer speaker-identification MLP. Two units `a` and `b` in the same hidden layer are fed *exactly the same* input vector. The only thing that can make them compute different values is their own incoming weights. Suppose you set the whole weight matrix to a single constant `c`. Then: 1. **Forward:** unit `a` and unit `b` compute the same pre-activation, so the same activation. Every unit in the layer produces one identical number. 2. **Backward:** the gradient of the loss with respect to unit `a`'s output equals the gradient with respect to unit `b`'s output, because they play interchangeable roles in everything downstream (their outgoing weights are also identical). The gradient with respect to their incoming weights is that shared signal times the shared input, so it is identical too. 3. **Update:** whatever the optimizer does, it applies the *same* number to both rows. After the step they are still equal. By induction this holds forever. The layer is permanently one unit wearing 256 hats. This is what "breaking symmetry" means: the initialization must supply the asymmetry, because gradient descent is a deterministic function of the current parameters and never manufactures a difference between two parameters that start equal and are treated identically. ## Why exactly zero is a worse case than a generic constant With `c` nonzero, the shared gradient is generally nonzero -- the units move, they just move together. With `c = 0` the network is more thoroughly stuck. In a plain stack `y = W x`, the gradient with respect to `W` in one layer is proportional to the activation coming in *and* to the weights above it. If every weight matrix is zero, the hidden activations are zero and the signal flowing back is multiplied by zero matrices, so the weight gradients are zero throughout and only the biases move. The network sits at a saddle point that the update rule cannot leave. ## What symptom this produces in a real run It does not look like a crash. It looks like a network that trains a little (the biases and the output layer's offsets adapt) and then plateaus at the accuracy a constant or near-constant predictor would get. If you inspect the hidden layer, every column of the weight matrix is the same vector, and every hidden activation in a batch is the same number repeated across the width. That is a distinctive fingerprint and worth recognising: it points at the initialization, not at the learning rate or the data. ## What you do instead Draw the weights from a zero-mean random distribution -- Gaussian or uniform -- with a spread chosen from the layer's fan-in (Xavier/Glorot for tanh-like activations, He's variant for ReLU). Two roles are being served at once and it helps to keep them apart: - **Randomness** breaks the symmetry. Almost any random draw does this. - **Scale** keeps the signal alive through depth. This is the part that needs a rule; too small and activations shrink layer after layer, too large and they blow up or saturate. ## The legitimate zeros Some parameters *should* start at zero or a fixed constant, and knowing which ones is the mark of someone who understands the argument rather than the slogan: - **Biases:** zero is the normal default. The weights in the same layer are already random, so the units are already distinct; the bias adds no symmetry problem. - **The output layer's bias in a badly imbalanced classifier:** setting it to the log-odds of the class priors starts the model already predicting the base rate, which is a real convenience. - **The final scale parameter of a residual branch:** deliberately zeroed so a very deep residual stack begins as the identity map. The units inside that branch are still randomly initialized, so no symmetry is created -- only the branch's contribution is switched off at step zero. The common thread: a constant is only dangerous when it makes two *interchangeable* units identical. A constant on a parameter that is not duplicated across a width is harmless, and sometimes useful.
- Does initializing all biases to zero cause the same problem?No, and it is the normal default. The symmetry argument bites when two interchangeable units in a layer become identical, and a random weight matrix already makes them distinct -- adding the same bias offset to distinct units keeps them distinct. The one bias worth setting deliberately is the output layer's, which can be started at the log-odds of the class priors so the model opens by predicting the base rate.
- Would a nonzero constant, say 0.1 everywhere, fix it?No. The units now move, because the shared gradient is nonzero, but they move together: every row of the weight matrix receives the same update at every step and the layer stays effectively one unit wide. Constant initialization fails for symmetry reasons, not for magnitude reasons, so no choice of the constant rescues it. You need randomness across the width, and separately a sensible scale.
- If you inspected a run suffering from this, what would you look at first?The hidden activations for a single batch. Under a constant initialization every unit in the layer emits the same value, so the activation matrix has identical columns and its across-width variance is exactly zero. Checking that the incoming weight matrix has identical rows confirms it in one line. The loss curve alone will not tell you -- it just plateaus early, which has a dozen other causes.
saying these in an interview costs you the question
- Says gradient noise will eventually separate identical units
- Claims the problem is only slower convergence, not permanent symmetry
- Thinks zeroing biases is equally harmful
- Believes any nonzero constant initialization fixes it
- Confuses symmetry breaking with choosing the right scale