skip to content

Vanishing and Exploding Gradients

Backprop multiplies one Jacobian per layer, so factors below one shrink the signal to nothing over depth and factors above one blow it up. Interviewers ask what makes each factor small or large.

on this pageshow

questions

3

Why can gradients vanish or explode as they backpropagate through 50 layers?

level: middleimportance: must knowfreq 82%

answer

  1. chain rule, applied repeatedly
  2. one Jacobian per layer
  3. a product, not a sum
  4. sigmoid derivative caps at 0.25
  5. factor above or below one

basics

~20 s

Backpropagation multiplies one Jacobian per layer, so the gradient reaching layer 1 is a product of 50 factors. If those factors average below one the gradient decays geometrically with depth; above one it grows the same way.

solid answer

~50 s

Backpropagation is the chain rule applied layer by layer: the gradient handed down to layer `l-1` is the gradient at layer `l` multiplied by that layer's Jacobian, which is the weight matrix composed with the diagonal of activation derivatives. Over depth D that is a product of D factors, so magnitudes compound geometrically instead of adding. A sigmoid's derivative peaks at 0.25, so a 50-layer plain sigmoid stack can shrink the signal by roughly `0.25^50`, about `1e-30` — layer 1 effectively never moves whatever the learning rate. Push the other way and it is the same arithmetic: if each layer's Jacobian has norm around 1.3, thirty layers multiply the gradient by `1.3^30`, roughly 2600, and the update overshoots anything the loss surface supports. Vanishing and exploding are one phenomenon with the per-layer factor on either side of one. Only Jacobians whose singular values sit near 1 keep the gradient the same order of magnitude at any depth.

go deeper

for a junior

Be ready to state that backpropagation multiplies a factor per layer, so many layers can shrink the gradient to nothing or blow it up, and that the early layers are the ones that stop learning.

for a middle

Expect to write the backward step as the weight matrix times the diagonal of activation derivatives, and to do the arithmetic out loud: 0.25 per sigmoid layer over fifty layers, or a gain of 1.3 over thirty.

for a senior

Show that you reason about the per-layer gain rather than about depth alone, and that you can say which measurable quantity you would look at to tell a compounding depth problem from a single bad layer.

for a principal

Own the framing that trainable depth is a spectral property of the composed Jacobian, not a hyperparameter, and be able to argue what a network design must guarantee for depth to remain free.

## What backpropagation actually computes A feed-forward network is a chain of transformations. For layer `l`, write the pre-activation as `z_l = W_l a_(l-1) + b_l` and the output as `a_l = phi(z_l)`, where `phi` is the elementwise activation function. Training minimises a scalar loss `L` computed from the final layer. Backpropagation is the chain rule applied to that chain. Let `delta_l = dL/dz_l` be the gradient of the loss with respect to layer `l`'s pre-activation. Then ``` delta_(l-1) = W_l^T * diag(phi'(z_l)) * delta_l ``` The matrix `J_l = W_l^T * diag(phi'(z_l))` is the layer's backward Jacobian: the weight matrix composed with a diagonal of activation derivatives. Getting from the loss to the first layer means applying `J_D`, then `J_(D-1)`, and so on — the gradient at the bottom is a **product** of D matrices applied to the gradient at the top. ## Why a product behaves so badly Sums are forgiving; products are not. Taking norms, ``` ||delta_0|| <= ( ||J_1|| * ||J_2|| * ... * ||J_D|| ) * ||delta_D|| ``` Rewrite the product as a sum of logs: the total log-gain is `sum_l log||J_l||`. If the *average* per-layer log-gain is some constant `c`, the total gain is roughly `exp(D * c)` — exponential in depth. Two consequences follow immediately. First, `c < 0` (average gain below one) gives a gradient that dies exponentially; `c > 0` gives one that grows exponentially. Second, `c = 0` is a knife-edge, not a basin: nothing pulls a plain stack back to it, so any systematic bias in the per-layer gain compounds. ## The vanishing side The sigmoid's derivative is `sigma'(z) = sigma(z) * (1 - sigma(z))`, whose maximum value is 0.25, attained at `z = 0`. So every diagonal factor in a sigmoid stack multiplies the backward signal by **at most** 0.25, in the best case, before the weight matrix is even considered. Fifty such layers give `0.25^50`, which is about `1e-30`. Even with weight matrices of norm exactly one, the gradient arriving at layer 1 is thirty orders of magnitude smaller than the one at the head. That layer will not move: to give it a meaningful step you would need a learning rate around `1e28` larger than the one the top layers can tolerate, which would destroy them. This is not a floating-point bug. Even in exact arithmetic, gradient descent takes steps proportional to the gradient, so a gradient that small means no learning. It is a conditioning property of the composed map. A tanh unit has a derivative peaking at 1, so it is better behaved than a sigmoid — but that peak is reached only exactly at zero pre-activation. Drive units away from zero and the factor drops below one and the same compounding takes over. ## The exploding side Weight scale is the other half of each factor: `||J_l|| <= ||W_l|| * max|phi'|`. If the weights are large enough that the per-layer Jacobian norm sits at 1.3, then thirty layers multiply the gradient by `1.3^30`, about 2600. The resulting step jumps far outside the region where the local linear model of the loss is valid; weights grow, so the next iteration's Jacobians are larger still, and the run diverges within a handful of steps. ## The near-identity condition Stated spectrally rather than architecturally: if a layer's backward Jacobian has all its singular values near 1, it acts like an isometry on the backpropagated vector — it rotates it without changing its length. A product of D such Jacobians still has norm near 1, so depth costs nothing. That spectral condition is what every structural remedy for deep training is buying, and it is a condition on the Jacobian, which is weights and activation derivatives *together*, and along the whole trajectory, not only at the first step. ## The interview probe A common probe is: "is a vanishing gradient a property of the loss, of the depth, of the activation, or of the weight scale?" None of them alone. Depth sets the exponent; the activation derivative and the weight scale together set the base; the loss only sets the starting vector at the head (a mismatched loss and output unit can contribute one saturating factor there, but cannot produce geometric decay). The quantity that decides the outcome is the product of all the layer Jacobians. ## Recognising each in a run Vanishing: the loss falls then plateaus early, early-layer weights are indistinguishable from their initial values, gradient magnitudes span many orders of magnitude between head and input, and adding depth makes results worse. Exploding: the loss is stable and then spikes, weight norms grow fast, and the run ends in overflow.

  • Does the loss function itself cause vanishing gradients?
    Not the geometric part. The loss only sets the gradient vector at the output, so a poorly matched pairing — squared error on a saturating output unit, say — can contribute a single small factor at the head. Decay that compounds over depth comes from the product of layer Jacobians, and would happen with any loss.
  • Can one network vanish in some places and explode in others?
    Yes. The gradient at a parameter is the product over the layers between it and the loss, so blocks with different weight scales or different saturation levels have different gains. It is common to measure gradient norms differing by orders of magnitude across a network, and to see the balance shift during training as weights grow.
  • Why is exactly one an unstable place for the per-layer gain to sit?
    Because the total gain is the per-layer gain raised to the depth. At exactly one the gradient is preserved, but nothing restores that value if it drifts: a per-layer gain of 1.05 over forty layers is already a factor of seven, and 0.95 is a factor of one-eighth. Depth amplifies any systematic bias.

It is compounding, not addition. Losing 3% once is nothing; losing 3% at each of fifty stages leaves under a quarter of what you started with, and the same rate in the other direction multiplies it four-fold.

saying these in an interview costs you the question

  • Says the gradient shrinks because the loss value is small
  • Calls it purely a learning-rate problem
  • Thinks depth alone causes it regardless of per-layer gain
  • Treats vanishing and exploding as unrelated failures
  • Says layer gradients add up rather than multiply
  • Claims it is only a floating-point precision issue

context

open as a page

What does a training loss that spikes and then becomes NaN say about gradients?

level: juniorimportance: should knowfreq 60%

basics

~20 s

It is the classic signature of exploding gradients: one update was large enough to push weights into a range where activations overflow to infinity, and infinity arithmetic then produces NaN. Once NaN reaches the weights, the run never recovers.

open as a page

A 40-layer network has gradient norms of 1e-2 at the head and 1e-9 at the input — what do you conclude?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Seven orders of magnitude across the depth means the gradient is decaying geometrically layer by layer. Only the top few layers learn; the rest still hold their initial weights, acting as a fixed random feature map. Effective depth is about four.

open as a page