skip to content

Why do gradients vanish across time steps in a simple RNN trained with BPTT?

level: middleimportance: must knowfreq 72%

answer

  1. the same matrix, every single step
  2. a product, not a sum
  3. matrix power means exponential in lag
  4. tanh derivative caps at one
  5. 0.9^200 versus 1.1^200

basics

~10 s

Backprop through time multiplies by the same recurrent Jacobian once per step. Factors below one in magnitude shrink that product geometrically, so gradients from distant steps arrive at zero; factors above one explode it.

solid answer

~50 s

In a simple recurrent net the hidden state is `h_t = tanh(W_h h_(t-1) + W_x x_t + b)`, with the *same* `W_h` at every step. To get the gradient of a loss at step T with respect to a hidden state at step k, backprop through time multiplies T-k copies of the step Jacobian `dh_t/dh_(t-1) = diag(tanh'(a_t)) W_h`. That is a matrix power, so it behaves exponentially in the lag. `tanh'` is `1 - tanh^2`, which never exceeds 1 and is far below 1 once units saturate, so if `W_h`'s largest singular value is also below 1 each backward step strictly contracts: an effective per-step factor of 0.9 over 200 steps gives about 7e-10. Push the recurrent spectrum above 1 and the same product runs the other way — 1.1^200 is roughly 2e8. Vanishing and exploding are one mechanism with the exponent's sign flipped.

go deeper

for a junior

Recall the shape of the answer: the same recurrent matrix is applied at every step, so the backward pass multiplies many copies of one factor, and repeated multiplication either shrinks or blows up.

for a middle

Be ready to write the step Jacobian as the recurrent matrix times a diagonal of tanh derivatives, and to show the exponential arithmetic in the lag. Interviewers at this level expect both failure modes from one equation.

for a senior

Demonstrate that you know what survives: the shared weight's gradient is a sum over lags, so training looks healthy while the long-lag terms are already zero. Talk about the effective memory horizon, not just 'gradients vanish'.

for a principal

Own the framing that vanishing is a silent failure and exploding is a loud one, and that the two demand completely different amounts of monitoring investment. Be able to argue what a team should instrument before it starts trusting long-sequence models.

## The setup A simple ("vanilla") recurrent network carries a hidden state across time with one reused weight matrix: ``` a_t = W_h h_(t-1) + W_x x_t + b h_t = tanh(a_t) ``` `W_h` is the recurrent matrix, `W_x` maps the input at step t, and `tanh` is applied elementwise. The crucial word is *reused*: the same `W_h` appears at step 1 and at step 200. ## What backprop through time actually multiplies Suppose the loss is computed at the final step T. The chain rule gives the sensitivity of that loss to a hidden state far in the past, at step k: ``` dL_T/dh_k = (dL_T/dh_T) * prod_{t=k+1..T} (dh_t/dh_(t-1)) ``` and each factor in that product is ``` dh_t/dh_(t-1) = diag(tanh'(a_t)) * W_h ``` where `tanh'(a) = 1 - tanh(a)^2`. So the long-range gradient is a product of (T - k) matrices that are all the *same* `W_h` scaled by a diagonal of activation derivatives. A product of many near-identical matrices is essentially a matrix power, and matrix powers behave exponentially. ## The two shrinking factors **The activation derivative.** `1 - tanh^2` peaks at exactly 1, and only when the pre-activation is exactly zero. Anywhere else it is strictly less than 1, and once a unit sits in the flat part of the curve it is close to 0. So this factor can only ever hold the product steady or shrink it — it can never amplify. **The recurrent matrix.** If the largest singular value of `W_h` is below 1, then combined with a derivative bounded by 1 every backward step is a strict contraction, and the product decays geometrically in the lag. This is the sufficient condition for vanishing. Concretely, an effective per-step factor of 0.9 over 200 steps is 0.9^200 ~ 7e-10; the gradient reaching step 1 is nine orders of magnitude smaller than the one reaching step 199. ## Exploding is the same equation If the recurrent spectrum is large the product runs the other way: 1.1^200 ~ 1.9e8. A spectral radius above 1 is a *necessary* condition for exploding gradients, not a sufficient one — saturating `tanh` derivatives can still contract a matrix whose radius exceeds 1. The practical point is that vanishing and exploding are not two problems; they are one mechanism, repeated multiplication, seen at the two signs of the exponent. The knife-edge where the product neither grows nor shrinks is a measure-zero regime that training does not stay in. ## What does *not* vanish A common confusion is to conclude that the gradient on `W_h` is zero. It is not. Because `W_h` is shared, its gradient is a **sum** of contributions over every (loss step, source step) pair: ``` dL/dW_h = sum_t sum_(k <= t) (dL_t/dh_t) (dh_t/dh_k) (dh_k/dW_h) ``` The lag-1, lag-2, lag-5 terms are perfectly healthy. Only the long-lag terms are annihilated. The sum is therefore dominated by recent history, and training proceeds happily — it just optimises a short-memory model. This is why the failure is described as *long-range credit assignment dying first*: the model's effective memory horizon is roughly where the per-step factor has driven the contribution below the noise floor of the other terms, and everything beyond that horizon is invisible to the optimiser. ## Why it is hard to notice Exploding gradients announce themselves: a loss spike, an overflow, a step that destroys the parameters. Vanishing gradients produce no error at all. The loss curve looks normal, the gradient norm looks normal (the short-lag terms supply it), and the only symptom is that a dependency the task genuinely requires never gets learned. Someone who has multiplied `tanh`'s peak derivative of 1 by itself 200 times and watched a healthy-looking loss curve at the same time understands why the diagnosis has to be done deliberately rather than read off the training log. ## The interview framing Say the mechanism (repeated multiplication by one Jacobian), name the two factors (`tanh` derivative bounded by 1, and the recurrent spectrum), give the exponential arithmetic, and note that vanishing and exploding are the same expression. Then add the subtlety that the total weight gradient survives while the long-lag component does not — that last point is what separates a memorised answer from an understood one.

  • What is the largest value tanh's derivative can take, and why does that matter here?
    The derivative is `1 - tanh(a)^2`, which equals 1 only when the pre-activation is exactly zero and is strictly smaller everywhere else. It therefore acts as a multiplier in [0, 1] at every backward step: it can hold the product steady at best, and shrinks it as soon as units move into the flat region. It can never rescue a shrinking product.
  • If the recurrent matrix has spectral radius 1.2, are exploding gradients guaranteed?
    No. A spectral radius above 1 is a necessary condition for exploding gradients, not a sufficient one. The step Jacobian is the recurrent matrix scaled by activation derivatives, and once units saturate those derivatives can pull the effective per-step factor back below 1. You can have a large recurrent spectrum and still see gradients that vanish.
  • If long-range gradients are zero, why isn't the gradient on the recurrent weight matrix zero?
    Because the matrix is shared, its gradient sums contributions over every lag. Lag-1 through lag-20 terms are full size; only the long-lag terms are annihilated. The sum is dominated by recent steps, so optimisation carries on normally and quietly fits a short-memory model.
  • What is meant by a recurrent model's effective memory horizon?
    It is the lag beyond which a step's gradient contribution has decayed below the magnitude of the other terms in the sum, so it can no longer influence the update. With a per-step factor near 0.9 that horizon is a few dozen steps regardless of how long the sequence actually is.

It is compound interest run backwards: a 10% loss per step is barely noticeable once, and annihilating after two hundred steps.

saying these in an interview costs you the question

  • Says gradients vanish because the sequence is too long to store
  • Treats vanishing and exploding as unrelated mechanisms
  • Claims a smaller learning rate cures vanishing gradients
  • Believes tanh's derivative can exceed one for large inputs
  • Says the total gradient on the recurrent weights is zero
  • Confuses a vanishing gradient with a bad local minimum

context