skip to content

Why can't a GRU learn a dependency 3,000 steps back even with well-behaved gates?

level: seniorimportance: should knowfreq 52%

answer

  1. gates fixed gradients, not retrieval
  2. fixed state, distance-independent capacity
  3. 0.999 to the power 3,000
  4. truncated windows never span it
  5. no way to address an earlier position

basics

~10 s

Gating stops gradients vanishing along an open path, but the fact must still survive in a fixed-size state through every intervening step, and truncated training windows usually never connect the two positions at all.

solid answer

~50 s

Gating solves one specific problem: the multiplicative decay of gradients through a long chain. It does not solve retrieval. To link a pronoun to an antecedent 3,000 tokens earlier, some unit must hold that antecedent in a fixed-size hidden state across all 3,000 steps, keeping its update gate near 0 the whole time while the same state is also being used for everything else the sentence needs — capacity is O(hidden size), independent of how far back the dependency is. Even a gate at 0.999 leaves a product over 3,000 steps that is far below 1. And in practice training truncates backpropagation to windows of a few hundred steps, so no gradient ever connects the pronoun to the antecedent; the model cannot learn the link because it is never shown one. There is also no content-based lookup: the state has no way to go back and address a specific earlier position.

go deeper

for a junior

Know the headline: gates make it possible for information and gradients to persist across many steps, but they do not give the network unlimited memory of everything it has read.

for a middle

Explain the mechanism on both sides — why a near-zero update gate gives a near-identity Jacobian, and why a fixed-size state plus a product of slightly-less-than-one factors still limits what survives.

for a senior

Show how you would diagnose it on a real model: error against cue-target distance, a short-distance control, and a check of the truncation length actually used in training.

for a principal

Frame the tradeoff for a team: whether to spend on wider states and longer truncation windows, or to accept that this class of dependency needs a different retrieval mechanism entirely.

## What gating actually fixed In an ungated recurrent cell, backpropagating a gradient from step `T` to step `t` multiplies together `T - t` Jacobians. If their typical magnitude is below 1, the product decays geometrically and long-range gradients vanish; above 1, it explodes. Gates change this because the state update becomes an interpolation, `h_t = (1 - z_t) * h_(t-1) + z_t * cand_t`. With the update gate near 0 for a unit, the derivative of `h_t` with respect to `h_(t-1)` along that unit is close to 1, so the gradient can travel many steps roughly undamped. That is the whole of what gating buys, and it is real. It is not the same as being able to *use* information from 3,000 steps back. Four separate obstacles remain. ## Obstacle 1 — fixed-size state, unbounded content The hidden state is a fixed vector of `H` numbers. Everything the model needs from the entire past — the antecedent, the current clause, the syntactic state, whatever else — must be encoded in those `H` numbers simultaneously. Capacity does not grow with distance or with sequence length. A long novel does not get a bigger memory than a short paragraph. So carrying a specific rare fact for thousands of steps means dedicating part of a scarce, shared resource to it for the whole span, and every intervening write competes for the same room. ## Obstacle 2 — the gate must stay shut for the entire span For a unit holding the antecedent, the update gate must remain near 0 at every one of the 3,000 steps. If it opens even briefly — because some intervening token looked relevant — the stored value is blended away and cannot be recovered. And "near 0" is not the same as 0: the carried coefficient is a product of `(1 - z_t)` terms. At 0.999 per step, the value after 3,000 steps is scaled by roughly `0.999^3000`, which is about 0.05. Gating slows the decay by orders of magnitude; it does not abolish it. ## Obstacle 3 — no content-based retrieval A gated recurrent cell reads only its own previous state. There is no operation that says "go find the position that best matches this query and bring it back." Anything the model wants at step 3,000 must have been continuously carried forward from step 0 through every intermediate state. That is fundamentally different from a mechanism that can address earlier positions by content, and it is the deepest of the four limits: it is architectural, not an optimisation problem. ## Obstacle 4 — the training signal is truncated and rare Two practical facts finish the job. **Truncation.** Backpropagation through time is almost always truncated to a window of tens to a few hundred steps, because storing activations for the full unroll is expensive. A dependency spanning 3,000 steps sits entirely outside that window: the gradient linking the pronoun's error to the weights that stored the antecedent is never computed. The model is not failing to learn the pattern — it is never shown a learning signal for it. **Sparsity of the pattern.** Even with a long enough window, dependencies at that range are rare in text. A handful of examples per corpus produce a tiny gradient contribution against millions of local-context examples, and the optimiser very reasonably spends capacity on the frequent pattern. ## How to diagnose it in practice If you suspect a long-range failure rather than a general capacity problem, the useful checks are: measure error as a function of the distance between the cue and the target and see whether it degrades with distance specifically; compare against a model given the cue at a short distance, which isolates retrieval from the task itself; and check the truncation length actually used in training against the distances you care about, because that number alone often explains everything. ## The honest conclusion Widening the hidden state helps a little — more room to store things — and lengthening the truncation window helps a little, at a memory cost. Neither turns a gated recurrent cell into something that can reliably retrieve an arbitrary fact from thousands of steps back. The right answer in an interview is to say which of the four limits you think binds for the case in front of you, and to be clear that gating addressed the gradient path only.

  • Does widening the hidden state to a few thousand units fix a 3,000-step dependency?
    It relieves only the capacity limit, and at quadratic parameter cost in the recurrent matrices. The gate must still stay shut for every step, the truncation window is unchanged, and there is still no way to address an earlier position by content. In practice you buy a modest improvement and a much more expensive model, not a solved problem.
  • How would you tell a long-range retrieval failure apart from plain underfitting?
    Plot error against the distance between the cue and the point where it is needed. Underfitting is roughly flat in distance; a retrieval failure degrades as distance grows and largely disappears when the same cue is placed a few steps away. Also check the truncation length used in training — if it is shorter than the distances you care about, that alone explains the failure.
  • Why does the reset gate not help here?
    The reset gate suppresses history on the way into the candidate — it is a mechanism for forgetting on purpose, such as restarting at a segment boundary, not for holding something longer. Long-range retention is governed by the update gate staying near 0 for the relevant units; the reset gate can only make the candidate more input-driven.

It is like relaying a message down a line of people who may each only whisper a fixed number of words: the chain can be careful, but nobody can walk back and re-read what the first person actually said.

saying these in an interview costs you the question

  • Says gates eliminate vanishing gradients entirely
  • Claims a bigger hidden state solves arbitrary-length dependencies
  • Ignores that truncated training windows never span the gap
  • Confuses gradient flow with the ability to retrieve a stored fact
  • Assumes a gate value just under 1 is lossless over thousands of steps

context