skip to content

Why does a sentiment RNN learn the final clause of a review but ignore the opening sentence?

level: seniorimportance: should knowfreq 46%

answer

  1. the last clause gets the signal
  2. gradient decays with lag, not with length
  3. the loss curve hides it completely
  4. sum over lags stays healthy
  5. measure sensitivity against lag

basics

~20 s

Gradients from early time steps decay geometrically on the way back, so the opening sentence gets almost no learning signal. Recent steps keep full-size gradients, so the model quietly becomes a short-memory model reading the ending.

solid answer

~40 s

This is long-range credit assignment dying first. The gradient connecting the label to a hidden state k steps earlier is a product of k per-step Jacobians, so it decays exponentially in the lag: the last clause ('...but the ending ruined it') gets a full-size signal and the opening sentence gets essentially none. The gradient on the shared recurrent weights is a *sum* over lags, and the short-lag terms are healthy, so the loss falls normally and nothing looks broken. The model fits whatever a few dozen steps of memory can explain, which on review text is quite a lot. To confirm it, I'd bucket the evaluation set by where the decisive cue sits and measure the gradient norm at input step k against the lag — I expect a clean exponential decay.

go deeper

for a junior

Recall that a recurrent model's learning signal weakens the further back in the sequence you go, so the end of an input influences training far more than the beginning.

for a middle

Explain why the aggregate gradient stays normal-sized: the shared weight's gradient sums over all lags, and the short lags are unaffected. That is the reason nothing in the training log flags the failure.

for a senior

Show the diagnostic instinct. Propose position-bucketed evaluation, a sensitivity-versus-lag plot, and occlusion of the early region, and state what result would confirm or refute a reach limit before you touch the model.

for a principal

Own the monitoring gap: a silent failure that survives every standard training metric needs a slice-based evaluation contract, not more loss curves. Be ready to argue what long-range evaluation slices a team should maintain permanently.

## The symptom A recurrent sentiment classifier trains cleanly: the loss falls, the gradient norm looks ordinary, validation accuracy is respectable. But inspection shows it is essentially reading the tail of the review. Reviews that end '...but the ending ruined it' are classified negative even when the first four sentences are glowing, and reviews whose verdict is stated up front and then qualified are classified badly. Nothing in the training log says anything is wrong. ## Why the early steps get no signal The gradient of the loss with respect to a hidden state k steps in the past is a product of k per-step Jacobians, each of which is the shared recurrent matrix scaled by a diagonal of activation derivatives. Multiplying that many near-identical factors makes the magnitude exponential in the lag. With an effective per-step factor below 1, the contribution from a token 200 steps back is many orders of magnitude below the contribution from a token 5 steps back. Gradient descent can only strengthen a pathway it receives signal about. A dependency whose gradient contribution is 1e-9 relative to its neighbours is, for optimisation purposes, not there. So the parameters that would make the opening sentence matter never move in a coordinated direction — not because the model decided the opening was uninformative, but because it never got told. ## Why the loss looks healthy anyway Two things conspire to hide it. **The gradient is a sum over lags.** Because one recurrent matrix is shared across all steps, its gradient aggregates contributions from every (loss, source-step) pair. The lag-1 to lag-30 terms are full size. The total gradient is therefore normal in magnitude and points somewhere sensible — it just contains no information about long-lag structure. Monitoring the global gradient norm will never surface this. **Short-range cues carry a lot of the label.** Natural review text is redundant: sentiment words cluster, and the last clause is often the summary judgement. A short-memory model can get most of the way on that alone. So the loss curve and the headline metric both reward the shortcut, and the model converges to it. The failure only shows up on the slice of examples where the shortcut is wrong. This asymmetry is worth stating explicitly in an interview: exploding gradients are a loud failure — a loss spike, an overflow, a wrecked parameter vector — while vanishing gradients are silent. The training log cannot distinguish 'learned the dependency' from 'never received a gradient about it'. ## How to confirm it Do not diagnose from the loss curve. Three checks that produce evidence: 1. **Position-bucketed evaluation.** Partition the evaluation set by where the decisive cue sits — reviews whose sentiment is determined in the first sentence versus in the last. If accuracy collapses on the early-cue bucket and is fine on the late-cue bucket, the model's memory horizon is the story. 2. **Sensitivity as a function of lag.** Measure the norm of the loss gradient with respect to the input representation at step k, and plot it against the lag from the loss. A clean exponential decay is a direct measurement of the effective horizon: the lag at which the curve hits the noise floor is roughly how far back the model can see. 3. **Occlusion.** Blank out the opening sentence at inference and see whether predictions change at all. If they are numerically identical, the model was never using it. ## What the horizon actually is The effective memory horizon is where a step's gradient contribution falls below the other terms in the sum. It is set by the per-step contraction factor, not by the sequence length you feed in. Feeding longer sequences does not extend it; it just adds steps whose gradient contributions are already negligible. That is the trap in 'we'll fix it by using the whole document' — the extra tokens cost compute and change nothing about what the optimiser can learn. It is also worth being precise that this is a *training* pathology, not an inference one. The forward pass really does carry information from step 1 into step 200 — imperfectly, but it is there. What is missing is the learning signal that would have taught the network to preserve the *useful* part of it. The network is capable of long memory in principle; it was never trained into it. ## The interview framing Name the mechanism (exponential decay of the per-lag gradient), explain why the aggregate loss and gradient norm hide it (the sum over lags is dominated by recent steps and short-range cues fit most of the data), and then propose a measurement rather than an opinion. Candidates who jump straight to 'the dataset is biased toward endings' without checking gradient reach are guessing.

  • Would feeding longer reviews help the model use the opening sentence?
    No. The horizon is set by the per-step contraction factor, not by the input length. Adding steps only adds contributions that were already negligible, so you pay more compute for exactly the same reachable window. Longer inputs address a truncation problem, not a decay problem.
  • Does a vanishing gradient mean the forward pass loses the early information too?
    Not necessarily. The forward pass does carry state from step 1 to step 200, imperfectly but genuinely. What is missing is the backward signal that would have taught the network to preserve the useful part. It is a training pathology, and the distinction matters because it rules out 'the hidden state is too small' as the automatic explanation.
  • How would you measure the model's effective memory horizon on this data?
    Plot the norm of the loss gradient with respect to the input representation at step k against the lag from the loss. The curve should decay roughly exponentially; the lag where it reaches the noise floor is the horizon. Cross-check with occlusion: blank out tokens at various lags and see where predictions stop changing.

saying these in an interview costs you the question

  • Concludes the dataset is biased without measuring gradient reach
  • Treats a falling training loss as proof the whole sequence is used
  • Blames hidden-state capacity and never considers gradient decay
  • Proposes reversing the sequence as a fix rather than a shuffle of the same limit
  • Expects the global gradient norm to reveal the problem

context