skip to content

Why does encoder-decoder RNN translation quality collapse as the source sentence gets longer?

level: middleimportance: must knowfreq 68%

answer

  1. fixed budget, growing content
  2. flat, then a cliff, not a slope
  3. written before anything is demanded of it
  4. the end of the source is best remembered
  5. detail goes first, gist survives

basics

~20 s

One fixed-size vector carries a fixed information budget, so a longer source forces the encoder to discard more, and to choose what to keep before the decoder emits a word. Quality stays flat, then falls steeply.

solid answer

~50 s

Everything the decoder will ever know about the source has to fit in a single vector of fixed width - on the order of a thousand floats - written once, when the encoder finishes reading. Short sentences fit inside that budget, so quality is roughly flat. Past a point, commonly reported around thirty source words for early English-to-French systems, the source carries more detail than the budget holds and BLEU falls off a cliff. Two things compound it. The encoder must decide what to keep *before* it knows which target word will be needed, so it cannot preferentially retain what matters. And a recurrent state has a recency bias: material read near the end of the source leaves a stronger trace than material read at the start. The symptoms are characteristic - dropped clauses, wrong or invented names and numbers, and generic, plausible-sounding output.

go deeper

for a junior

Be ready to state the core fact: one fixed-size vector must hold the whole source, so longer sources lose more. Naming the flat-then-falling shape of the quality curve is enough at this level.

for a middle

Explain all three mechanisms - fixed capacity, a summary committed before the decoder makes any demand, and the recency bias of a state that is overwritten every step - and say which kinds of content are lost first.

for a senior

Show how you would confirm the diagnosis: bucket held-out data by source length, look for flat-then-cliff, and run a truncated-input control to separate lost information from genuine difficulty before proposing an architecture change.

for a principal

Own the argument that this is an architectural limit, not a data or tuning limit, and be able to say what that implies about where a team should spend: more data and longer training move the level, only changing the channel changes the shape.

## The information budget Start with the arithmetic of the constraint. The only channel between encoder and decoder is one fixed-size vector - say a thousand floating-point numbers. That budget is set at design time and does not move. A six-word source has to be represented in a thousand numbers; so does a sixty-word source; so does a two-thousand-word article for a summarizer. The amount of content the model must carry grows with input length, and the space it has to carry it in does not. Up to some length the budget is not binding: a thousand numbers is plenty for a short clause, and quality is limited by ordinary things like vocabulary coverage and training data. Past that length it binds, and every extra word of source displaces something. That is why the length-quality curve is not a gentle slope - it is flat and then falls. For early English-to-French systems built on a single roughly thousand-dimensional context vector, the reported breakpoint sat around thirty source words. ## Three mechanisms, not one **Capacity.** The obvious one. Fixed width, growing content, so information is lost. Empirically what goes first is fine-grained, positionally-local detail - a rare surname, a number, a subordinate clause - while global properties survive: the topic, the register, roughly how long the output should be. That is exactly the failure signature people report. Output that reads fluent, sits on the right subject, and quietly gets a name or a figure wrong, or drops a clause entirely. **Commitment before demand.** The encoder writes the vector when it finishes reading, and it must do so without knowing what the decoder will need. When the decoder is about to emit the target word for source position 4, it would like the details of position 4; but the encoder had to allocate its budget across all positions in advance. There is no way to spend more of the budget on the part that turns out to matter, because at write time nothing knows which part that is. A fixed summary is a *lossy compression chosen blind*. **Recency and path length.** A recurrent state is overwritten at every step: `h_t = f(h_(t-1), x_t)`. Information from step 1 has to survive being passed through the same nonlinearity `n` times before it reaches `h_n`. Gating in an LSTM or GRU makes that survival far more likely than in a plain RNN, but it does not make it free. The practical consequence is a bias toward the end of the source: the last things read are the best represented. Symmetrically, during training the gradient signal that would teach the model to preserve early-source detail has to travel the longest way back through the recurrence, so that is the part that is learned worst. ## What it looks like when you measure it The diagnostic is to bucket a held-out set by source length and score each bucket separately, rather than reporting one aggregate number. A bottleneck-limited model shows a flat region followed by a steep decline. If instead every bucket is equally bad, the problem is not the bottleneck - suspect data, vocabulary handling or optimization. A useful control is to feed truncated versions of the long inputs: if quality on the truncated prefix is fine while the full-length input fails, the model can handle the content and is losing it to the budget, not to difficulty. Another way to see it: watch what a summarizer does with a long article. Asked to compress two thousand words, it tends to produce a summary that is topically right, drawn disproportionately from the article's opening and closing, and specific in the wrong places - it invents a plausible number rather than reporting the real one, because the real one is no longer in the vector. ## Why more data does not rescue it A common wrong answer is that the model simply has not seen enough long sentences. More data does raise the whole curve - the model gets better at everything - but the cliff stays, because the constraint is architectural rather than statistical. No amount of data lets a thousand numbers represent an unbounded amount of source detail. The same is true of training longer or of a better optimizer: they move the level, not the shape. ## The fix, in one line The direction the field took was to stop insisting on a single summary written once, and let the decoder consult the source at each output step instead - the mechanics of that scoring belong to alignment-based attention. In an interview, name the diagnosis crisply and then name that direction; the point being tested is that you understand *why* one fixed vector was the thing that had to go.

  • Early systems reversed the source sentence before encoding it. Why did that help?
    Reversing puts the first source words last, so they sit closest to the handoff and to the first target words that depend on them. Average source-target distance is unchanged, but the *minimum* lag for the earliest correspondences shrinks, which makes those dependencies far easier to learn and gets the output started correctly. It is a symptom-level hack: it redistributes which part of the source is best preserved rather than enlarging the budget.
  • What does the encoder tend to discard first as the source grows?
    Fine-grained, positionally-local detail: rare names, numbers, and subordinate clauses, especially from earlier in the source. Global properties survive better - topic, register, rough output length. That is why the failure mode is fluent, on-topic output with wrong specifics rather than obvious gibberish.
  • Would training on far more long sentences remove the degradation?
    No. More data lifts the whole length-quality curve but leaves the cliff, because the limit is a fixed-width channel rather than an estimation problem. The same goes for training longer or switching optimizer. Only changing what crosses the encoder-decoder boundary changes the shape of the curve.

saying these in an interview costs you the question

  • Blames vanishing gradients alone and never mentions fixed capacity
  • Says more training data would fix the long-sentence cliff
  • Claims the context vector grows with input length
  • Assumes degradation is linear rather than flat-then-cliff
  • Says an LSTM's gating removes the length limit entirely

context