Why does a caption decoder score well teacher-forced but degenerate when it feeds itself?
answer
- compare training inputs with generation inputs
- the model never sees its own mistakes
- one bad token, unseen conditioning state
- errors grow along the sequence
- perplexity measured with a crutch
basics
~20 sExposure bias. The decoder was only ever trained on correct prefixes, so once it emits one wrong token it is conditioning on a prefix the training data never contained, and errors compound. Teacher-forced perplexity never measures that regime.
solid answer
~50 sTraining and generation ask the decoder two different questions. Trained with teacher forcing, it estimates `p(y_t | true prefix, x)`; at generation it is queried at `p(y_t | its own prefix, x)`. Those coincide only while the model is right. One wrong token at step 3 shifts the decoder into a state distribution it never saw during training, where its estimates are unconstrained extrapolation — so the next token is more likely wrong too, and the error rate grows along the sequence. Repetition loops like "a man a man a man" are a stable attractor of that off-distribution regime: the repeated prefix keeps reinforcing itself. Teacher-forced validation perplexity cannot see any of this, because it hands the model a correct prefix at every step. The fix in evaluation is to run free-running generation on held-out inputs and score whole sequences, then watch the gap between the two numbers.
go deeper
Recall the name and the one-line cause: the decoder trained only on correct prefixes, so it has never seen the self-generated prefixes it must handle at generation time.
Explain the mechanism step by step — a wrong token moves the decoder into a prefix region with no training support, so the next estimate is extrapolation and the error rate grows along the sequence.
Demonstrate the diagnosis: rule out a harness bug and a train/serve preprocessing mismatch first, then quantify the gap between teacher-forced perplexity and a free-running sequence-level score, and show you track both during training.
Frame it as an evaluation-design problem your organisation owns. Decide which metric gates a release, insist that the likelihood-versus-generation gap is reported as a first-class number, and judge how much the mismatch actually costs this product before spending on a fix.
## The two conditional queries An autoregressive decoder is a single function, but training and deployment probe it at different inputs. - **Training (teacher forced):** the model is asked, at every position, what follows a *ground-truth* prefix. The prefixes it sees are drawn from the data distribution of real target sequences. - **Generation (free running):** the model is asked what follows *its own* prefix. Those prefixes are drawn from the model's own output distribution. Maximum likelihood under teacher forcing fits the first query. Nothing in the objective constrains the second. This mismatch is **exposure bias**: the model was never exposed to the inputs it will actually face. ## Compounding error, stated precisely Suppose the decoder emits a wrong token at step 3. Every later step now conditions on a prefix containing that token. If such a prefix has essentially zero probability under the training distribution of target sequences, the model has no data-driven constraint there at all — its output is whatever the learned function extrapolates to. The conditional estimate at step 4 is therefore less reliable than at step 3, which makes another wrong token more likely, which pushes the prefix further from anything seen in training. The per-step error probability is not constant along the sequence; it grows, because the conditioning context itself is drifting. That is the compounding part. It is why the degradation is worst at the tail of long generations and why short outputs often look fine. ## The symptoms **Degenerate repetition.** An image-captioning decoder that scores well with ground-truth prefixes can collapse to "a man a man a man" as soon as it feeds itself. Once a repeated bigram is in the prefix, the strongest local pattern in that off-distribution context is the repetition itself, and the loop is self-sustaining. The model is not confused about the image; it is stuck in a region where its own history dominates the conditioning. **Unrecoverable structural violations.** A SMILES molecule generator trained on valid strings emits an unbalanced opening parenthesis at step 12. Every valid continuation in the training data closed the brackets it opened, so the model has never learned what to do from a state where the count cannot be repaired, and the rest of the string is garbage. The constraint is global, but the model only ever learned it from prefixes that already satisfied it. **Drift and hallucinated detail.** In longer generation the output stays fluent but slowly stops corresponding to the source, because fluency is locally reinforced by the model's own prefix while grounding in the source input is not. ## Why your metrics missed it Teacher-forced validation loss and perplexity are computed exactly like the training loss: at every position the model is handed the correct prefix. They are honest measurements of the teacher-forced conditional and they are useful — they will catch under-fitting, a broken data pipeline, or over-fitting to the training set. What they cannot do is measure behaviour on self-generated prefixes, since by construction the model never generates anything during that evaluation. A model can therefore improve its validation perplexity for several epochs while its free-running output gets worse. The operational fix is to make free-running evaluation a first-class metric, not an afterthought: - Run the decoder autoregressively on held-out inputs and score the *whole sequence* against the reference with a sequence-level metric appropriate to the task, plus cheap structural checks such as validity rate, output-length distribution, and repeated-n-gram rate. - Track it on the same cadence as the loss so you can see the two curves diverge instead of discovering it after a release. - Report the gap between teacher-forced perplexity and the free-running score as its own quantity. A widening gap is the signature of exposure bias; a model with a small gap is one whose likelihood numbers you can trust as a proxy. ## Diagnosing it in a real run When someone reports "great validation loss, terrible outputs", separate three candidate causes before blaming exposure bias: 1. **An evaluation harness bug** — the generation path silently still receives ground-truth tokens, or feeds the wrong shift, or starts from the wrong token. Verify that the free-running path is genuinely self-fed. 2. **A training/serving discrepancy** — different tokenisation, different special tokens, or different preprocessing between the two paths. 3. **Genuine exposure bias** — which you confirm by seeding generation with a correct prefix of increasing length and watching how far the model gets before degenerating. If a longer forced prefix reliably buys more good tokens, and quality falls off after the forcing stops, the diagnosis is the distribution shift rather than a plain capability failure. ## Perspective Exposure bias is a real and well-defined mismatch, but its practical severity varies. It is worst where outputs are long, where the target has global structure that a single early token can violate, and where the model is small or the data thin. It shrinks — though it does not vanish — as models get more capable, because a stronger model makes fewer early mistakes and generalises better to prefixes it has not seen. That is why the first response to a degeneration report should be measurement of the gap, not an immediate change of training objective.
- How would you confirm the problem is exposure bias rather than a bug in the generation path?First verify the generation path is genuinely self-fed and uses the same tokenisation and special tokens as training — a harness that silently still supplies ground truth is a common cause. Then seed generation with correct prefixes of increasing length. If a longer forced prefix reliably buys more good tokens and quality falls off after forcing stops, the distribution shift is the cause.
- Why is degradation usually worse at the end of a long generation than at the start?Because the conditioning context drifts cumulatively. Early steps still condition on a near-correct prefix, so their error rate is close to the teacher-forced one. Each mistake pushes the prefix further from anything the training distribution contained, so the per-step error probability rises as the sequence grows. Short outputs often look fine for exactly this reason.
- Does exposure bias disappear as models get larger?It shrinks but does not vanish. A stronger model makes fewer early mistakes and extrapolates better to prefixes it has not seen, so the gap between teacher-forced likelihood and free-running quality narrows. The mismatch in what the objective optimises is still there, which is why long-horizon and structurally constrained outputs remain the failure cases.
saying these in an interview costs you the question
- Blames overfitting when training and validation loss both look healthy
- Says teacher-forced perplexity is simply a wrong or broken metric
- Thinks repetition loops mean the learning rate was too high
- Cannot explain why later tokens are worse than earlier ones
- Proposes only more training data without measuring the free-running gap