skip to content

Teacher Forcing

Feeding the true previous token while training makes the decoder converge fast but never rehearse its own mistakes, so errors compound at inference. Interviewers probe that train/inference mismatch.

on this pageshow

questions

3

In seq2seq training, what is teacher forcing and why is it the default?

level: juniorimportance: must knowfreq 70%

answer

  1. what does the decoder eat each step
  2. ground truth, not its own guess
  3. inputs are the target shifted by one
  4. training only; generation has no teacher
  5. buys stability, sells train/test match

basics

~20 s

Teacher forcing feeds the decoder the ground-truth previous token at every training step instead of its own prediction. It keeps training stable and fast, because each step is conditioned on a correct prefix and no generation loop is needed.

solid answer

~40 s

An autoregressive decoder predicts token t from the tokens before it. Under teacher forcing, the input at step t during training is the ground-truth token at position t-1 of the target, not whatever the model itself predicted at step t-1. The loss is still next-token cross-entropy at every position, so the model is trained to estimate `p(y_t | y_<t, x)` on true prefixes. This is the default for three reasons: the decoder's whole input sequence is known before the forward pass, so no sampling or argmax loop is required inside training; there is no non-differentiable discrete choice in the path; and early-training garbage predictions cannot poison every later step, which would otherwise leave the model with almost no usable learning signal. The cost is that training and generation now see different input distributions.

go deeper

for a junior

Be ready to say in one breath that the decoder is fed the previous ground-truth target token during training and its own previous token during generation, and that the labels are the target shifted by one.

for a middle

Explain the mechanics: why no gradient has to flow through a discrete token choice, why known inputs make the forward pass cheaper, and why masking future positions is required to avoid leaking the label.

for a senior

Show that you treat the train/generate asymmetry as an operational fact, not trivia: teacher-forced validation numbers measure the teacher-forced conditional, so your evaluation plan has to include a self-fed run before anyone trusts the model.

for a principal

Own the trade explicitly. Teacher forcing buys a stable, well-conditioned, parallelisable training problem at the price of a distribution shift you must handle elsewhere, and you should be able to argue where in the stack that shift is cheapest to absorb.

## The setting A sequence-to-sequence decoder is autoregressive: it produces one target token at a time, and each prediction is conditioned on the source input `x` and on the target tokens already emitted. Formally it models the factorisation `p(y | x) = prod_t p(y_t | y_1..y_(t-1), x)`. That definition leaves one practical question open: at training time, where do the tokens `y_1..y_(t-1)` that the decoder consumes actually come from? There are two possible answers, and they are not the same thing. - **Teacher forcing**: feed the ground-truth target token from position `t-1`. The prefix the decoder conditions on is always the true one from the dataset. - **Free running**: feed whatever the model itself produced at step `t-1`. The prefix is the model's own output. Teacher forcing is the near-universal default for training. Free running is what necessarily happens at generation time, because no ground truth exists then. The whole topic exists because those two regimes differ. ## What the loss looks like With teacher forcing, the target sequence is shifted by one to form the decoder's inputs. If the target is `<start> a cat sat <end>`, the decoder inputs are `<start> a cat sat` and the labels at those four positions are `a cat sat <end>`. The per-position loss is cross-entropy between the predicted distribution over the vocabulary and the one-hot true token, summed or averaged over positions. Nothing exotic: it is ordinary maximum likelihood on the factorised conditional. An important detail juniors sometimes get backwards: the input at step `t` is the ground-truth token at position `t-1`, never the token at position `t`. Feeding position `t` would be label leakage — the model would learn the identity function and collapse. Any decoder that can see the whole target at once must therefore also mask future positions so that position `t` cannot attend to `t` or beyond. ## Why it is the default **Learning signal.** Early in training the model's own outputs are noise. If the decoder is fed its own noise, then by step five the conditioning context is meaningless and the gradient at step five teaches almost nothing about the real task. Ground-truth prefixes guarantee that every position is a well-posed prediction problem from the very first update, so the model converges dramatically faster. **No discrete sampling in the training path.** Choosing the model's own token requires an argmax or a draw from a categorical distribution. Both are non-differentiable, so gradients do not flow back through the fed token; you would need a surrogate estimator to train through it. Teacher forcing sidesteps the problem entirely — the fed tokens are constants from the dataset. **Known inputs, cheaper computation.** Because the decoder's entire input sequence is known before the forward pass, training does not need a step-by-step generation loop. For a decoder without recurrence this makes all positions computable in one parallel pass; even for a recurrent decoder, whose hidden state is still sequential, it removes the per-step sampling round trip and makes batching over a whole target sequence straightforward. **Stability.** The loss surface is far better behaved when the conditioning context is fixed data rather than a moving function of the current parameters. Free-running training is closer to a moving-target problem: change the parameters, and the inputs the model sees change too. ## The asymmetry that is left behind The price is a train/generate mismatch. During training the model only ever answers the question "given a correct prefix, what comes next?" During generation it is asked "given the prefix I myself produced, what comes next?" Those are different conditional queries, and the second one is never practised. The model's estimates on self-generated prefixes are therefore whatever the function happens to extrapolate to, not something the data constrained. This is called exposure bias, and it is the reason a decoder can look excellent on teacher-forced validation numbers and still produce degenerate text when it runs on its own. That asymmetry is not a bug in teacher forcing so much as its trade: you buy a tractable, stable, well-conditioned training problem, and you pay for it with a distribution shift at inference that has to be managed separately. ## Practical notes - Teacher forcing is a **training-time** technique. At generation time there is nothing to force with, so the decoder always feeds itself. - The same mechanism applies to any autoregressive target — text, molecular strings, symbol sequences — not just translation. - A validation loss computed with teacher forcing measures the same quantity as the training loss. It is a genuine measurement, but of the teacher-forced conditional, not of generation quality. Any evaluation you want to trust as a proxy for deployment must let the model feed itself.

  • Why is teacher forcing not used at generation time?
    Because there is no ground truth to feed. At generation the model is given only the source input and whatever it has already emitted, so the decoder is necessarily free-running. Teacher forcing is a training-time convenience that exists only while labelled targets are available; the moment the model is deployed, the crutch is gone.
  • What would go wrong if you fed the decoder the ground-truth token at position t rather than t-1?
    That is label leakage. The model would see the answer it is being asked to predict and could reach near-zero loss by copying its input, learning nothing about the sequence. This is exactly why any decoder that consumes the whole target in one pass must mask future positions, so position t is conditioned only on positions strictly before it.
  • How does teacher forcing change what a decoder's training loop has to compute?
    The decoder's entire input sequence is known before the forward pass, so training needs no step-by-step generation loop and no argmax or sampling inside the graph. A non-recurrent decoder can then score every position in one parallel pass; a recurrent one still unrolls its hidden state sequentially but avoids the sampling round trip at every step.

saying these in an interview costs you the question

  • Says the decoder input at step t is the ground-truth token at position t
  • Thinks teacher forcing is also applied at generation time
  • Claims teacher forcing changes the loss function rather than the inputs
  • Believes it is a regularisation technique that reduces overfitting
  • Cannot state any downside, treating it as free improvement

context

open as a page

Why does a caption decoder score well teacher-forced but degenerate when it feeds itself?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Exposure bias. The decoder was only ever trained on correct prefixes, so once it emits one wrong token it is conditioning on a prefix the training data never contained, and errors compound. Teacher-forced perplexity never measures that regime.

open as a page

How does scheduled sampling mitigate exposure bias, and what does it cost?

level: middleimportance: should knowfreq 40%

basics

~20 s

Scheduled sampling flips a coin at each decoder step: feed the ground-truth previous token, or the model's own. The self-feeding probability is annealed upward during training, so the decoder practises recovering from its own mistakes.

open as a page