skip to content

What does the causal mask do in a decoder-only transformer, and what breaks without it?

level: middleimportance: must knowfreq 72%

answer

  1. Look-ahead is forbidden
  2. Lower-triangular weight matrix
  3. Applied before the softmax, not after
  4. Negative infinity becomes exactly zero weight
  5. One pass trains every next-token prediction

basics

~20 s

The causal mask sets attention scores for future positions to negative infinity before the softmax, so each position can attend only to itself and earlier tokens. Without it, training leaks the answer: a position can see the token it is meant to predict.

solid answer

~50 s

In a decoder-only stack, every position is trained to predict the next token, and the whole sequence is trained in a single forward pass. That only works if position i is forbidden from seeing positions after i. The causal mask enforces it: before the softmax, entries of the score matrix above the diagonal are set to negative infinity, so after exponentiation their weight is exactly zero. The resulting weight matrix is lower-triangular — position 4 may attend to 1 through 4 and never to 5. Drop the mask and training collapses in a specific, recognizable way: the loss falls almost to zero because each position can simply copy the next token from its own input, yet the model generates nonsense at inference, where future tokens genuinely do not exist. The mask is also what makes the train and inference conditions match, since generation is strictly left-to-right. An encoder used for classification deliberately omits it, because there every position is allowed to see the whole input.

code

python · 9 lines
python
import numpy as np

n = 5
allowed = np.tril(np.ones((n, n), dtype=bool))
scores = np.zeros((n, n))                 # pretend all matches are equal
scores[~allowed] = -np.inf
w = np.exp(scores - scores.max(-1, keepdims=True))
w /= w.sum(-1, keepdims=True)
print(np.round(w, 2))                     # lower-triangular, rows sum to 1

go deeper

for a junior

Know that in a text-generating transformer each position may only look at itself and earlier tokens, and that this is enforced by a mask applied before the softmax. Be able to say why looking ahead would be cheating.

for a middle

Explain the mechanics: upper-triangular entries set to negative infinity, weights becoming exactly zero, rows still summing to one. Then connect it to parallel training — one forward pass supplying every next-token prediction in the sequence.

for a senior

Be ready to diagnose leakage from symptoms: implausibly fast loss decline with incoherent free-running generation, and evaluation that hides the bug because it is also teacher-forced. Know how causal and padding masks are combined in real batching code.

for a principal

Own the framing that the mask defines the model's objective, not just its plumbing — bidirectional and causal stacks are different products of the same architecture. Be able to argue which one a given application needs and what it costs to pick wrong.

## What the mask is Self-attention as defined lets every token attend to every token. For a model that generates text left to right, that is wrong: predicting token 5 while looking at token 5 is not a prediction. The causal mask (also called the autoregressive or look-ahead mask) is the restriction that turns unrestricted attention into next-token prediction. Mechanically it is applied to the score matrix, after the scaled dot product and before the softmax. For a sequence of length n, build a lower-triangular pattern: entry (i, j) is allowed if j <= i and forbidden otherwise. Forbidden entries are set to negative infinity. The softmax exponentiates, e to the minus infinity is zero, so those positions receive exactly zero weight and contribute nothing to the weighted sum of values. Each output row is a mixture over the current token and its predecessors only. Note it is applied to the scores, not to the weights after the softmax. Zeroing weights after normalization would leave the remaining weights summing to less than one; masking before the softmax renormalizes over the allowed positions automatically. ## Why it makes parallel training possible This is the part worth saying explicitly in an interview, because it is the real payoff. Given a training sequence of n tokens, the model is asked to predict token 2 from token 1, token 3 from tokens 1-2, and so on — n-1 prediction problems. Without masking you would have to run n-1 separate forward passes over progressively longer prefixes. With masking, one forward pass over the full sequence computes all of them at once, because row i of the attention output was computed as if only the first i tokens existed. The loss is then averaged over all positions. That single trick is a large part of why transformers train efficiently at scale. ## What breaks if you drop it The failure is diagnostic and easy to recognize. Training loss falls implausibly fast, often approaching zero within a small fraction of an epoch, because the task has become trivial: position i can attend to position i+1 and read off the answer it is being scored on. Perplexity on the training data looks superb. Then generation produces incoherent output, because at inference time there is no token i+1 to copy — the model is being asked to do a task it never actually learned. A held-out evaluation that also runs teacher-forced will *not* catch this, since it leaks the same way; the tell is the gap between teacher-forced metrics and free-running generation. A subtler version of the same bug appears when the mask is off by one — allowing position i to see i+1 — or when it is applied in some layers and not others. The symptom is the same shape: too-good loss, bad generation. ## Causal mask versus padding mask These are different masks that are often combined into one tensor, and conflating them is a common interview stumble. The causal mask depends only on position and is identical for every sequence in the batch. A padding mask depends on the data: when sequences of different lengths are batched together and short ones are padded, the padding positions must be masked out so real tokens do not attend to filler. In practice the two boolean patterns are combined and applied together, but they exist for unrelated reasons — one enforces the direction of information flow, the other cleans up a batching artifact. ## Where causal masking does not apply Encoder stacks used for embedding or classification are bidirectional by design: to represent a sentence for retrieval or sentiment, every token should see the whole sentence, so no causal mask is applied. The same architecture with the mask removed is a fundamentally different objective. This is one of the clearest ways to state the encoder/decoder distinction in terms of mechanism rather than diagram. ## Consequences at inference Because the mask enforces that a token's representation depends only on itself and its predecessors, appending a new token never changes the representations already computed for earlier ones. Generation is therefore a strictly left-to-right process, and the training condition matches the deployment condition exactly. That property is also what a well-behaved implementation relies on when it processes a long prompt and then emits tokens one at a time. ## The one-line version "Set the upper triangle of the score matrix to negative infinity before the softmax. Position i then attends only to positions 1 through i, which is what lets one forward pass train every next-token prediction in the sequence — and without it the model just copies the answer and produces garbage when generating."

  • Why is the mask applied to the scores before the softmax rather than zeroing weights afterwards?
    Masking before the softmax renormalizes over the allowed positions automatically: the exponentials of the forbidden entries are zero, and the remaining weights still sum to one. Zeroing after the softmax would leave each row summing to less than one, shrinking the output vector by an amount that varies with position — early tokens would be scaled down far more than late ones.
  • How does a causal mask differ from a padding mask?
    The causal mask depends only on position and is the same for every sequence: it forbids attending to the future. A padding mask depends on the data — it blocks attention to filler positions added so that variable-length sequences can be batched. They are frequently combined into a single boolean tensor, but they solve unrelated problems, and one does not substitute for the other.
  • Why do encoder models used for embeddings or classification omit the causal mask?
    Their objective is to represent the whole input, not to predict the next token. For sentiment, retrieval or entailment, a token's representation is better when it can see everything on both sides, so attention is bidirectional. Removing the mask is not an optimization on the same task — it changes the task the model is trained to do.
  • What would you check first if training loss dropped near zero within the first few hundred steps?
    Suspect label leakage through masking. Verify the mask is lower-triangular and not off by one, that it is applied in every attention layer rather than some, and that it is combined correctly with any padding mask. Then compare teacher-forced evaluation against free-running generation: leakage looks excellent under teacher forcing and incoherent when the model must generate.

saying these in an interview costs you the question

  • Saying the mask is applied to attention weights after the softmax
  • Claiming masking is needed at inference only, not during training
  • Confusing the causal mask with the padding mask for variable-length batches
  • Believing a masked model still trains one position per forward pass
  • Thinking a bidirectional encoder simply forgot to include the mask

context