skip to content

Why must a batch of variable-length sequences be padded, and what does the mask do?

level: juniorimportance: must knowfreq 72%

answer

  1. one array, one time dimension
  2. pads are fed through, not skipped
  3. a zero input still updates the state
  4. loss, pooling, final-state readout
  5. divide by summed mask, not padded width

basics

~20 s

A batch must be one rectangular array, so shorter sequences are padded out to the longest length. The mask marks which steps are real, keeping pad steps out of the loss, out of pooling, and out of the state you read.

solid answer

~50 s

A batch of sequences is stored as one array with a single time dimension, so every sequence in it must show the same number of steps; padding fills the shorter ones up to the batch maximum with a reserved pad id or a zero vector. Those pad steps are still fed through the model, so they must be neutralised explicitly, and that is what the mask is for: a per-step boolean, `1` for real and `0` for pad, applied in three places. The per-step loss is dropped at pad positions and averaged over the summed mask, not over `batch * T`. Pooling over time becomes `sum(x_t * m_t) / sum(m_t)`. And the "final" state means the state after the last real step, not after the trailing pads. Zeros are no substitute for a mask, because a zero input still drives a recurrent update.

go deeper

for a junior

Be ready to say why a batch has to be rectangular and what the mask marks. Naming the three places pad steps must be excluded - the loss, any pooling over time, and the final-state readout - is enough here.

for a middle

Explain why a zero input is not a no-op for a recurrent update, and write the masked mean out loud: sum of masked vectors divided by the summed mask, not by the padded width.

for a senior

Show how you would catch this in production: representation norms correlating with sequence length, metrics that move when the batch's padded width changes, loss scale that shifts with reshuffling.

for a principal

Own the convention. Per-token versus per-sequence loss normalisation decides whether long sequences dominate the objective, so fix it once for the team and make training and evaluation agree on it.

## Why padding exists at all A recurrent model consumes one step at a time, but training runs on batches, and a batch is a single array with shape roughly `(batch, time, features)`. That array has exactly one time dimension, so every sequence inside it must present the same number of steps. Real data does not cooperate: a batch of customer-support chat turns can run from a 4-token "ok thanks" to a 900-token pasted error report. Padding resolves this by filling every shorter sequence out to the batch's longest length with a designated pad value — a reserved pad id for token sequences, a zero vector for continuous features such as audio frames. The consequence that trips people up is simple: the pad steps are *inside the array*, so the model runs over them like any other step. Nothing about padding is self-cancelling. Whatever you do not exclude on purpose, you have included. ## What a mask actually is A mask is a companion array over the time dimension, `1` where the step is real and `0` where it is pad. It changes no arithmetic by itself; it is an instruction used at three specific places. 1. **The loss.** For per-step targets, every position produces a loss term, including positions whose target is a filler. Masking multiplies those terms by the mask and normalises by `sum(mask)` — the count of real steps — rather than by `batch * T`. Without it the model is trained to predict a fake label many times per sequence. 2. **Aggregation over time.** Mean or max pooling over the time axis must see only real steps. A masked mean is `sum(x_t * m_t) / sum(m_t)`. 3. **The state readout.** "The final hidden state" has to mean the state after the last *real* step of each sequence, which is at a different index for each row of the batch. ## Zeros are not neutral The standard misconception is that padding with zeros makes a mask unnecessary, because zero contributes nothing. For a recurrent cell that is false. The update is `h_t = tanh(W h_(t-1) + U x_t + b)`. Setting `x_t = 0` leaves `h_t = tanh(W h_(t-1) + b)` — still a nonlinear map applied to the state, still a change at every pad step, and over a long pad run the state drifts toward a fixed point of that map, forgetting whatever the real tokens put there. For token inputs it is worse: a pad id indexes an embedding row like any other id, and that row is a learned vector unless you deliberately hold it at zero. Pads change the state, produce outputs, and produce loss terms unless you stop them. ## The silent-averaging failure Take the support-chat batch padded to 900 steps and mean-pool each sequence into a sentence vector with no mask. A 30-token turn contributes 30 real vectors and 870 pad vectors, so over 96% of its "sentence representation" is the pad embedding. Every short turn collapses toward the same point, and the classifier on top ends up separating sequences by length instead of by content — an error that is invisible in the loss curve and shows up only as short inputs being systematically misclassified. The second half of that failure is the denominator. Suppose you do mask, zeroing the pad vectors, but still divide the sum by 900. The 30-token turn's pooled vector is scaled by `30/900`, so the *magnitude* of the representation encodes length. A downstream layer will exploit that happily, and the whole thing shifts the moment a batch is padded to a different width — so a model that looked fine in training behaves differently at serving time, where batches are shaped differently. Mask the numerator, and normalise by the true lengths. ## Normalising the loss The same choice appears in the loss and it is a real modelling decision, not a formality. Dividing by the total number of real steps in the batch (a per-token mean) weights long sequences more heavily, because they contribute more terms. Averaging within each sequence first and then across sequences (a per-sequence mean) gives every example equal weight regardless of length. Dividing by `batch * T` does neither: the effective gradient scale then depends on how much padding this particular shuffle happened to produce, so the same data reordered trains differently. ## Padding also costs compute Correctness aside, pad steps are work. When a batch runs from 4 to 900 steps, more than 90% of the recurrent computation in that batch is spent on padding that then gets masked away. Two mitigations exist: skip pad steps in the recurrence so the state is carried through unchanged instead of updated, and batch sequences of similar length together so the padded width stays close to the real lengths. ## What the interviewer is listening for They want to hear that padding is a storage requirement, that masking is the correction, and that the correction has both a numerator and a denominator. A candidate who says "I pad with zeros so it does not matter" has told the interviewer they have never checked what the state does during the pad run.

  • You already zero the pad vectors before mean pooling. Why is that still not enough?
    Because the denominator matters too. If the sum of masked vectors is divided by the padded width rather than by the sequence's true length, each pooled vector gets scaled by roughly length over padded width, so its magnitude encodes length. Short sequences are damped hardest, and the representation changes whenever a batch happens to be padded to a different width. Divide by the summed mask.
  • For a per-step loss, what is the difference between normalising by total real steps and normalising per sequence?
    Normalising by the total number of real steps in the batch is a per-token mean: long sequences contribute more terms and so pull the gradient more. Averaging within each sequence and then across sequences gives every example equal weight regardless of length. They are different objectives; pick the one that matches how you will be evaluated, and keep it consistent between training and validation.
  • Does it matter what value you pad with if everything is properly masked?
    For the loss and for pooled outputs, no — masked terms are dropped whatever they contain. It still matters for anything the mask does not cover: if pad steps are allowed to run through the recurrence, the pad embedding decides how the state drifts, and a learned pad row can drift a long way. Keeping a fixed pad value and skipping pad steps entirely removes the question.

Padding is like typing spaces to make every line in a column the same width; the mask is the note saying which characters were actually typed, so nobody counts the spaces as content.

saying these in an interview costs you the question

  • Says zeros are neutral so no mask is needed
  • Thinks padding only changes the array shape, not the outputs
  • Masks the pooled sum but still divides by the padded width
  • Leaves pad positions in the per-step loss with a filler target
  • Assumes the last index of the batch is the end of every sequence

context