Why does packing several SFT examples into one sequence risk cross-example contamination?
answer
- padding is paid-for nothing
- glue examples, share attention
- block-diagonal, not one long causal mask
- restart the position counter
- loss curve looks better, not worse
basics
~20 sPacking concatenates unrelated examples into one full-length sequence to avoid padding waste. Without boundary-aware attention masking and position resets, tokens of the third example attend to the first, so the model learns to condition answers on irrelevant preceding text.
solid answer
~50 sShort SFT examples waste compute: pad a 300-token transcript out to a 4,096-token window and most of the batch is padding. Packing fixes that by concatenating several examples end to end until the window is full. The catch is that a transformer attends across the whole sequence by default, so the tokens of the last packed example can attend to the first one — examples that have nothing to do with each other. The model then learns spurious conditioning: answers that depend on unrelated preceding text, and a bias toward whatever tends to sit earlier in a pack. The fix is a block-diagonal attention mask that confines each example to its own span, plus resetting position indices at each boundary so example three starts at position zero rather than at 2,000. Modern training stacks implement this, but it must be switched on deliberately — naive concatenation silently trains a subtly wrong objective.
go deeper
Know that packing means several short training examples share one full-length sequence to avoid wasting compute on padding, and that the examples must be kept separate from each other.
Explain the mechanism: causal attention spans the whole sequence, so without a block-diagonal mask a later example attends to an earlier unrelated one, and position ids must restart per example. Note that labels remain masked per example.
Show that you would verify it rather than assume it: confirm the stack passes document boundaries to the attention kernel, avoid mid-completion truncation, shuffle before packing, and treat a suspiciously good loss curve as a signal to check for leakage.
Frame it as a throughput-versus-correctness decision. Decide when packing is worth enabling given the tooling you can actually audit, and be willing to prefer length bucketing on a stack whose boundary handling you cannot verify.
## The efficiency problem packing solves SFT datasets are usually very uneven in length. A batch is a rectangle, so every sequence is padded to the longest one in it, and padding positions do work that is thrown away. With a 4,096-token window and a median example of 400 tokens, an unpacked batch can be 80–90% padding — you are paying for compute that contributes nothing to any gradient. **Sequence packing** concatenates examples end to end into a single full-length sequence, so each training token is a real token. On short-example datasets this is often a multiple-times speedup for identical learning, which is why it is default hygiene in production SFT rather than an exotic optimisation. ## What breaks if you just concatenate A causal transformer lets every position attend to all earlier positions in its sequence. Once you have glued unrelated examples together, "earlier positions" includes other people's examples. Two things go wrong. **Cross-example attention.** The answer tokens of the fourth packed example can attend to the transcript and answer of the first. The gradient therefore rewards using that context. You have quietly changed the task from "answer given this prompt" to "answer given this prompt and some unrelated text that happens to precede it." The damage is subtle: the model learns weak spurious correlations, and it becomes sensitive to whatever text precedes a request at inference — where, of course, no such preceding pack exists. **Position drift.** Positional information is assigned by offset within the sequence. Under naive packing, the same short example is at positions 0–400 when it lands first in a pack and at positions 3,100–3,500 when it lands last. The model sees the same content at wildly different positional encodings, which adds noise and, with rotary-style encodings, can push short examples into positional regimes they never occupy at inference. ## The correct implementation Two mechanisms, applied together: **Block-diagonal (document-boundary) attention.** Instead of a single causal mask over the whole packed sequence, build a mask that is causal *within* each example and zero across examples. Every token attends only to earlier tokens of its own example. Efficient attention kernels support this by taking the cumulative sequence lengths of the packed examples, so it costs essentially nothing — in fact it does less work than the dense mask. **Position-id resets.** Restart the position counter at each example boundary, so each packed example sees the same positions it would have seen alone. And, independently, the label mask still applies per example: each packed example contributes loss only on its own completion span, so a pack of five yields five separate supervised spans separated by ignored prompt regions. ## Second-order effects worth naming **Loss weighting shifts.** Averaging cross-entropy over all supervised tokens in a pack means examples with long completions contribute more than examples with short ones. Unpacked batches, averaged per example, weight examples equally. Neither is wrong, but the effective weighting changes when you turn packing on, so a like-for-like comparison against an unpacked baseline is not exactly like-for-like. **Truncation at the boundary.** A naive packer that fills to exactly the window length will cut the last example mid-completion, training a truncated target with no stop token. Packers should either move a straddling example to the next pack or, if splitting, do so knowingly. **Ordering artefacts.** Packing examples in dataset order groups similar examples into the same pack. Even with correct masking this correlates the examples inside a gradient step; shuffling before packing avoids it. **Detection.** Contamination from bad packing is hard to see in the loss curve — it typically shows as slightly *better* training loss, since extra context is extra information. The practical check is a controlled comparison: a short run with packing off, or with masking explicitly enabled versus disabled, evaluated on held-out prompts presented alone. If quality on single, isolated prompts is worse than the training curve implies, cross-example leakage is a prime suspect. ## The judgement call Packing is close to free when the stack supports boundary-aware attention, and it is a real correctness risk when it does not. If you are running on tooling where you cannot confirm that document boundaries are respected, the honest options are to leave packing off and accept the padding cost, or to bucket examples by length so padding waste shrinks without any concatenation. The one thing you should not do is concatenate and hope, because the failure is silent and the loss curve will actively reassure you.
- Besides the attention mask, what else must reset at a packed-example boundary?Position indices. Each packed example should start its position counter at zero, otherwise identical content gets very different positional encodings depending on where it lands in the pack, and short examples occupy positional regimes they never see at inference. Labels also reset per example: every packed example contributes loss only on its own completion span.
- How does packing change the effective weighting of examples in a gradient step?Averaging loss over all supervised tokens in a pack weights examples by completion length, so a long answer dominates several short ones. Unpacked, per-example averaging weights them equally. This makes packed and unpacked runs not strictly comparable, which matters when you are attributing a quality change to packing rather than to the weighting shift it introduced.
- When would you deliberately not pack?When the training stack cannot guarantee boundary-aware attention, since silent cross-example conditioning is worse than wasted padding. Also when examples are already near the window length, where packing buys almost nothing. Length bucketing is the middle path: group similar-length examples into batches to cut padding without concatenating unrelated examples at all.
saying these in an interview costs you the question
- Assumes concatenating examples is safe because attention is causal
- Thinks packing changes what the model learns only by speeding it up
- Forgets to reset position ids at example boundaries
- Lets the packer truncate the final example mid-completion
- Reads a lower training loss as proof packing was implemented correctly