skip to content

How does replaying a small buffer of source examples reduce catastrophic forgetting?

level: middleimportance: should knowfreq 44%

answer

  1. put the old task back in the objective
  2. mix into every batch, not a separate phase
  3. a couple of percent often suffices
  4. buffer defines what is protected
  5. eval set must be disjoint from the buffer

basics

~20 s

Interleaving a small sample of the original task's data into every batch puts an old-task term back into each gradient. The optimizer can no longer lower the new loss at unlimited cost to the old one.

solid answer

~50 s

Forgetting happens because the objective contains no term from the old distribution, and rehearsal fixes that at the source. Keep a buffer of original examples — a fraction as small as 2 percent of the source set is often enough — and interleave it into every batch rather than training on it in a separate phase. Each update is then a gradient on a mixture, so a direction that would wreck the old task raises the loss it is scored on and gets pushed back. Two choices decide whether it works. Sampling: the buffer should be class-balanced and cover every capability you promised to retain, since anything absent from it is unprotected. Size and weighting: a tiny buffer is revisited many times and can be memorised, so the retained evaluation set must be disjoint from it. The cost is storage plus a data-retention question, and if the source data cannot be kept, rehearsal is simply unavailable.

go deeper

for a junior

Know that mixing some original examples into training preserves the old task, and that they belong in the same batches as the new data rather than in a separate pass afterwards.

for a middle

Explain the gradient argument: a mixed batch gives every update an old-task term that resists directions destroying old performance. Be able to discuss buffer size, the mixing ratio and why a small buffer risks memorisation.

for a senior

Show you treat buffer contents as the retention contract — balanced, covering every promised capability, versioned, and disjoint from the retained evaluation set. Tune the ratio against both evaluations rather than fixing it by habit.

for a principal

Own the tradeoff between retention and storage or data-retention policy, and decide when a capability is retired and leaves both the buffer and the gate. Otherwise the buffer grows without limit and every adaptation cycle gets slower.

## The idea If forgetting is caused by an objective that mentions only the new task, the most direct repair is to put the old task back into the objective. **Rehearsal** — also called replay — does exactly that: a stored buffer of original examples is mixed into the training stream during adaptation, so every update is computed on a mixture of old and new data. It is the strongest and simplest mitigation available, and everything else in this area exists because keeping the old data is sometimes impossible. ## Why it works With a mixed batch, the gradient at each step is a weighted sum of a new-task gradient and an old-task gradient. A direction that lowers the new loss while destroying old performance now increases the old-task term, and the sum resists it. The optimizer is pushed toward the intersection of the two low-loss regions rather than into whatever nearby point minimises the new task alone. The crucial detail is *interleaving*, not alternating phases. Training on new data for an epoch and then on old data for an epoch reproduces the original problem twice — each phase forgets the other. Every batch should carry both. ## How much data Surprisingly little. Because the pretrained weights still encode most of the old structure, the buffer's job is to *anchor*, not to reteach: a buffer holding on the order of a couple of percent of the source examples, interleaved into every batch, is often enough to hold a retained source score near its original level. That asymmetry — a small fraction of data preventing a large loss of capability — is the practical reason rehearsal is the default recommendation. The mixing ratio is a separate knob from buffer size. You can hold 2 percent of the source set and still make old examples 10 percent of every batch by sampling them more often; you can also weight the old-task loss term up or down. More old signal buys retention and costs plasticity on the new task, and the retained evaluation is how you find the point you actually want. ## What goes in the buffer The buffer defines what is protected. Anything not represented is unprotected, so selection matters more than raw count. - **Cover every retained capability.** In class-incremental training, keep examples per class rather than sampling uniformly from a stream, or rare classes vanish from the buffer and then from the model. - **Keep it balanced.** An imbalanced buffer reproduces its imbalance as a prediction bias, on top of the recency bias that class-incremental training already induces. - **Prefer diverse, representative examples over exclusively hard ones.** Selecting only the hardest cases produces a buffer that is unrepresentative of the old distribution and can anchor the model to its edge cases. - **Fix it before the run and version it**, so a retention number means the same thing across releases. ## The failure modes - **Buffer memorisation.** A small buffer is revisited hundreds of times more often than any new example. The model can fit those specific items while the underlying capability still decays. This is why the retained evaluation set must be strictly disjoint from the buffer — otherwise the tripwire reports the buffer's training accuracy and never fires. - **Ratio set and forgotten.** Too little old data and forgetting proceeds barely slowed; too much and the new task never really lands. It is a tunable with a measurable cost on both sides, so tune it against both evaluations. - **Phase separation.** Replaying old data as a warm-up or a final repair pass instead of interleaving it wastes most of the benefit. - **Stale buffer under distribution shift.** If the world moved, examples kept from two years ago anchor the model to a distribution that no longer exists. For a news-topic tagger refreshed monthly, the buffer must be curated to represent label semantics you still want, not merely old rows. ## What it costs Rehearsal buys retention with storage and with permission. You must keep the source examples, which raises retention-policy, licensing and privacy questions that are often the real blocker — a vendor dataset you are no longer licensed to hold, or user data you promised to delete, cannot sit in a replay buffer. It also lengthens each run, since part of every batch is spent on data whose task is already learned. When the data genuinely cannot be kept, rehearsal is off the table and you fall back to methods that operate on the parameters instead — a penalty that makes the weights the old task depended on expensive to move — accepting weaker retention in exchange for storing no examples at all. ## The judgment call Rehearsal is not free plasticity. Every batch slot given to old data is one not spent on the new task, and every retained capability enlarges the buffer. At some point the honest answer is that a capability is retired and should be dropped from both the buffer and the retained evaluation — a decision worth making explicitly, since the alternative is a buffer that grows forever and an adaptation that gets slower every cycle.

  • Why interleave the replayed examples into every batch instead of training on them in a separate phase?
    Because a separate phase recreates the original failure. Training on new data alone forgets the old task, then training on old data alone drifts back off the new one, and you oscillate rather than converge. Only a mixed batch produces a gradient containing both terms at every step, which is what forces a solution acceptable to both tasks.
  • How do you choose which examples go into the buffer?
    Cover every capability you intend to retain, keep it class-balanced so it does not import a prediction bias, and prefer diverse representative examples over only the hardest ones. Fix and version it before the run. Whatever is missing from the buffer is unprotected, so selection defines the retention guarantee more than raw count does.
  • What happens if the retained evaluation set is drawn from the same pool as the replay buffer?
    The gate stops working. Buffer items are seen hundreds of times during training, so the score reflects memorisation of those specific examples rather than retained capability, and it stays high while the underlying ability decays. The evaluation slice has to be carved out and frozen before the buffer is built.
  • The source data cannot be retained for legal reasons. What changes?
    Rehearsal is unavailable, so retention has to come from the parameters rather than the data — a penalty that makes the weights the old task relied on expensive to move, computed once while the old data is still accessible and then stored as statistics rather than examples. Expect weaker retention and plan the acceptable-loss threshold accordingly.

Keeping a few old exam questions in every practice set: you are not relearning last year's course, just refusing to let this year's cramming crowd it out.

saying these in an interview costs you the question

  • Trains on old data in a separate later phase
  • Thinks the buffer must be a large share of the source set
  • Evaluates retention on examples that are in the buffer
  • Samples the buffer uniformly and starves rare classes
  • Ignores the storage and data-retention cost entirely

context