skip to content

How do you detect catastrophic forgetting while adapting a model to a new task?

level: seniorimportance: should knowfreq 46%

answer

  1. the failure has no dashboard
  2. freeze a slice of the old task
  3. same cadence as new-task validation
  4. per class, not aggregate
  5. best-so-far minus current, as a gate

basics

~20 s

Keep a frozen evaluation set from the original task and score it on the same cadence as the new task's validation set. Forgetting shows up as a falling retained-source score, so make it a release gate.

solid answer

~50 s

The failure is silent by construction: every metric you watch during adaptation belongs to the new task, and all of them look healthy. So detection means a retained source evaluation set — frozen, versioned, never trained on — scored each epoch alongside the new-task validation set and wired as a tripwire. Report it per class: in class-incremental factory defect inspection with five stages of two defect classes each, overall accuracy can look acceptable while task-1 recall has fallen from 0.91 to 0.42, because later classes carry the volume. Track the forgetting measure for each earlier stage — its best score so far minus its current score. If a drop is confirmed, a linear probe refitted on the frozen backbone tells you whether the representation drifted or only the head was re-aimed. Make the retained score a hard promotion criterion, or the regression ships behind a green dashboard.

go deeper

for a junior

Know that the old task needs its own held-out evaluation set and that it must be scored during the run, not only at the end. Recognise that healthy new-task metrics say nothing about what was lost.

for a middle

Explain why the retained set must be frozen, versioned and disjoint from anything replayed into training, and why per-class reporting beats an aggregate when later classes dominate the volume.

for a senior

Demonstrate that you design the tripwire before the run and wire it into promotion criteria. Be able to compute a forgetting measure per earlier stage and to localise the damage with a linear probe on the frozen backbone.

for a principal

Own where the acceptable-loss threshold sits for each retained capability and who signs off when it is crossed. Argue for a standing retention budget across releases rather than a per-run judgment call made after the numbers are in.

## Why detection is the whole problem Catastrophic forgetting is not hard to fix once you know it is happening. It is hard to *notice*. During adaptation, every number on the screen — training loss, new-task validation accuracy, the new task's confusion matrix — comes from the new distribution, and all of them are improving. Nothing in the ordinary training loop looks at the capability being lost. A news-topic tagger refreshed monthly on the current month's articles will quietly stop recognising last year's label set while its monthly evaluation reports its best score ever. So detection is a design decision made *before* the adaptation run, not a diagnosis made after a complaint. ## The retained source evaluation set The primitive is simple: carve a held-out evaluation set from the original task and freeze it. The properties that matter: - **Never trained on, at any stage.** If old examples are also being replayed into training, the evaluation slice must be disjoint from the replayed pool, or you are measuring memorisation of the buffer rather than retained capability. - **Versioned and immutable.** A retained eval that quietly gets regenerated each cycle cannot support a trend across releases. Its whole value is comparability over months. - **Scored on the same cadence as the new-task validation set** — each epoch during a run, and each release across runs. Forgetting can appear within a single epoch, so per-run checkpoints matter. - **Covering every capability you claim to still have**, including ones nobody has asked about recently. Anything not represented in the retained set is unprotected by definition. ## Read it per class, not in aggregate Aggregate accuracy is the most common way teams miss a live regression. Under class-incremental training, later stages usually contribute most of the current evaluation volume and are also the classes the model currently favours, so their strength masks the earlier classes' collapse. The concrete shape: five sequential stages of factory defect inspection, two defect classes per stage. After stage 5, headline accuracy across all ten classes may still look respectable, while stage-1 recall has fallen from 0.91 to 0.42. Per-class recall shows it immediately; a single averaged number does not. Reporting a per-stage or per-class table, and alerting on the worst cell rather than the mean, is the fix. ## Metrics that summarise a sequence When adaptation happens repeatedly, you want a summary that survives more than two stages. - **Forgetting measure per task.** For each earlier task, the difference between the best score it ever achieved and its score now. It is per-task and directly interpretable as damage. - **Average performance over all tasks seen so far.** The headline number for a continual pipeline; it balances plasticity on new tasks against retention of old ones. - **Backward transfer.** The change in earlier tasks' scores caused by learning later ones. Negative backward transfer is forgetting; the occasional positive value means the later task actually helped an earlier one. Tracking these across releases turns a one-off surprise into a trend line you can act on before it crosses a threshold. ## Localising the damage Once a drop is confirmed, the useful next question is *what* moved. Freeze the current backbone and fit a fresh linear classifier on old-task examples. If that probe recovers most of the old score, the representation is largely intact and the output layer has simply been re-aimed at the new labels — a much cheaper problem, and one that a calibrated or rebalanced final layer can often address. If the probe cannot recover the score, the features themselves have drifted and you need to change the training procedure rather than the head. A second, cheaper signal is parameter movement: the norm of the change in each layer's weights relative to its starting point. Large movement concentrated in upper layers is consistent with feature drift; movement confined to the head is not. ## Wiring it as a gate, not a chart A detection mechanism nobody blocks on is documentation, not a control. The retained source score belongs in the promotion criteria: if per-class recall on any retained capability falls more than an agreed margin below its recorded best, the candidate does not ship. That margin is a product decision — for a retired capability it may be generous, for a safety-relevant one it may be zero — but it has to be written down before the run, because after the run there is always a reason why this particular drop was acceptable. The same gate is what makes the mitigations measurable. Whether you interleave old examples into training or penalise movement of important parameters, the retained evaluation is how you tell whether it worked and how much new-task performance it cost. ## Common detection mistakes - Evaluating the old task only at the end of the run, after the checkpoints that would have shown the onset were discarded. - Building the retained set from the same pool that is being replayed, so it reports buffer memorisation. - Watching aggregate accuracy across an imbalanced class set. - Assuming the old task is safe because the new task's loss curve looks well behaved. - Regenerating the retained set each cycle, destroying comparability with earlier releases.

  • Why is aggregate accuracy on a mixed evaluation set a poor forgetting detector?
    Because the classes learned most recently usually dominate both the evaluation volume and the model's predictions, so their strength offsets an early class collapsing. Ten-class accuracy can look acceptable while the first stage's recall has halved. Reporting per class or per stage and alerting on the worst cell rather than the mean is what makes the regression visible.
  • Your retained evaluation set overlaps the pool of old examples being replayed into training. What breaks?
    You stop measuring retained capability and start measuring memorisation of the replayed examples. The retained score stays high because those exact items are in the training stream, which is the one outcome the gate is supposed to rule out. The evaluation slice must be disjoint from the replay pool and frozen before either is used.
  • How would you tell whether the features drifted or only the output layer was re-aimed?
    Freeze the current backbone and fit a fresh linear classifier on old-task examples. If it recovers most of the old score, the representation still encodes what the old task needed and the head is the casualty, which is cheap to repair. If it cannot, the features themselves moved and the training procedure has to change.
  • How often should the retained evaluation run during an adaptation?
    At least every epoch, and for short aggressive runs more often. Forgetting can develop within a single pass over the new data, so an end-of-run evaluation tells you that damage happened but not when, and by then the intermediate checkpoints that would have shown the onset are usually gone.

saying these in an interview costs you the question

  • Checks only the new task's validation metrics
  • Reports one aggregate accuracy over imbalanced classes
  • Evaluates the old task once, after training finishes
  • Builds the retained eval from the replayed examples
  • Treats the retained score as a chart, never a gate
  • Regenerates the retained set every release

context