skip to content

How does elastic weight consolidation protect a network against catastrophic forgetting?

level: seniorimportance: nice to knowfreq 26%

answer

  1. a spring per parameter, not per layer
  2. anchored at the old solution, not at zero
  3. stiffness comes from Fisher information
  4. squared gradients of the old log-likelihood
  5. store two vectors, discard the examples

basics

~20 s

Elastic weight consolidation adds a quadratic penalty pulling every parameter toward its old-task value, scaled by that parameter's Fisher information. Weights the old task depended on become stiff, while unimportant ones stay free to fit the new task.

solid answer

~50 s

Elastic weight consolidation replaces the data with a constraint. Training on the new task minimises the new loss plus a penalty of the form (lambda/2) * sum over i of F_i * (theta_i - theta_star_i) squared, where theta_star is the parameter vector at the end of the old task and F_i is the i-th diagonal entry of the Fisher information matrix estimated at theta_star. Fisher entries are average squared gradients of the old-task log-likelihood, so a parameter whose small perturbation would sharply change that likelihood gets a stiff spring, while one the old task barely used stays free. After that single estimation pass you store only theta_star and the Fisher diagonal — no examples — which is why it is the fallback when source data cannot be retained. The limits: the diagonal ignores parameter correlations, lambda trades retention against new-task performance, penalties accumulate across tasks and progressively stiffen the network.

go deeper

for a junior

Know the shape of the idea: a penalty keeps important weights near their old values while unimportant ones stay free, and importance is estimated rather than declared by hand.

for a middle

Be able to write the objective as the new loss plus a Fisher-weighted squared distance from the old parameters, and to say what the Fisher diagonal measures and why it is not the same as weight decay.

for a senior

Explain the estimation pass, the storage argument for keeping two vectors instead of a dataset, and how you would sweep lambda against both a new-task validation score and a retained source evaluation.

for a principal

Own the choice between buying retention with stored data and buying it with a parameter constraint, given data-retention policy and how long the adaptation sequence will run. Name the point at which accumulated stiffness makes continual adaptation no longer worth it.

## The premise Rehearsal fights forgetting by putting the old data back into the objective. **Elastic weight consolidation (EWC)** fights it without the data, by adding a term that makes the *parameters* the old task relied on expensive to move. It is the canonical regularisation-based approach to continual learning and the one to reach for when the source examples cannot be kept. ## The objective During adaptation to a new task the loss becomes: `total_loss = new_task_loss + (lambda / 2) * sum_i F_i * (theta_i - theta_star_i)^2` where: - `theta_star` is the parameter vector at the end of training on the old task — a snapshot, - `F_i` is the i-th entry of the diagonal of the Fisher information matrix, evaluated at `theta_star` on old-task data, - `lambda` sets the overall strength of the constraint. Read it as a spring per parameter. Every weight is attached to where it was, and `F_i` is that spring's stiffness. Where stiffness is high the weight can barely move; where it is near zero the weight is effectively unconstrained and available for the new task. ## What the Fisher diagonal means and how it is estimated The Fisher information measures how sharply the model's log-likelihood responds to a change in a parameter. In practice the diagonal is estimated by running old-task inputs through the trained model and averaging the **squared gradient of the log-likelihood** with respect to each parameter: `F_i = average over samples of (d log p(y | x; theta_star) / d theta_i)^2` Two variants exist and it is worth being precise about which you mean. The true Fisher samples the label `y` from the model's own predictive distribution; the *empirical* Fisher uses the observed labels from the data. The empirical version is the cheaper and more common choice, and it coincides with the true Fisher when the model fits the old task well. The intuition: a large squared gradient means small perturbations of that weight change what the model believes about old-task data, so moving it is likely to cost old-task performance. A near-zero value means the old task is indifferent to that weight. EWC turns that indifference into permission. The estimation is a single pass over old-task examples, done **once**, at the moment the old task finishes. Afterwards you store `theta_star` and the Fisher diagonal — two vectors the size of the model — and the examples can be discarded. That is the whole storage argument for EWC, and it is also its honest caveat: EWC is not a no-data method, it is a *no-retention* method. You need access to old-task data at consolidation time. ## Why the penalty is not weight decay A frequent confusion. Weight decay pulls parameters toward **zero**, uniformly, and expresses a preference for small weights. EWC pulls parameters toward the **old solution**, with a per-parameter strength derived from how much the old task cared. They have the same quadratic form and completely different anchors, and swapping one for the other does nothing for retention. It is also not freezing. A frozen parameter cannot move at all; an EWC-constrained parameter can move if the new task's gradient is strong enough to pay the quadratic price. That graded, negotiable stiffness is the reason the method is called *elastic*, and it is what lets the network keep learning where the old task has no stake. ## Choosing lambda Lambda is the retention/plasticity dial and it has failure modes at both ends. - Too small: the penalty is negligible, the parameters drift, and forgetting proceeds much as it would without EWC. - Too large: the model is anchored to the old solution and cannot fit the new task — retention is perfect and the adaptation is pointless. There is no default that transfers between problems, because the scale depends on the loss magnitude and on the Fisher scale. Tune it against both evaluations: the new task's validation score and a retained source evaluation set, sweeping lambda and reading the frontier between the two. ## Where it falls short - **The diagonal approximation.** Only per-parameter importance is modelled; correlations between parameters are ignored. Real networks have strongly coupled weights, and a pair that must move together to preserve a feature is not represented by two independent springs. - **Accumulation across tasks.** With several past tasks you either keep a penalty per task, which grows the bookkeeping linearly, or fold them into a single running anchor. Either way the constrained fraction of the network grows over a long sequence and plasticity declines — the network gradually stiffens until new tasks stop landing. - **Task boundaries required.** The snapshot and the Fisher pass have to happen at a known moment. A continuously drifting stream with no clean boundary does not offer one. - **Weaker than replay.** When old data is available, interleaving it typically retains more for less tuning effort. EWC earns its place when the data cannot be kept, or as a complement that lets a smaller buffer go further. - **A local, quadratic picture.** The penalty is a second-order approximation of the old loss around `theta_star`. It is trustworthy near that point and progressively less so as the parameters travel far from it, which is exactly the regime a large new task creates. ## How to present it in an interview State the objective, say what the Fisher diagonal measures and how it is estimated, distinguish it from weight decay and from freezing, and name the diagonal approximation plus the accumulation problem as the limits. If you also say that you would still measure the result on a retained source evaluation rather than trusting the penalty, you have covered what the question is really probing: whether you understand that the method is an approximation with a knob, not a guarantee.

  • How is the Fisher diagonal actually estimated in practice?
    Run old-task inputs through the model at the end of the old task and average the squared gradient of the log-likelihood with respect to each parameter. Using the observed labels gives the empirical Fisher, the common and cheaper choice; sampling labels from the model's predictive distribution gives the true Fisher. It is one pass, done once, and afterwards only the resulting vector is kept.
  • How does the EWC penalty differ from ordinary weight decay?
    Both are quadratic, but the anchors differ. Weight decay pulls every parameter toward zero with a single uniform coefficient, expressing a preference for small weights. EWC pulls each parameter toward its old-task value with a strength set by that parameter's Fisher information, expressing which weights the old task depended on. Only the second preserves a previous solution.
  • What goes wrong when EWC is applied across a long sequence of tasks?
    The constraints accumulate. Each consolidation stiffens more of the network, so the fraction of parameters free to change shrinks and later tasks stop landing — retention is bought with steadily falling plasticity. You also face bookkeeping growth if you keep a separate penalty per task rather than folding them into one running anchor.
  • How would you choose lambda for an EWC run?
    Sweep it and read the frontier between two measurements: the new task's validation score and a retained source evaluation. Too small and forgetting proceeds almost unchecked; too large and the model cannot fit the new task at all. No default transfers between problems, because the right scale depends on the loss magnitude and the Fisher scale.

Setting some bolts on a machine to high torque and leaving others finger-tight: nothing is welded shut, but the parts that mattered before now resist being moved.

saying these in an interview costs you the question

  • Says EWC freezes the important weights outright
  • Describes the penalty as pulling weights toward zero
  • Claims EWC needs no old-task data whatsoever
  • Ignores that only the Fisher diagonal is modelled
  • Treats lambda as having a universal default value
  • Assumes it retains as much as replaying real data

context