skip to content

Why can a Q-network's performance on already-mastered states regress as training continues?

level: seniorimportance: should knowfreq 46%

answer

  1. generalisation and interference, same mechanism
  2. one update moves distant states' values
  3. the agent's own data distribution keeps moving
  4. no rehearsal of regions no longer visited
  5. track a frozen probe set of states

basics

~20 s

Shared weights mean fitting one region of the state space silently moves the values of other regions. Because an agent's own improving policy keeps shifting which states it visits, the regions it stopped visiting are overwritten rather than rehearsed.

solid answer

~50 s

This is catastrophic interference. A network stores every state's value in the same parameters, so a gradient step that fixes one region also perturbs regions that share features with it. In supervised learning a fixed shuffled dataset keeps rehearsing every region, so the perturbations cancel out. In reinforcement learning the agent generates its own data, and as its policy improves the visited distribution moves, so previously-mastered regions simply stop appearing in the updates and drift. A city traffic-signal controller shows it clearly: fitting the value of one intersection's phase pattern silently degrades near-identical phases at neighbouring intersections, and nothing pulls them back until the agent happens to revisit them. Bootstrapping compounds it, because a degraded estimate is immediately used as a learning target elsewhere. The diagnosis is a fixed probe set of states whose predicted values you track over training, not the return curve alone.

go deeper

for a junior

Know that a network keeps all state values in one shared set of weights, so an update for one situation can change the answer for another.

for a middle

Explain why supervised training with a shuffled dataset escapes this and an agent generating its own shifting data does not.

for a senior

Demonstrate the diagnosis: a frozen probe set, per-scenario evaluation, and a checkpoint comparison, rather than reading the pooled return curve.

for a principal

Frame the tradeoff for the team - representations that generalise more interfere more - and decide what evaluation regime makes silent regression visible before it reaches production.

### The symptom Training looks healthy, average return climbs, and then behaviour that was reliable a hundred thousand steps ago quietly stops working. Nothing crashed, no loss spiked, and re-running the same evaluation confirms that the agent has genuinely got worse at a situation it used to handle. This is not noise and it is not the usual exploration wobble; it is interference. ### The cause: shared parameters A Q-network does not store one number per state. It stores one parameter vector that produces every state's values. Two states with similar features push through overlapping paths in the network, so a gradient step taken to correct one of them also moves the other. Generalisation and interference are the same mechanism seen from opposite sides: helpful when the neighbouring state's true value really is similar, destructive when it is not. A city traffic-signal controller makes it concrete. Intersections have near-identical phase patterns but different demand profiles, so their true values differ while their feature representations barely do. Fitting the value of one intersection's phase pattern drags the estimates for the neighbouring ones with it. Nothing in the update notices the damage, because the loss only measures the transition currently being fitted. ### Why reinforcement learning suffers more than supervised learning Interference exists in supervised learning too, under the name catastrophic forgetting, but it needs a specific setup to bite: training on task A and then on task B without revisiting A. Ordinary supervised training with a fixed shuffled dataset is immune, because every region of the input space is rehearsed on every epoch, so perturbations from one batch are corrected by the next. Reinforcement learning removes that protection in three ways. - **The agent generates its own data.** As the policy improves, the distribution of visited states moves. Regions the agent has learned to leave quickly, or has learned to avoid entirely, stop appearing in the updates - so nothing corrects the drift in their values. - **The data is temporally correlated.** Consecutive experience comes from one part of the state space, so a stretch of training can consist almost entirely of one region. - **Bootstrapping propagates damage.** Because targets are built from the network's own estimates, a value that has drifted in a region nobody is checking is immediately used as the learning target for the states that lead into it. Interference does not stay local. There is a second-order consequence too: if a degraded value makes a state look better than it is, the greedy policy will steer toward it, changing the data distribution again. The failure is self-reinforcing rather than self-correcting. ### How to diagnose it Average return is the wrong instrument, because gains in a region the agent visits often can mask losses in one it visits rarely. Better evidence: - **A fixed probe set.** Freeze a set of representative states before training and log their predicted action values at intervals. Interference shows up as movement in the predictions for probe states that recent experience never touched. - **Per-region evaluation.** Evaluate returns broken down by starting condition or scenario type rather than pooled. Regression in one scenario while the pooled average rises is the signature. - **Held-out transitions.** Keep a fixed set of transitions and track the magnitude of their prediction errors over time; interference makes errors on stale regions grow while errors on fresh ones shrink. - **Compare against a frozen checkpoint.** Replaying an older checkpoint's policy on the regressed scenario distinguishes real forgetting from exploration variance. ### What reduces it The two levers are the data distribution and the representation. Keeping the update stream diverse across regions - rather than letting it track whatever the current policy happens to be doing - restores something like the rehearsal that supervised learning gets for free. On the representation side, features with less overlap interfere less: sparse or locally-scoped representations confine an update's effect to a smaller neighbourhood, at the cost of generalising less. Smaller step sizes slow the damage without removing it. Normalising inputs and rewards keeps any single update from making a large parameter move. None of these is a cure. Interference is the price of the parameter sharing that made the network worth using, so the goal is to keep it beneath the rate at which useful learning happens, not to eliminate it. ### Answering well Say the mechanism (shared weights), say why RL is the hard case (the agent's own shifting data distribution removes rehearsal), say how bootstrapping spreads the damage, and give a concrete diagnostic. A candidate who has actually lived through this reaches for a fixed probe set instead of staring at the return curve.

  • How do you distinguish interference from ordinary exploration noise?
    Freeze the current parameters and evaluate the regressed scenario deterministically, with exploration switched off, several times. Exploration noise disappears; interference does not. Then replay an older checkpoint on the same scenario - if it still succeeds where the current one fails, values genuinely degraded rather than the agent merely acting randomly. Predictions on a fixed probe set moving while nothing in recent data touched those states is the confirming evidence.
  • Does catastrophic interference also affect supervised learning?
    Yes, under the name catastrophic forgetting, but only in continual or sequential training. With a fixed shuffled dataset, every region is rehearsed each epoch, so perturbations from one batch are corrected by others. Reinforcement learning is the hard case because the agent generates its own data, so the visited distribution shifts as the policy improves and stale regions get no rehearsal at all.
  • Why does bootstrapping make interference worse than it would otherwise be?
    The learning target for a state is built from the network's own estimate of a successor. If that successor's value has drifted through interference, the drift is copied into the target and spreads to every state that leads into it. A degraded value can also make a state look attractive, so the greedy policy steers toward it and shifts the data distribution again, which is a self-reinforcing rather than self-correcting loop.

Tuning one string on a guitar with a shared bridge nudges the others out of tune. Nobody notices until they play a chord they had not played in a while.

saying these in an interview costs you the question

  • Blames overfitting rather than interference between shared parameters
  • Assumes rising average return proves nothing has regressed
  • Thinks a lower learning rate eliminates interference entirely
  • Treats generalisation and interference as unrelated phenomena
  • Diagnoses only from the training loss curve

context