Why can a network's test error worsen for dozens of epochs and then improve for hundreds more?
answer
- training time acts like capacity
- the middle of the run is the peak
- hard and noisy examples fitted last
- recovery needs hundreds more epochs
- a rising curve is not proof to stop
basics
~10 sEpoch-wise double descent lets a fixed-size network's test error fall, rise while it memorizes the hardest and most mislabelled training examples, and then fall again over a much longer training budget.
solid answer
~50 sTraining time behaves like capacity. Early on the network has only fitted coarse structure, and test error falls. In the middle of the run it acquires just enough effective capacity to start memorizing the hard and mislabelled examples, and test error climbs - a run degrading between epoch 20 and epoch 60 looks exactly like ordinary overfitting. Keep going and the network passes its own interpolation point: training loss bottoms out, but continued updates keep reshaping which zero-error solution it sits in, toward a smoother one, and test error can improve for another few hundred epochs. The practical consequence is that a rising validation curve is not by itself proof that you trained too long. But the recovery needs a large model, noisy or hard examples and a long budget, so treat it as a hypothesis to test rather than a default.
go deeper
Know that test error over epochs is not always one descent followed by a rise. With a large model on noisy data it can dip, rise, and dip again without anything being changed.
Explain why training time behaves like capacity: broad structure is fitted first and the contradictory examples much later, and it is the memorizing of those that produces the middle bump.
Demonstrate the diagnosis: check whether training loss has flattened, judge the label quality and model size, and run one long baseline before declaring a run over-trained.
Decide when the team should pay for the long budget at all. The second descent costs hundreds of epochs, and a well-regularized shorter run often reaches the same validation point for far less compute.
## The observation Hold the model, the data and the schedule fixed and plot test error against epoch. The textbook picture is a descent followed by a rise once the model starts overfitting. Epoch-wise double descent is the observation that for a sufficiently large network on data with appreciable label noise, the curve can go **down, up, and down again**: it improves early, degrades across a middle stretch - say from epoch 20 to epoch 60 - and then improves steadily for another four hundred epochs, often ending better than the pre-degradation best. Nothing about the model changes across this run. Same width, same depth, same data, same optimizer. ## Why time acts like capacity The key idea is **effective capacity**: how much of the network's expressive power the training procedure has actually recruited so far. At initialization the function is nearly featureless; after a few epochs it captures broad, high-signal structure that most examples agree on; only much later does it have the fitted detail needed to force the rare, contradictory, mislabelled examples into place. So a single long run traverses the same three regimes that a capacity sweep traverses: - **Early**: effective capacity is below what is needed to interpolate. The fit is coarse, generalization improves - this is the classical descent. - **Middle**: effective capacity approaches the amount needed to interpolate this training set. The network starts memorizing the hard and corrupted examples, and to do so with barely enough recruited capacity it distorts the function around them. Test error rises. This is the epoch-wise analogue of sitting on the interpolation threshold. - **Late**: training loss is essentially at the floor and the network is interpolating. Updates no longer trade off fitting errors; they keep moving within the space of solutions that all achieve near-zero training loss, drifting toward smoother, lower-norm ones. Test error descends a second time. ## What it means at the console This is a real diagnostic trap. You look at a run where validation error has been climbing for forty epochs, conclude the model is overfitting, kill it and shrink the model. Sometimes that is correct. Sometimes you have stopped in the middle of the bump and thrown away the best model the configuration could have produced. The distinguishing evidence, in order of cost: 1. **Is the training loss still falling steeply?** The second descent only becomes possible once training loss has essentially bottomed out. A run whose training loss is still dropping fast is nowhere near its interpolation point, and the rise is more likely to be ordinary overfitting. 2. **How noisy are the labels?** With clean, well-curated labels there is far less to memorize, and the middle bump is usually shallow or absent. If you know the corruption rate is meaningful, the bump becomes a live hypothesis. 3. **How big is the model relative to the dataset?** A small model never gets past its interpolation point at all, so it can only overfit. 4. **Run the long baseline once.** The cheapest decisive test is a single run continued far past where you would normally stop, with the same seed and schedule, to see whether the curve turns over. Do it once for the configuration family, not for every experiment. ## Interactions with the rest of the recipe A schedule that anneals the learning rate to zero quickly ends the run inside the first phase, so the second descent never gets the chance to appear. Strong explicit regularization suppresses the bump in the same way it suppresses the model-wise peak: it prevents the exact interpolation of noisy labels that creates the instability in the first place. That gives the honest framing for a senior answer. Epoch-wise double descent is not a reason to train every model for a thousand epochs. It is a reason to know **why** your validation curve is rising before you act on it, and to recognize that 'train much longer' and 'regularize properly' are two different routes out of the same bump - with very different compute bills. ## The claim to avoid Do not claim that longer training always recovers. On clean data, with a modest model, or under strong regularization, a rising validation curve after the training loss has flattened usually means exactly what it looks like. The phenomenon is conditional, and stating the conditions is most of the answer.
- How would you check whether a rising validation curve is epoch-wise double descent rather than plain overfitting?Check the training loss first: the second descent only becomes possible once training loss has essentially bottomed out, so a still-plunging training curve argues for ordinary overfitting. Then check label quality and model size. If both point the other way, run one long baseline on the same seed and schedule, far past your usual stopping point, and see whether the curve turns over.
- Which conditions make epoch-wise double descent likely to appear at all?A heavily over-parameterized model, appreciable label noise or otherwise contradictory examples, weak or absent explicit regularization, and a budget long enough to run well past the point where training loss flattens. Weaken any one of those - clean labels, a strong weight-decay coefficient, a small model, a short schedule - and the curve is usually monotone after its first descent.
- What does this imply about a schedule that anneals the learning rate to zero within the first phase?It ends the run before a second descent could occur, so you will never observe one. Whether that costs accuracy is empirical: chasing the recovery is only worth it when the extra hundreds of epochs are affordable and the bump is deep. Often a shorter, properly regularized run reaches a comparable validation point for a fraction of the compute.
saying these in an interview costs you the question
- Insists any rising validation curve proves the model can never recover
- Claims longer training always eventually helps
- Blames the mid-run rise on the learning rate without checking training loss
- Ignores label noise as a precondition for the bump
- Expects the recovery from a small, strongly regularized model