What is double descent, and why does test error peak at the interpolation threshold?
answer
- three regimes, not two
- peak where training error hits zero
- just enough capacity to fit noise
- many interpolating solutions past the peak
- label noise creates the bump
basics
~20 sDouble descent is the pattern where test error falls, then rises to a peak at the interpolation threshold - the capacity at which a model can just barely fit every training label - and then falls again as capacity grows further.
solid answer
~40 sPlot test error against model capacity and you can see three regimes rather than two. Below the interpolation threshold the curve looks classical: error drops, then starts climbing as the model begins chasing noise. The peak sits at the threshold itself - the capacity at which the parameters first suffice to drive training error to zero, mislabelled examples included. There the fit is essentially forced: roughly one solution interpolates the data, and it does so with large, wildly oscillating weights. Push capacity well past that point and many different functions interpolate the same training set, so gradient descent's implicit preference for small-norm solutions can pick a smoother one, and test error descends a second time. The peak is large when the training labels carry noise and is often shallow or invisible on clean data.
go deeper
Be ready to say that test error can rise and then fall again as models get bigger, and that the bump sits near the size where training error first reaches zero.
An interviewer expects the three regimes and a precise definition of the interpolation threshold: the capacity at which the model first fits every training label exactly, corrupted ones included.
Show that you can tell from a sweep which regime you are in and say what you would change - scale decisively past the peak, tune regularization, or fix the label quality that created the bump.
Own the judgment call: whether a capacity sweep is worth running at all on this problem, what compute a second descent would cost, and how label quality changes the curve the team will get.
## The shape being described Take one model family, one dataset and one training recipe, and sweep a single capacity knob - say the width of every layer. For each width, train to convergence and record final test error. The classical expectation is a single-minimum curve: too little capacity underfits, too much overfits. Double descent is the observation that for modern over-parameterized networks the curve has **three** segments: a first descent, a rise to a peak, and then a **second descent** that can take test error below anything achieved on the classical side. A concrete version of the experiment: a 50,000-image, 10-class benchmark in which 15% of the training labels have been deliberately corrupted. Sweep width from tiny to very large. Test error improves, then gets clearly worse as width approaches a particular value, then improves again - monotonically - for every width beyond it. The widest model is the best model on the sweep, and it is better than the best model on the classical side. ## The interpolation threshold The peak is not at an arbitrary place. It sits at the **interpolation threshold**: the smallest capacity at which the model, trained with this optimizer for this budget, can drive *training* error to zero - fitting every training label exactly, including the 15% that are wrong. Two things about this definition matter in an interview. First, it is defined by what the model can *fit*, not by what it *should* fit. Interpolating a noisy training set is not a good thing; it is simply the event that marks the location of the peak. Second, for simple model families it lands roughly where the number of trainable parameters matches the number of training examples, but for deep networks that count is a poor proxy. The honest working definition is empirical: train models of increasing size and find the first one whose training error hits (essentially) zero. That point moves when you change the dataset size, the label quality, the architecture, the optimizer or the length of the schedule. ## Why the peak is there Just at the threshold, the model has *exactly* enough freedom to interpolate and none to spare. The interpolating solution is close to unique, and it is forced to pass through every corrupted label. To do that while remaining a smooth-ish function elsewhere is impossible, so the fitted function develops enormous swings around each mislabelled point. Small changes in the training sample produce completely different fitted functions - the estimator is extremely unstable - and test error is correspondingly bad. ## Why the second descent happens Past the threshold, the set of parameter settings that interpolate the training data stops being a point and becomes a large space. Every one of them has zero training error, so the training objective no longer distinguishes them. What breaks the tie is the **implicit bias** of the training procedure: gradient descent started near zero tends to land on solutions of small norm - the flattest, least contorted way of threading the same points. The bigger the model, the richer the set of interpolating functions, and the smoother the one the optimizer can find. Test error therefore falls again even though training error has been pinned at zero the whole time. This is why the phenomenon is real rather than an artifact: nothing about the training objective changed, only the size of the solution set the optimizer gets to choose from. ## The role of label noise Label noise is the ingredient that makes the bump visible. Interpolation means fitting the wrong labels too, and the cost of doing that with barely enough capacity is exactly the instability that creates the peak. Remove the corruption and the same sweep typically shows a mild bump or none at all. Any account of double descent that never mentions how clean the labels are is incomplete. The same is true of regularization: the sweeps that display a dramatic peak are usually run with weak or no explicit regularization. Tuning the regularization strength separately at each capacity can flatten the peak substantially, leaving a curve much closer to monotone. ## What this does not mean It does not mean overfitting has been repealed, that bigger is always better, or that you can read your own model's regime off somebody else's plot. It means capacity sweeps on over-parameterized models can be non-monotone, the worst place to sit is right on the threshold, and where that threshold is for *your* data is an empirical question.
- Where does the interpolation threshold sit in terms of parameters and training examples?For simple families it lands near the point where the number of trainable parameters matches the number of training examples, but that count is a weak proxy for deep networks. The usable definition is empirical: the smallest model in the family that drives training error to essentially zero, with this optimizer and this budget. It shifts when you add data, change architecture, or train longer.
- Why does label noise make the peak so much larger?Interpolation means fitting every label exactly, wrong ones included. Near the threshold the model has just enough capacity to do that and no slack left, so it distorts the fitted function violently around each corrupted point, and the fit becomes hypersensitive to the training sample. With clean labels there is far less to contort around, so the bump is shallow or absent.
- Can adding more training data ever make test error worse near the threshold?Yes, and it is the most counterintuitive corollary. More data pushes the interpolation threshold to a larger capacity, so a fixed-size model that was comfortably past the peak can end up sitting on top of it and get worse despite seeing more examples. The effect disappears once the model is scaled well beyond the new threshold.
Right at the threshold there is exactly one suit in the shop that fits, and it fits only by contorting around every wrinkle in the measurements. With a whole warehouse of suits, many of them fit - and the tailor can hand you the one with the cleanest lines.
saying these in an interview costs you the question
- Claims test error is always U-shaped in model capacity
- Places the peak wherever validation loss happens to be lowest
- Says double descent means overfitting cannot happen
- Never mentions label noise or regularization strength
- Treats a second descent as guaranteed on any dataset