Your team cites a double descent plot to argue for dropping regularization and just scaling up. What do you say?
answer
- one curve, one setting
- the sweep was under-regularized
- tuned regularization flattens the peak
- the threshold moves when data grows
- two routes out, not one
basics
~10 sA double descent plot does not show that scale replaces regularization: those sweeps are run under-regularized on noisy labels, and tuning the penalty at each capacity can flatten the peak away.
solid answer
~50 sI would separate what the plot shows from what it is being used to claim. It shows that on that dataset, that model family and that budget, test error was non-monotone in capacity and the largest model beat the peak. It does not show regularization was unnecessary: those sweeps are usually run with weak or no explicit regularization on data with a meaningful share of corrupted labels, and tuning the regularization strength separately at each capacity can flatten the bump into a near-monotone curve. It also does not transfer - the interpolation threshold moves with dataset size, label quality, architecture and schedule, so a size that is safely past the peak today can sit on top of it after the next data refresh. My counter-proposal is to keep regularization as a tuned axis in the sweep and budget one long run to locate our own threshold.
go deeper
Know that one plot showing a bigger model winning is not a general rule. The result depends on the dataset, the label quality and how the smaller models were trained.
Be able to name what was held fixed in such a sweep - regularization strength, schedule, label noise - and explain why leaving those untuned makes the comparison unfair to the smaller models.
Describe the experiment you would run instead: per-capacity tuned regularization on the real data, with compute and serving cost measured, before committing to a scaling plan.
Own the tradeoff between compute spent scaling past the threshold and effort spent on label quality, and set the expectation that the threshold is a dated estimate that must be revisited as data grows.
## Separate the observation from the inference The plot is real. A capacity sweep on an over-parameterized family, trained on data with corrupted labels, produces test error that is non-monotone in model size: a first descent, a peak at the interpolation threshold where training error first reaches zero, then a second descent in which the largest models are the best on the sweep. The inference on the table - 'therefore stop regularizing and scale' - imports three things the plot never established. ## Claim 1: that regularization was shown to be unnecessary The striking sweeps are typically run with weak or no explicit regularization; that is what makes the bump dramatic enough to plot. It is a demonstration of a phenomenon, not a tuned comparison of two strategies. When the regularization strength is instead **tuned separately at each capacity**, the peak is largely suppressed and the resulting curve is close to monotone in model size - and the tuned models are generally no worse, at any size, than the untuned ones. The mechanism is direct: the peak exists because a model with barely enough capacity is forced to interpolate corrupted labels exactly, and regularization is precisely the thing that prevents exact interpolation of noise. So the honest reading is: there are **two** routes out of the bad region. Scale decisively past the threshold, or regularize properly. The plot compares one of them against doing neither. That is the flaw in the comparison the team is drawing on - a tuned large model measured against untuned small ones. ## Claim 2: that a plot from one setting transfers to ours The interpolation threshold is not a property of an architecture. It is a property of an architecture *with* a dataset size, a label-noise level, an optimizer and a training budget. Change any of them and it moves. The operational consequence is uncomfortable and worth saying out loud: **adding data moves the threshold to a larger capacity.** A model that was comfortably past the peak can land on the peak after a data refresh and get measurably worse, despite training on more examples. A scaling policy justified by 'we are past the threshold' therefore needs re-checking whenever the corpus grows, which is exactly when nobody re-runs the sweep. ## Claim 3: that scale is free Even where the second descent is real, reaching it means training models substantially larger than the threshold, which is a compute, latency and serving-cost decision, not just an accuracy one. 'The biggest model on the sweep won' is only an argument if the biggest model is deployable. ## What I would propose instead 1. **Keep regularization in the sweep as a tuned axis.** Each capacity gets its own tuned strength. Anything less is not a comparison of strategies. 2. **Spend one long run to locate our threshold** on our data with our label quality, so we know which side we are on rather than assuming. 3. **Attack the ingredient rather than the symptom.** The bump is driven by label noise. Effort spent on label quality shrinks the phenomenon and helps every model on the curve, usually more cheaply than scaling past it. 4. **Avoid sitting near the threshold.** That is the one clear, actionable lesson. If the deployable size lands near the estimated threshold, either scale decisively past it or step back to a capacity where tuned regularization controls the fit - do not park on the peak. 5. **Re-check after data changes.** Make the threshold estimate a dated artifact with an owner, not a one-time slide. ## What I would concede The team is right about something important: the classical instinct that 'bigger will overfit' is not a reliable guide for over-parameterized networks, and refusing to try larger models on that instinct alone is a real failure mode too. The disagreement is not about whether to try scale - it is about deleting a tuned, cheap control on the strength of one uncontrolled comparison.
- What evidence would actually support choosing scale over regularization on our problem?A sweep on our data and our label quality where every capacity is trained with its own tuned regularization strength, showing the large model matching or beating the best regularized smaller one at an acceptable serving cost. Anything less compares a tuned large model against untuned small ones, which is not the claim being made.
- If we cannot afford models well past the threshold, what does double descent tell us to do?Stay clearly on the classical side of the peak and regularize properly. Sitting near the threshold is the worst place to be: the model has just enough capacity to interpolate the noisy labels and no slack to stay smooth. Either scale decisively past it or step back to a capacity where a tuned penalty controls the fit.
- Why does a growing training set complicate a scaling policy justified by double descent?More data pushes the interpolation threshold to a larger capacity. A model size that was safely past the peak can end up sitting on it after the corpus grows, and test error can regress even though nothing about the model changed. The threshold estimate has to be revisited whenever the dataset does.
saying these in an interview costs you the question
- Treats double descent as proof that regularization is obsolete
- Assumes a second descent appears on every dataset
- Ignores that the threshold moves as the dataset grows
- Compares a tuned large model against untuned small ones
- Cites the phenomenon without mentioning label noise
- Counts accuracy only, not the serving cost of the larger model