A penalised model's CV-error curve is flat across a whole decade of lambda — why is the lambda with the lowest mean error a poor pick?
answer
- compare the dip to the bars
- reshuffle the folds and watch it move
- smallest of many noisy means
- selected on noise, reported as truth
- a plateau is a tie, not an optimum
basics
~10 sOn a flat curve the winner is decided by fold noise, not by any real difference: reshuffle the folds and the minimum jumps to a neighbouring lambda. That lowest mean is also optimistically low.
solid answer
~50 sEach point on the curve is a noisy estimate: the mean of a handful of fold scores, with an error bar around it. When the curve is flat over a decade of lambda, the gaps between neighbouring points are smaller than those bars, so the argmin is essentially selected by which fold split you drew — re-run with different fold assignments and it hops to a different grid point. The reported minimum is also biased low: taking the smallest of a hundred noisy means picks up whatever favourable noise is in the grid, so it understates that lambda's true error. The right response is to treat the plateau as a tie set rather than pretend to resolve it, which is what the one-standard-error band does: it makes the tie explicit and breaks it toward larger lambda, giving a choice that barely moves across re-splits.
go deeper
Remember that every point on the curve is an estimate with a margin of error, not an exact number. If neighbouring points differ by less than that margin, the lowest one is not meaningfully the best.
Explain the two distinct effects: the selection is unstable because it ranks differences smaller than the noise, and the winning score is biased low because it was chosen for being low. Do not merge them into one hand-wave.
Show the diagnosis you would actually run: repeat the sweep under different fold assignments, tabulate where the minimum lands each time, compare it with the one-standard-error selection, and report a range rather than a single decimal.
Set the reporting norm. Decide in advance that tuning results are published as a selection rule plus an indistinguishable range, so no team member is free to quote a noise-driven argmin as the optimum and defend it after the fact.
## What flatness actually means A CV-error-versus-lambda curve is a hundred-ish noisy estimates plotted side by side. Each is the mean of the per-fold scores at one lambda, and each carries a standard error that tells you how much that mean would wobble under a different partition of the rows. A **flat** stretch means the vertical spread of the means across that stretch is comparable to, or smaller than, the error bars attached to them. This is common, not exotic. It happens when predictors are correlated, so extra shrinkage merely redistributes weight among near-duplicates without changing predictions much; when there is a broad range of lambda over which the effective complexity of the fit changes slowly; and whenever the dataset is large enough that the fit is not very sensitive to modest shrinkage. ## Two separate problems, often conflated **Instability of the selection.** If the differences you are ranking on are smaller than the noise in the ranking, the ordering is arbitrary. The concrete symptom: re-run cross-validation with a different fold assignment and record where the minimum lands each time. On a flat curve it wanders across a decade. Nothing about your data changed; only the noise did. Reporting `optimal lambda = 0.0137` after that is false precision. **Optimism of the reported minimum.** Even setting stability aside, the height at the argmin is a biased estimate of that lambda's true error. Suppose every grid point had exactly the same true error; each estimate would still land above or below it by chance, and by construction you are reading off the one that landed lowest. The minimum of many noisy estimates is systematically below the truth — you have selected on the noise, and the number you quote contains that selection. It is a self-scored result: the same fold errors both chose the winner and reported its score. One nuance worth having ready: the grid points are not independent estimates. Every lambda is scored on the same folds, so a fold split that is easy for one lambda tends to be easy for all of them, and the curve shifts up or down almost as a whole. That correlation damps the optimism relative to a hundred independent draws — but it does not remove it, because the *shape* of the curve still jitters and the argmin still follows the jitter. ## How to diagnose it in a few minutes - Compare the depth of the dip with the height of the bars. If the minimum sits well inside its neighbours' bars, you have a plateau, not an optimum. - Re-run the whole tuning sweep with two or three different fold assignments and tabulate the argmin from each. If they disagree by a factor of ten, say so out loud rather than picking one. - Do the same for the one-standard-error selection. It typically moves by one grid point or not at all, which is the point of using it. ## What flatness does not mean It does not mean the penalty is doing nothing. The plateau is the interior of the curve; the ends still tell a story — error climbs as lambda goes large enough to crush the model toward a constant prediction, and usually climbs on the small-lambda side too. Nor does it mean your grid is too coarse: refining a grid inside a plateau resolves noise, not signal, and gives you a more precisely stated arbitrary answer. ## What to do instead Stop asking the noise to choose. Convert the plateau into an explicit tie set — every lambda whose mean is within one standard error of the best — and break the tie on a criterion that does not depend on the fold split. Parsimony is the usual criterion, and on a lambda axis that means the largest lambda in the set. Three things improve at once: the selection is reproducible across re-splits, the deployed model is simpler, and you can state the rule you used before running the sweep rather than after seeing the plot. When you report, report the reading rather than a single decimal: the selected lambda, the range of lambda that was indistinguishable from the best, and the fact that the choice inside that range was made on simplicity. A reviewer who sees the band understands the model; one who sees only `lambda = 0.0137` has to take your word for it.
- How would you show a sceptical colleague that the minimum is unstable?Re-run the whole tuning sweep three or four times with different fold assignments and put the results in one small table: the argmin from each run, and the one-standard-error selection from each run. The argmin column typically jumps across a decade while the one-standard-error column moves by a grid point or not at all. That table settles the argument faster than any explanation.
- Would a finer lambda grid fix a flat curve?No. Refining the grid inside a plateau only adds more estimates whose differences are smaller than their error bars, so the argmin becomes more precise and no less arbitrary. A finer grid helps when the dip is genuinely sharp and you are stepping over it, which you can check by comparing the depth of the dip with the width of the bars.
- Where does the flatness usually come from?Most often from correlated predictors: over a wide range of lambda, extra shrinkage shuffles weight among near-duplicate features without changing the fitted predictions much, so error barely responds. It also appears when the sample is large relative to the number of predictors, so the unpenalised fit was already stable and the penalty has little to correct.
saying these in an interview costs you the question
- Quotes the argmin lambda to three decimals as optimal
- Believes averaging over folds removes the noise entirely
- Tries to fix a flat curve with a finer lambda grid
- Treats the minimum CV error as an unbiased performance number
- Concludes a flat region means the penalty is useless