skip to content

One-Standard-Error Rule

The cross-validated error curve over lambda is noisy, so its bare minimum overfits the folds; the rule instead takes the largest lambda within one error bar of it. Interviewers ask why.

on this pageshow

questions

3

Under the one-standard-error rule, with 5-fold CV mean error 2.41 at lambda = 0.01 (standard error 0.09), 2.49 at 0.05 and 2.58 at 0.10, which lambda wins?

level: middleimportance: should knowfreq 45%

answer

  1. start from the best mean, not the best lambda
  2. one error bar sets a tolerance band
  3. everything inside the band counts as tied
  4. break the tie toward more shrinkage
  5. 2.41 plus 0.09 is the cutoff

basics

~10 s

Lambda = 0.05. Take the lowest mean error, 2.41, add the standard error at that point, 0.09, to get a threshold of 2.50, then pick the largest lambda whose mean still sits under it.

solid answer

~50 s

The rule anchors on the minimum: mean `2.41` at `lambda = 0.01`, standard error `0.09`, so the tolerance band is everything at or below `2.41 + 0.09 = 2.50`. Walking toward heavier shrinkage, `0.05` has mean `2.49` and still clears the threshold, while `0.10` at `2.58` does not. So the selected value is `lambda = 0.05` — five times the minimising lambda, at a mean error only `0.08` worse, a gap smaller than the noise in the estimate itself. The direction matters: among candidates that are indistinguishable given that noise, you take the most heavily penalised one, because a bigger lambda means more shrinkage and a simpler, steadier model. If you were tuning on a higher-is-better score you would flip the band to best minus one standard error and again take the largest lambda inside it.

go deeper

for a junior

Be ready to do the arithmetic out loud: lowest mean plus its standard error gives a cutoff, and the answer is the largest lambda still under that cutoff. Saying 0.05 with the two-line calculation is a full pass here.

for a middle

An interviewer expects the mechanics: where the error bar comes from, why the threshold is anchored at the minimum rather than recomputed per candidate, why larger lambda counts as the simpler model, and how the band flips for a higher-is-better score.

for a senior

Say what the selection buys in production: a lambda that barely moves when the folds are reshuffled, and a model with smaller coefficients, at a cost you can quote in the metric's own units — here 0.08 days of mean absolute error.

for a principal

Own it as a team default rather than a per-project taste call. Decide before the tuning run which reading you will honour, write it down, and be explicit that the constant one is a convention you may replace with a tolerance stated in business units.

## The picture behind the numbers Tuning a penalty means fitting the same model at every value of lambda on a grid — say a hundred values running from almost no shrinkage up to a lambda that flattens the model — and scoring each fit by cross-validation. At each grid point you get one score per fold, and two numbers summarise them: the **mean**, which is the height of the curve, and the **standard error of that mean**, drawn as a vertical error bar. Plotted against lambda (usually on a log axis) the means form the CV-error curve. It normally dips somewhere in the middle: very small lambda leaves the fit free to chase noise, very large lambda shrinks the model toward a near-constant prediction. The naive reading is *take the lowest point*. The one-standard-error rule reads the same plot differently. ## The rule, stated exactly 1. Find the grid point with the lowest mean CV error. Call its mean `m` and the standard error of that mean `s`. 2. Form the threshold `m + s`. 3. Among all grid points whose mean error is at or below `m + s`, choose the one with the **largest** lambda. Two details carry most of the meaning. The threshold is built **once**, from the minimising point's own error bar — not from a bar recomputed at whichever candidate you are examining. And the tie-break runs toward the heavier penalty, because the whole intent is to prefer the simpler model among candidates you cannot actually tell apart. ## Worked on the numbers in the question Suppose the full table for a hospital length-of-stay fit, scored by mean absolute error in days, reads: ``` lambda 0.005 0.01 0.02 0.05 0.10 mean 2.43 2.41 2.44 2.49 2.58 ``` with the standard error at the minimising lambda equal to `0.09`. Threshold = `2.41 + 0.09 = 2.50`. Which grid points clear it? `0.005` (2.43), `0.01` (2.41), `0.02` (2.44) and `0.05` (2.49) all sit at or below `2.50`; `0.10` at `2.58` does not. The band is `{0.005, 0.01, 0.02, 0.05}` and the largest lambda inside it is **0.05**. You accept a mean error of `2.49` instead of `2.41` — eight hundredths of a day — in exchange for a model shrunk five times harder. Read it as a walk: start at the minimum, step toward larger lambda, and stop at the last point before the threshold is broken. ## Why the largest lambda is the simplest model For both an L2 and an L1 penalty, lambda multiplies the penalty term added to the loss, so raising it shrinks the coefficients harder; with L1 it also drives more of them to exactly zero. Along this axis, larger lambda means more shrinkage means lower model complexity. The rule is really the general heuristic *prefer the simplest candidate that is statistically indistinguishable from the best*, and on a lambda grid the simplest candidate is the largest lambda. If your tuning axis ran the other way — a complexity knob where a bigger number means a richer model — the same principle would send you to the **smaller** value. Carry the rule as "the most-regularised candidate inside the band", never as "the rightmost point on the plot". ## Metrics where higher is better If you tuned on mean AUC or mean R-squared, invert the band. Take the maximum mean `M` and its standard error `s`, keep every grid point whose mean is at least `M - s`, and again choose the largest lambda in that set. Flipping the sign by accident is one of the more common slips; the safe phrasing is "within one error bar of the best score, whichever direction best runs in". ## What the band is, and is not The band is a **tolerance**, not a hypothesis test. One standard error is a convention — nothing derives the constant, and there is nothing special about the number one. What it encodes is that differences inside the band are about the size of the noise in the estimate itself, so choosing among those candidates on the strength of that noise is not really choosing at all. The rule then settles the tie on a criterion you can defend out loud — parsimony — instead of on whichever fold split you happened to draw. It is also narrow in scope: it picks a value of lambda, and nothing more. It does not tell you what error to report for the chosen model, and it does not change how the folds were constructed. ## Traps that cost people the question - Anchoring the threshold at the candidate lambda's own error bar rather than the minimum's. - Walking left into smaller lambda, which selects a *more* complex model and inverts the rule's purpose. - Forgetting the sign flip on higher-is-better metrics, so the band comes out empty or covers the whole grid. - Claiming "within one standard error" means "not significant at the 5% level". - Reporting the minimising lambda anyway after describing the rule correctly — interviewers notice when the arithmetic and the conclusion disagree.

  • The error bar drawn at lambda = 0.10 is wide enough to reach 2.41 — does that make 0.10 eligible?
    No. The band is defined once, from the mean and standard error at the minimising lambda, and then every candidate's mean is compared against that single threshold. Here the threshold is 2.50 and the mean at 0.10 is 2.58, so it is out regardless of how wide its own bar is. Comparing each candidate's bar to the minimum is a different, much looser criterion that would push you toward a nearly null model.
  • Why does the rule move toward larger lambda rather than smaller?
    Because larger lambda means heavier shrinkage: smaller coefficients, and with an L1 penalty fewer non-zero ones. Inside the band the candidates are indistinguishable on error, so the deciding criterion is simplicity, and simplicity lies to the right. Walking left would pick a more complex model on the strength of noise, which is exactly what the rule exists to avoid.
  • How does the reading change if you tune on mean AUC instead of mean error?
    Flip the band. Find the maximum mean AUC and the standard error at that point, keep every lambda whose mean AUC is at least maximum minus one standard error, and choose the largest lambda among them. The logic is identical — within one error bar of the best — only the direction of best changes.

Several runners finish inside the timing system's margin of error. You do not crown the one whose clock read lowest; you call it a tie and apply a tie-break you can justify.

saying these in an interview costs you the question

  • Always takes the lambda with the lowest CV mean error
  • Adds the standard error measured at the candidate, not at the minimum
  • Picks the smallest lambda inside the band
  • Calls the one-standard-error band a 5% significance test
  • Thinks a larger lambda means a more complex model
  • Keeps the error threshold unflipped when scoring by AUC

context

open as a page

A penalised model's CV-error curve is flat across a whole decade of lambda — why is the lambda with the lowest mean error a poor pick?

level: seniorimportance: should knowfreq 34%

basics

~10 s

On a flat curve the winner is decided by fold noise, not by any real difference: reshuffle the folds and the minimum jumps to a neighbouring lambda. That lowest mean is also optimistically low.

open as a page

When would you ship the CV-minimum lambda instead of the larger one-standard-error lambda?

level: principalimportance: nice to knowfreq 16%

basics

~10 s

Take the minimum when accuracy is the entire objective and the dip is deep relative to the error bars. The one-standard-error rule spends a little expected accuracy to buy a simpler, more reproducible model.

open as a page