When two tree depths score within one standard error in CV, why pick the shallower tree?
answer
- a tie is not a win
- the cheapest candidate that ties
- one standard error of slack
- simplest model inside the noise band
basics
~20 sA gap smaller than one standard error is indistinguishable from fold noise, so the two are tied on evidence. Among tied candidates the simpler one usually varies less on new data and is cheaper to run, explain and maintain.
solid answer
~50 sThis is the one-standard-error rule. You find the best cross-validated score, compute the standard error of that estimate across folds, then choose the **simplest** candidate still within one standard error of the best — for an error curve, the simplest model with error at or below `min_error + 1 SE`. If depth 4 scores 0.808 and depth 9 scores 0.812 with a standard error around 0.027, the 0.004 gap is a fraction of the noise, so take depth 4. Tied on measured skill, the simpler model tends to have lower variance, degrade more gracefully on shifted data, and cost less to serve and explain. It also pulls you off the exact maximum of the curve, which is its most noise-inflated point. It is a heuristic, not a theorem: one SE is a convention, and it needs a meaningful complexity ordering to apply at all.
go deeper
Know the idea in plain terms: if two settings score within the measurement noise of each other, they are tied, and the smaller or simpler one is the safer default.
Be able to state the rule mechanically and get the direction right: compute the best score's standard error, take the simplest candidate inside that band, moving up from a minimum error or down from a maximum score.
Demonstrate the judgment around it — inspect the whole CV curve before trusting the band, know that the k-fold standard error is understated, and be able to say when a flat curve makes the rule underfit.
Own the exchange rate the rule silently assumes: it trades a fraction of a point of accuracy for robustness and operating cost, and only you can say whether that trade is right for this product.
## The rule, stated precisely Given a set of candidates ordered by complexity, each scored by k-fold cross-validation: 1. Find the best mean CV score (or lowest mean CV error). 2. Compute the standard error of that best candidate's CV estimate — the standard deviation of its fold scores divided by the square root of the number of folds. 3. Among **all** candidates whose mean score is within one standard error of the best (equivalently, whose error is at or below `min_error + 1 SE`), select the **simplest**. Note the direction carefully: with an error metric you move *up* one SE from the minimum and take the simplest candidate under that ceiling; with a score metric you move *down* one SE from the maximum and take the simplest candidate above that floor. Getting this backwards selects the most complex tied model, which is the opposite of the rule's intent. ## Why a tie is not a win Suppose depth 4 scores 0.808 and depth 9 scores 0.812 across five folds, with a fold standard deviation of 0.06 — a standard error near 0.027. The observed 0.004 advantage is under a sixth of the estimate's own noise. Re-run the whole thing with a different fold split and the ordering will often flip. Declaring depth 9 the winner is reading a coin flip as a measurement. Once the two are tied on the evidence, the choice must be made on something else — and complexity is the natural tiebreaker, for three reasons. **Variance.** A deeper tree fits finer structure, much of which is sample-specific. Two models with the same expected accuracy on this data distribution are not equally safe: the higher-capacity one has the wider spread of possible fitted models and typically the worse behaviour when the deployment distribution drifts a little from the training one. **Selection bias.** The maximum of the CV curve is, by construction, the point most flattered by noise — the same winner's-curse logic that inflates the best of any large candidate set. Deliberately stepping off the argmax and toward the simple end trades a fraction of an SE of apparent performance for an estimate that is less a product of the fold split. **Cost.** Shallower trees are faster to train and serve, easier to explain to a reviewer or a regulator, and less prone to surprising behaviour on rare inputs. If accuracy is tied, these are free wins. ## Where the rule is weak - **It needs a complexity ordering.** "Simplest" is only defined if the candidates line up on a single axis — depth, number of leaves, number of features, number of components. Across heterogeneous families (a shallow tree versus a linear model versus a boosted ensemble) there is no canonical ordering, so you must state the axis you are using — parameter count, inference latency, explainability — or the rule does not apply. - **One SE is a convention, not a derivation.** There is nothing sacred about one; it is a pragmatic amount of slack chosen because it is roughly the resolution of the measurement. - **The SE itself is understated.** For k-fold, the fold scores are correlated because the training sets overlap, so `s / sqrt(k)` is a floor on the real uncertainty. That cuts both ways: the true tie region is wider than one nominal SE, which is an argument for the rule's spirit rather than against it. - **Flat curves let it slide too far.** If the CV curve is nearly flat over a wide complexity range, one SE of slack can carry you deep into underfitting territory. Always look at the curve; if the simplest candidate within the band is drastically simpler than the best, sanity-check its absolute performance rather than accepting it mechanically. - **Sometimes small gains are worth money.** If a fraction of a point maps to real revenue or real harm avoided, and the complex model is affordable to run and maintain, the rule's implicit preference for simplicity is the wrong exchange rate. Say so explicitly rather than following the heuristic off a cliff. ## What it does and does not give you It gives you a defensible, reproducible tiebreak that biases toward robustness and away from the noisiest point on the curve. It does **not** give you an unbiased estimate of the chosen model's performance — the score of the model you selected is still a selection statistic, and if you need to quote a number, it has to come from data that took no part in the selection. ## The interview answer in one breath "A difference smaller than the estimate's own standard error is not a difference. I take the simplest candidate inside the one-SE band because tied models should be separated on variance and cost, not on noise — and because the exact argmax is the most noise-inflated point on the curve. It is a heuristic: it needs a complexity axis, and I'd override it if the gain, however small, were worth real money."
- How do you define simplest when the tied candidates come from different model families?You cannot, unless you supply the axis yourself. The rule assumes a one-dimensional complexity ordering. Across families I name the criterion explicitly — inference latency, number of features to maintain, explainability to a reviewer — and say I am tie-breaking on that, rather than pretending there is a canonical notion of simpler between a shallow tree and a linear model.
- Does the rule fix the optimism in the score you finally report?No. It moves the choice off the argmax, which is the most noise-inflated point, so the selected candidate's own CV score is a little less flattered. But that score was still produced by data used in the selection, so it remains a selection statistic. Any number you publish still has to come from data that played no role in choosing the model.
- When would you deliberately override the rule and take the complex model?When the gain, even if small and uncertain, maps to material value and the complexity is genuinely affordable — you can serve it, monitor it and explain it. Also when the CV curve is so flat that one SE of slack drags you into a candidate that is obviously underfit in absolute terms. In both cases I record the override and the reason.
Two job candidates finish within the margin of error on a noisy test. You do not crown the half-point winner; you break the tie on the things you can actually measure reliably, like cost and reliability.
saying these in an interview costs you the question
- Reads a 0.004 CV gap as a real improvement
- Applies the rule without ever computing a standard error
- Claims the rule guarantees better test performance
- Picks the simplest candidate overall, ignoring the one-SE band
- Moves the band the wrong way and keeps the most complex tied model