skip to content

Cross-Validation and Model Selection

You will learn to estimate generalization honestly with the right CV scheme and to tune hyperparameters without touching the test set. Interviewers probe leakage: it silently inflates offline results.

on this pageshow

explore

questions

page 2 of 2

Why does a per-city average demand computed over the full three-year history leak, even with forward-chaining folds?

level: seniorimportance: should knowfreq 44%

basics

~20 s

The aggregate summarises the entire timeline, so a row from year one already carries year three's demand. Every training row sees the future and every validation row sees its own outcomes, no matter how correctly the folds are ordered.

open as a page

Would you choose Bayesian search or Hyperband for a four-hour run on 16 parallel workers?

level: principalimportance: should knowfreq 30%

basics

~20 s

Prefer Hyperband when a cheap fidelity ranks configurations like full training and workers are plentiful; prefer Bayesian search when evaluations are expensive and few. Hybrids sample configurations with a surrogate and schedule them with brackets.

open as a page

With a fixed 60-trial budget and six hyperparameters, how do you decide what to search?

level: principalimportance: should knowfreq 38%

basics

~20 s

Start from the compute you have, not the size of the space. Fix the hyperparameters that rarely matter at defaults and give the whole budget to the two or three that plausibly move the score.

open as a page

When is skipping nested cross-validation for one fixed tuning split defensible?

level: principalimportance: should knowfreq 38%

basics

~20 s

Skip nesting when a single large tuning split is already precise — millions of rows, few candidates, a decision with a wide margin. Keep it when data is scarce or wide and many settings are compared, where the optimism is largest.

open as a page

Your team has scored the same hold-out test set 40 times this quarter - what now?

level: principalimportance: should knowfreq 40%

basics

~20 s

That hold-out is no longer a test set: forty scored comparisons made it a validation set, and its number is optimistic by an unmeasurable amount. Get fresh untouched rows for the real verdict, and budget future use.

open as a page

How do you stratify cross-validation folds when the regression target is continuous?

level: middleimportance: nice to knowfreq 30%

basics

~20 s

Turn the continuous target into a temporary discrete label — quantile bins such as quartiles — and stratify the folds on that bin. Every fold then carries a similar spread of target values, steadying fold metrics on small or skewed data.

open as a page

Should a weekly hyperparameter search be warm-started from last week's best configuration?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

Warm-start by seeding the surrogate with past trials, not by trusting last week's winner. Old scores do not transfer to new data, so re-evaluate the incumbent, keep some random exploration, and cold-start periodically to catch a moved optimum.

open as a page

Why is bootstrap left-out error pessimistic, and what does the .632 estimator correct?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Each bootstrap replicate trains on only about 63% distinct rows, so its models are weaker and the error on left-out rows comes out too high. The .632 estimator blends that pessimistic error with the optimistic training error, 0.632 to 0.368.

open as a page

How do you search a space where an RBF kernel has a gamma but a linear kernel has none?

level: seniorimportance: nice to knowfreq 25%

basics

~20 s

Treat the space as conditional rather than rectangular: sample the kernel first, and only draw the width parameter when the kernel that uses it was chosen. A flat cross-product instead fits the same linear model many times over.

open as a page

What do you do when 5-fold cross-validation scores swing by 0.08 between random partitions?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Treat that swing as a measurement of your own noise floor. Repeat the whole cross-validation with several different random partitions and average across all runs, and refuse to call any model difference smaller than the swing you just observed.

open as a page

Each outer fold of a nested cross-validation picks different hyperparameters — is it broken?

level: seniorimportance: nice to knowfreq 25%

basics

~10 s

No. Nested cross-validation estimates a tune-then-fit procedure, not one fixed setting, so outer folds are free to disagree. Disagreement usually means the score is flat across nearby candidates or the inner estimates are noisy.

open as a page

When two tree depths score within one standard error in CV, why pick the shallower tree?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A gap smaller than one standard error is indistinguishable from fold noise, so the two are tied on evidence. Among tied candidates the simpler one usually varies less on new data and is cheaper to run, explain and maintain.

open as a page

Should you refit on train plus validation once the hyperparameters are chosen?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Usually yes. Once the configuration is frozen, the validation rows are just labelled data, and training on more rows generally gives a better model. Report the test score for the refit model, and never redo selection after folding validation in.

open as a page

Instead of refitting after cross-validation, should you average the k fold models into one predictor?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

A defensible option, not the default. Averaging the fold models gives a variance-reduced ensemble at no extra training cost, but costs k times the inference work, leaves no single auditable artifact, and was never scored as a unit.

open as a page

A nightly-rebuilt snapshot table reflects today's state, not decision-time state. How do you train on it safely?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Treat every mutable column as leaked until proven otherwise. A nightly rebuild overwrites fields with post-outcome values, so train only on columns that cannot change after the decision, or reconstruct earlier values from an append-only change history.

open as a page

When a dataset has both repeated customers and a time order, how do you design the validation split?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Start from what the model faces at inference. Break the axis production does not repeat: hold out unseen customers if it scores strangers, hold out later time if it re-scores known customers, and hold out both when it does both.

open as a page

showing 31–46 of 46