Cross-Validation and Model Selection
You will learn to estimate generalization honestly with the right CV scheme and to tune hyperparameters without touching the test set. Interviewers probe leakage: it silently inflates offline results.
on this pageshowhide
explore
- Resampling Schemes16 questions
- Train/Validation/Test Splits5 questions
- K-Fold and LOOCV4 questions
- Stratified and Grouped Folds4 questions
- Bootstrap Error Estimates3 questions
- Honest Generalization Estimates11 questions
- Nested Cross-Validation4 questions
- Selection Bias and Reporting3 questions
- Refitting the Final Model4 questions
- Hyperparameter Search9 questions
- Grid and Random Search4 questions
- Bayesian Search and Hyperband5 questions
- Data Leakage10 questions
- Target Leakage3 questions
- Preprocessing Before Splitting3 questions
- Temporal and Group Overlap4 questions
questions
page 2 of 2Why does a per-city average demand computed over the full three-year history leak, even with forward-chaining folds?
basics
~20 sThe aggregate summarises the entire timeline, so a row from year one already carries year three's demand. Every training row sees the future and every validation row sees its own outcomes, no matter how correctly the folds are ordered.
Would you choose Bayesian search or Hyperband for a four-hour run on 16 parallel workers?
basics
~20 sPrefer Hyperband when a cheap fidelity ranks configurations like full training and workers are plentiful; prefer Bayesian search when evaluations are expensive and few. Hybrids sample configurations with a surrogate and schedule them with brackets.
With a fixed 60-trial budget and six hyperparameters, how do you decide what to search?
basics
~20 sStart from the compute you have, not the size of the space. Fix the hyperparameters that rarely matter at defaults and give the whole budget to the two or three that plausibly move the score.
When is skipping nested cross-validation for one fixed tuning split defensible?
basics
~20 sSkip nesting when a single large tuning split is already precise — millions of rows, few candidates, a decision with a wide margin. Keep it when data is scarce or wide and many settings are compared, where the optimism is largest.
Your team has scored the same hold-out test set 40 times this quarter - what now?
basics
~20 sThat hold-out is no longer a test set: forty scored comparisons made it a validation set, and its number is optimistic by an unmeasurable amount. Get fresh untouched rows for the real verdict, and budget future use.
How do you stratify cross-validation folds when the regression target is continuous?
basics
~20 sTurn the continuous target into a temporary discrete label — quantile bins such as quartiles — and stratify the folds on that bin. Every fold then carries a similar spread of target values, steadying fold metrics on small or skewed data.
Should a weekly hyperparameter search be warm-started from last week's best configuration?
basics
~20 sWarm-start by seeding the surrogate with past trials, not by trusting last week's winner. Old scores do not transfer to new data, so re-evaluate the incumbent, keep some random exploration, and cold-start periodically to catch a moved optimum.
Why is bootstrap left-out error pessimistic, and what does the .632 estimator correct?
basics
~20 sEach bootstrap replicate trains on only about 63% distinct rows, so its models are weaker and the error on left-out rows comes out too high. The .632 estimator blends that pessimistic error with the optimistic training error, 0.632 to 0.368.
How do you search a space where an RBF kernel has a gamma but a linear kernel has none?
basics
~20 sTreat the space as conditional rather than rectangular: sample the kernel first, and only draw the width parameter when the kernel that uses it was chosen. A flat cross-product instead fits the same linear model many times over.
What do you do when 5-fold cross-validation scores swing by 0.08 between random partitions?
basics
~20 sTreat that swing as a measurement of your own noise floor. Repeat the whole cross-validation with several different random partitions and average across all runs, and refuse to call any model difference smaller than the swing you just observed.
Each outer fold of a nested cross-validation picks different hyperparameters — is it broken?
basics
~10 sNo. Nested cross-validation estimates a tune-then-fit procedure, not one fixed setting, so outer folds are free to disagree. Disagreement usually means the score is flat across nearby candidates or the inner estimates are noisy.
When two tree depths score within one standard error in CV, why pick the shallower tree?
basics
~20 sA gap smaller than one standard error is indistinguishable from fold noise, so the two are tied on evidence. Among tied candidates the simpler one usually varies less on new data and is cheaper to run, explain and maintain.
Should you refit on train plus validation once the hyperparameters are chosen?
basics
~20 sUsually yes. Once the configuration is frozen, the validation rows are just labelled data, and training on more rows generally gives a better model. Report the test score for the refit model, and never redo selection after folding validation in.
Instead of refitting after cross-validation, should you average the k fold models into one predictor?
basics
~20 sA defensible option, not the default. Averaging the fold models gives a variance-reduced ensemble at no extra training cost, but costs k times the inference work, leaves no single auditable artifact, and was never scored as a unit.
A nightly-rebuilt snapshot table reflects today's state, not decision-time state. How do you train on it safely?
basics
~20 sTreat every mutable column as leaked until proven otherwise. A nightly rebuild overwrites fields with post-outcome values, so train only on columns that cannot change after the decision, or reconstruct earlier values from an append-only change history.
When a dataset has both repeated customers and a time order, how do you design the validation split?
basics
~20 sStart from what the model faces at inference. Break the axis production does not repeat: hold out unseen customers if it scores strangers, hold out later time if it re-scores known customers, and hold out both when it does both.
showing 31–46 of 46