In nested cross-validation, what do the inner and outer loops each do?
answer
- two jobs, one dataset
- selection inside, scoring outside
- the held-out fold never votes
- fits multiply by the outer fold count
basics
~20 sThe inner loop picks hyperparameters using only the training part of each outer split. The outer loop then scores that already-tuned model on data the inner loop never saw, giving an honest estimate of the whole tuning procedure.
solid answer
~50 sNested cross-validation splits the data twice. The outer loop makes k evaluation folds; for each one you hold out the outer test fold and pass only the remaining data into an inner loop, which runs its own cross-validation over the candidate settings and returns the best one. You then train on the whole outer training portion with that setting and score once on the untouched outer test fold. Repeating for every outer fold gives k scores whose mean estimates how well the entire tune-then-fit procedure generalises. The load-bearing property is that no outer test fold ever influences the choice it is later used to judge: selection happens strictly inside, evaluation strictly outside. Flat cross-validation collapses the two loops into one, so the same folds both choose the winner and grade it, which pushes the reported score upward.
go deeper
Be ready to say that the settings are chosen on one part of the data and the score is taken on another part that had no say in the choice. Knowing that much already puts you ahead of most screening answers.
Explain the loop structure end to end without prompting: outer split, inner search on the outer training portion only, fit the winner, score the held-out fold, repeat. Be ready to count the fits and to say why the cost multiplies.
Specify it for a concrete dataset: fold counts for each loop, stratification and grouping honoured in both, the metric each loop optimises, and the compute you are committing to. Interviewers probe for the implementation slips that quietly break the isolation.
Own the call about when this estimate is worth several times the compute, which number the organisation is allowed to publish from it, and how you keep the honest estimate from being replaced by a prettier one further down the reporting chain.
## Two jobs people ask one loop to do Cross-validation gets used for two different jobs. **Selection**: of the candidate settings I am considering, which performs best? **Estimation**: how well will the thing I end up with perform on data nobody has seen? One cross-validation loop can answer either question honestly. It cannot answer both at once with the same folds, because the answer to the first is obtained by *maximising over* the very numbers you would then quote for the second. Nested cross-validation gives each job its own loop. ## The procedure, step by step 1. Split the data into `k_outer` folds. 2. For outer fold `i`: hold out fold `i` as the **outer test fold**; everything else is the **outer training portion**. 3. Run a complete cross-validation *inside that outer training portion only*: split it into `k_inner` folds, score every candidate setting on those inner folds, average, and keep the best setting. Nothing here has ever seen outer fold `i`. 4. Fit one model on the entire outer training portion using the winning setting. 5. Score it once on outer fold `i`. 6. Repeat for every `i`. You finish with `k_outer` scores. ## What those outer scores actually measure Each outer score grades a model produced by running your *entire* selection procedure — search included — on a training set of that size. Their mean therefore estimates the generalisation performance of the **procedure**, not of one specific hyperparameter vector. This is where most confusion starts: nested cross-validation does not hand you a setting or a model to deploy, it hands you a defensible number for the recipe. Producing the model you actually ship is a separate step. ## Why the isolation is load-bearing Suppose you skip the inner loop, score all candidates on the same `k` folds and report the best average. Every fold score carries sampling noise. Taking the maximum over candidates partly selects the setting whose noise happened to land favourably on those particular folds, so the reported number contains a positive selection component that will not reappear on fresh data. Isolation removes the mechanism: fold `i` cannot inflate a score for a choice it took no part in. ## Cost accounting For `c` candidate settings, the inner loops perform `k_outer * k_inner * c` fits, plus `k_outer` refits of each fold's winner. A 5-by-5 nesting over 20 candidates is `5 * 5 * 20 = 500` inner fits plus 5 refits — 505, against the 100 a flat 5-fold search over the same candidates would cost. The multiplier is exactly `k_outer`. ## Choosing the two fold counts They are independent choices. The outer count drives the precision of the reported estimate and the size of each training portion. The inner count only has to rank candidates well enough to pick a good one, so an economy such as 3 inner folds inside 5 or 10 outer folds is common and usually harmless. Anything structural must be honoured in **both** loops. Stratify both if the classes are imbalanced. Keep records that belong to the same group — the same patient, user or session — together in both loops if the data is clustered. Respect time order in both if the data is a time series. A grouping constraint applied only to the outer split leaks inside the inner search and quietly destroys the honesty you paid for. ## Common ways it is implemented wrongly - Running the search once on all the data and then "nesting" only the final evaluation. The search has already seen every outer test fold. - Reusing one inner split across outer folds to save time, when that split overlaps the outer test folds. - Reporting the best outer fold rather than the mean of the outer folds. - Letting the inner loop optimise one metric while the outer loop reports another, then presenting the outer number as though the procedure had targeted it. ## When both loops agree anyway If training sets are enormous relative to what the model can absorb and you are comparing a handful of settings, the inner estimates are so precise that maximising over them barely bites; flat and nested estimates then agree to within a rounding error, and the extra `k_outer`-fold compute buys nothing. The gap is widest exactly where data is scarce, the feature table is wide and the candidate set is large.
- How many model fits does a 5-by-5 nesting cost for 20 candidate settings?Each outer fold runs 5 inner folds over 20 candidates, so 100 inner fits, plus one refit of the winner on the outer training portion — 101 per outer fold. Across five outer folds that is 505 fits, versus 100 for a flat 5-fold search over the same candidates. The multiplier is the outer fold count.
- Do the inner and outer loops have to use the same number of folds?No, they are independent. The outer count sets the precision of the reported estimate; the inner count only has to rank candidates well enough to pick a good one, so three inner folds inside five or ten outer folds is a common economy. Stratification and grouping constraints, however, must be applied identically in both loops.
- What goes wrong if you tune once on all the data and then run a plain cross-validation to report the score?Nothing is nested at all. The search has already seen every fold that will later be used as a test fold, so each reported fold score is graded by data that helped choose the setting. The number is optimistic in exactly the way nesting exists to prevent.
The inner loop is the audition and the outer loop is opening night with a critic who was not at the audition. If the critic also picked the cast, the review tells you less than it appears to.
saying these in an interview costs you the question
- Says the outer loop tunes and the inner loop evaluates
- Expects nested cross-validation to return one final hyperparameter setting
- Reuses the outer test fold inside the inner search to save time
- Applies grouping or stratification only to the outer split
- Reports the best outer fold instead of the mean across outer folds