Each outer fold of a nested cross-validation picks different hyperparameters — is it broken?
answer
- what exactly is being estimated?
- a recipe, not a single setting
- flat objective or noisy inner scores
- never grade the modal winner
basics
~10 sNo. Nested cross-validation estimates a tune-then-fit procedure, not one fixed setting, so outer folds are free to disagree. Disagreement usually means the score is flat across nearby candidates or the inner estimates are noisy.
solid answer
~50 sDisagreement across outer folds is expected, not a bug. Each outer fold runs the whole selection procedure on its own training portion, and the number you take away is an estimate of that procedure's performance — not of any single hyperparameter vector. Two benign causes dominate: the objective is flat, so several near-tied settings win by luck of the fold, or the inner estimates are noisy enough that the ranking is unstable. I would look at the inner scores to tell those apart, and at whether the outer scores themselves are stable: settings that vary while outer performance stays steady is a flat objective and harmless; both varying wildly means the pipeline is fragile and the mean is a weak summary. What you must not do is pick the setting that won most outer folds and quote the outer mean as its score — that setting was chosen using the outer scores, which puts the optimism straight back in.
go deeper
Remember that the outer loop measures rather than chooses, so different folds ending up with different settings is normal. You are not expected to diagnose why, only not to call it a bug.
Explain what the outer mean actually estimates — the performance of the search-and-fit procedure — and name the two ordinary causes of disagreement, a flat objective and noisy inner rankings.
Demonstrate the diagnosis using the numbers you already have: compare the spread of inner candidate scores against fold noise, check whether the outer scores are stable too, and refuse the trap of grading the modal winner with the outer mean.
Own what gets communicated: a single headline number hides a procedure whose output varies, so decide what your organisation reports alongside it and when fold-level instability should block a launch rather than be averaged away.
## Nested cross-validation estimates a recipe The most common misunderstanding about nested cross-validation is that it is a fancier way to choose hyperparameters. It is not. The inner loops choose; the outer loop measures. What the outer scores measure is the performance of the *whole procedure* — "take a training set of about this size, run this search over this candidate set, fit the winner" — applied to fresh data. A procedure is allowed to produce different outputs on different inputs. If it did not, you would be evaluating a fixed setting, which is a different and easier experiment. So five outer folds returning five different winners is a description of your search's behaviour, not a defect in the evaluation. ## Two benign causes and one worrying one **A flat objective.** Several candidates sit within noise of one another on the criterion you are optimising. Whichever one wins is decided by fold-level luck. This is the usual case, and it is good news: it means the choice does not matter much, and it means your reported performance is robust to it. **Noisy inner estimates.** Small inner folds, a rare positive class, or a high-variance metric make the inner ranking unstable even when the candidates genuinely differ. Here the disagreement says your inner loop is under-powered; more inner folds, repeated inner cross-validation or a lower-variance criterion would stabilise it. **Genuine instability.** If both the winning settings and the outer scores swing widely, the whole pipeline is sensitive to which rows it sees. That is a real finding and it deserves attention before anything ships: the mean of the outer scores is then a poor summary of a wide distribution, and the honest report is the spread as well as the centre. Distinguishing them costs nothing: you already computed the inner score tables and the outer scores; look at them. ## What you must not conclude - **Do not pick the modal winner and quote the outer mean for it.** The moment you use the outer scores to choose, they stop being an unbiased estimate of that choice. You have rebuilt the flat scheme with extra steps. - **Do not rerun with different seeds until the folds agree.** Agreement obtained by search is not evidence; it is the same selection bias applied to fold assignment. - **Do not conclude the model is unusable.** A flat objective producing different winners is the most common outcome of a well-behaved search over a sensible candidate set. ## What the disagreement is worth as diagnostic information It tells you which hyperparameters matter. If the tree depth is identical in every fold but the learning rate wanders, depth is doing real work and the learning rate is in a flat region — useful for shrinking the candidate set on the next iteration. If one fold picks a setting from the edge of the range while the others cluster in the middle, the range may be too narrow, or that fold may contain an unusual subgroup worth inspecting. If a fold's winner is systematically more regularised, look at whether that fold's training portion is smaller or its class balance different. ## The reporting question What you report from a nested run is the outer estimate of the procedure, together with how much it varied across folds. What you deploy is decided separately, after the estimate is in hand; producing the shipped model is its own step with its own considerations, and nested cross-validation deliberately does not answer it. Keeping those two questions apart in an interview answer is the clearest signal that you have actually run this rather than read about it. ## A short script for the interview Say: the outer loop grades the procedure, so disagreement is expected; here is how I tell a flat objective from an unstable pipeline; here is the trap of grading the modal winner with the outer mean; and here is what the pattern of disagreement tells me about which hyperparameters actually matter. That covers the mechanism, the diagnosis and the failure mode in under two minutes.
- How do you tell a flat objective from a genuinely unstable pipeline?Look at both score tables. If the inner scores of the top candidates fall within the fold-to-fold noise and the outer scores stay steady, the objective is flat and the disagreement is harmless. If the outer scores swing widely as well, the pipeline is sensitive to which rows it sees, and the mean of the outer folds is a weak summary of a wide distribution.
- Can you report the outer mean as the score of whichever setting won the most outer folds?No. That setting was identified by looking at the outer results, so the outer scores are no longer independent of the choice and the estimate becomes optimistic again — the exact failure nesting exists to prevent. The outer mean belongs to the procedure. A specific setting needs its own untouched evaluation data.
- What does it mean when one hyperparameter is identical across all outer folds and another varies?The stable one is doing real work: the data supports the same value regardless of which rows are held out. The varying one sits in a flat region of the objective, so its exact value has little effect. That is a useful signal for shrinking the candidate set on the next iteration and for deciding what to tune carefully.
saying these in an interview costs you the question
- Calls nested cross-validation broken because folds disagree
- Quotes the outer mean as the score of the modal winning setting
- Reruns with new seeds until the outer folds agree
- Assumes disagreement always means the dataset is too small
- Thinks the outer loop is supposed to select hyperparameters