In scikit-learn, how do you run nested cross-validation around a GridSearchCV?
answer
- a search is itself an estimator
- the maximum of noisy estimates is biased
- inner picks, outer measures
- folds may disagree on best_params_
- estimates the procedure, not one model
basics
~20 sPass the unfitted search as the estimator to an outer cross-validation: cross_val_score(GridSearchCV(est, grid, cv=inner), X, y, cv=outer). Each outer fold tunes on its own training portion and is scored on data no tuning decision ever saw.
solid answer
~50 s`GridSearchCV` and `RandomizedSearchCV` are estimators, so they can be the estimator argument of another cross-validation: ``` inner = StratifiedKFold(5, shuffle=True, random_state=0) outer = StratifiedKFold(5, shuffle=True, random_state=1) search = GridSearchCV(pipe, grid, cv=inner, scoring='roc_auc') scores = cross_val_score(search, X, y, cv=outer, scoring='roc_auc') ``` On each outer fold the search is cloned, runs its whole inner search on that fold's training rows, and the winner is scored on the outer test rows — data that took part in no selection decision. The mean of `scores` estimates the **procedure**, not one model, and the outer folds may well choose different hyperparameters, which is informative in itself. The reason to bother: `search.best_score_` is the maximum over many candidates evaluated on the same folds, so selection noise inflates it. Cost is the product of the two loops — five outer times five inner times the grid — which is why people reserve it for small data, where a single hold-out would be too noisy to trust.
code
python · 19 linesfrom sklearn.datasets import load_breast_cancer
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import (GridSearchCV, StratifiedKFold,
cross_validate)
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
X, y = load_breast_cancer(return_X_y=True)
pipe = Pipeline([("scale", StandardScaler()), ("clf", LogisticRegression(max_iter=5000))])
grid = {"clf__C": [0.01, 0.1, 1, 10, 100]}
inner = StratifiedKFold(5, shuffle=True, random_state=0)
outer = StratifiedKFold(5, shuffle=True, random_state=1)
search = GridSearchCV(pipe, grid, cv=inner, scoring="roc_auc")
res = cross_validate(search, X, y, cv=outer, scoring="roc_auc", return_estimator=True)
print(res["test_score"].mean(), res["test_score"].std())
print([e.best_params_ for e in res["estimator"]])go deeper
Know that the score a hyperparameter search reports is the score of the candidate the search itself chose, so it flatters the model, and that a separate untouched set is needed for an honest number.
Write the construction: a search object passed as the estimator to cross_val_score with a different outer splitter, and explain that inner folds select while outer folds measure.
Judge when the cost is justified, read the fold-to-fold disagreement in best_params_ as a signal, and insist the estimator be a Pipeline so no transform is fitted across the outer boundary.
Decide the evaluation protocol for the organization: how much compute honest estimation is worth, what claims are permitted from a search score versus a nested estimate, and how model comparisons are made defensible on small data.
## The problem nested CV solves Run a search over 200 candidates with 5-fold CV and take `best_score_`. That number is a maximum over 200 estimates that share the same folds. Even if every candidate were equally good, the winner would be the one whose noise happened to be most favourable on those particular folds. The maximum of noisy estimates is biased upward, and the bias grows with the number of candidates and shrinks with dataset size. On a few hundred rows with a wide grid it can be several points. So `best_score_` answers "how well did the winner do on the folds that chose it" — which is not a generalization estimate. It is the same fallacy as tuning against a test set and then reporting that test score. ## The construction Nested cross-validation puts the *entire selection procedure* inside a cross-validation loop. The inner loop picks hyperparameters; the outer loop measures what the picking produces. ``` from sklearn.model_selection import GridSearchCV, StratifiedKFold, cross_val_score inner = StratifiedKFold(n_splits=5, shuffle=True, random_state=0) outer = StratifiedKFold(n_splits=5, shuffle=True, random_state=1) search = GridSearchCV(pipeline, param_grid, cv=inner, scoring='roc_auc', n_jobs=-1) scores = cross_val_score(search, X, y, cv=outer, scoring='roc_auc') print(scores.mean(), scores.std()) ``` Because `cross_val_score` clones its estimator per fold, each outer fold gets a fresh search that has never seen that fold's test rows. Nothing needs to be wired by hand — this is the whole implementation. Use `cross_validate(..., return_estimator=True)` instead when you want to inspect what each outer fold chose; each returned object is a fitted search with its own `best_params_`. ## Interpreting the output Three things the result tells you, and one it does not: 1. **The mean is an estimate of the procedure.** "If I run this preprocessing, this model family, this grid and this inner CV on data like this, expect roughly this score." That is genuinely what you deploy, since you will re-run the tuning on the full data. 2. **The spread across outer folds is a stability signal.** Wide spread on small data means any single hold-out estimate would have been close to arbitrary. 3. **Disagreement between outer folds' `best_params_` is diagnostic.** If five folds pick five different depths and score the same, depth does not matter much and you can simplify. If they agree, the choice is real. 4. **It does not hand you a model.** There is no single winner to ship, and averaging hyperparameters is meaningless. After the estimate, refit *one* search on all the data and deploy its `best_estimator_`; nested CV is the honest number attached to that object, not a substitute for producing it. ## Cost, and the alternative The fit count multiplies: outer folds times inner folds times candidates, times the cost of each fit. Five by five over a 200-point grid is 5,000 fits. That is fine for a few thousand rows and a linear model, and impossible for a large gradient-boosted model on millions of rows. When data is plentiful, the simpler protocol is usually better: split off a test set once, tune with cross-validation on the remainder, evaluate once at the end. With enough rows a single hold-out is precise enough that nested CV's variance reduction is not worth the compute. Nested CV earns its cost precisely where hold-outs are too small to trust — a few hundred rows, medical or scientific datasets, model-family comparisons where a couple of points decide the argument. ## Details that make it correct - **The estimator must be a `Pipeline` if any preprocessing is fitted.** A scaler, imputer, feature selector or target encoder fitted outside the pipeline is fitted on data including the outer test rows, and the leak survives the nesting entirely. Nested CV protects against selection bias, not against a transform fitted at the wrong scope. - **Use different splitters for the loops.** Sharing one shuffled splitter object between inner and outer is not a correctness bug — the inner splitter only ever sees the outer training rows — but distinct seeds make the independence obvious to a reader. - **Group and time structure still apply** at both levels: if the data is grouped, both loops must be group-aware, and `groups` must be forwarded to the outer call. - **`scoring` should match** between the loops unless you deliberately want to select on one criterion and report another. - **`error_score`** defaults to NaN, so a broken candidate can turn an outer fold's score into NaN; set it to `'raise'` while developing.
- After nested CV, which model do you actually deploy?None of the outer-fold models. Refit a single search on the entire dataset and ship its `best_estimator_`; the nested mean is the honest performance estimate you attach to it. Averaging hyperparameters across folds is meaningless, and picking the best-scoring outer fold reintroduces exactly the selection bias the nesting removed.
- The five outer folds choose five different values of the same hyperparameter. What does that tell you?That the score surface is flat in that direction — the differences the inner loop is resolving are within noise. Practically, simplify: fix it at a defensible value, or narrow the grid so the search spends its budget on parameters that move the score. Divergent choices with divergent scores are a different story and point at heterogeneity across folds.
- When is a single train/validation/test split preferable to nested CV?When data is plentiful. With tens of thousands of rows a hold-out estimate is already precise, the protocol is far cheaper and much easier to explain, and it maps onto how the model is actually retrained. Nested CV pays off on small data where a single hold-out's variance would swamp the difference you are trying to measure.
- Does nested CV protect you from a StandardScaler fitted before the outer split?No. Nesting only removes bias from hyperparameter selection. A transform fitted over all rows has already absorbed statistics from every outer test fold, and that leak passes straight through. The fix is orthogonal: put every fitted transform inside a Pipeline so it is refitted within each fold.
saying these in an interview costs you the question
- Reports GridSearchCV.best_score_ as the generalization estimate
- Thinks nested CV produces a single model to deploy
- Averages the chosen hyperparameters across outer folds
- Fits preprocessing outside the pipeline and calls the nesting leak-proof
- Uses nested CV on a large dataset where a hold-out would do