What splitter does scikit-learn use when you pass cv=5 to cross_val_score?
answer
- integer cv is expanded for you
- classifier plus discrete y is special
- check_cv decides, not you
- shuffle defaults to False
- row order becomes fold structure
basics
~20 sAn integer cv is expanded for you: StratifiedKFold when the estimator is a classifier and the target is binary or multiclass, plain KFold otherwise. Neither shuffles — folds are contiguous blocks of the data in its current row order.
solid answer
~50 s`cv` accepts an integer, a splitter object, or an iterable of `(train_idx, test_idx)` arrays. When you pass an integer, scikit-learn resolves it through `check_cv`: if the estimator is a classifier *and* `y` is binary or multiclass, you get `StratifiedKFold(n_splits=cv)`; in every other case — regressors, clustering, multilabel or continuous targets — you get `KFold(n_splits=cv)`. The default when `cv=None` is 5. The part that bites is that both are constructed with `shuffle=False`. Folds are consecutive slices of the rows exactly as they arrive, so any ordering in your file — sorted by label, by date, by customer — becomes the fold structure. If you want shuffling you must pass the object yourself, `KFold(n_splits=5, shuffle=True, random_state=0)`, and then the seed becomes part of your reported score. Shuffling is also exactly the wrong move for time-ordered or grouped data.
code
python · 12 linesimport numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import check_cv, cross_val_score, StratifiedKFold
X = np.random.RandomState(0).randn(200, 4)
y = np.repeat([0, 1], 100) # sorted by label on purpose
print(type(check_cv(5, y, classifier=True)).__name__) # StratifiedKFold
print(type(check_cv(5, y, classifier=False)).__name__) # KFold
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
print(cross_val_score(LogisticRegression(), X, y, cv=cv))go deeper
Know that cv can be an integer or a splitter object, that 5 is the default fold count, and that classifiers get stratified folds automatically.
Explain the check_cv rule precisely — classifier plus binary/multiclass target gives StratifiedKFold, everything else KFold — and that neither shuffles, so row order becomes fold structure.
Diagnose a suspicious cross-validation number from the data's ordering, and insist on one shared, seeded splitter across all candidates so comparisons are paired rather than confounded by partition noise.
Own the evaluation protocol: which splitter the team uses, how seeds and fold assignments are recorded alongside results, and when an integer cv is simply not admissible for the data you hold.
## The three things cv accepts Every scikit-learn helper that cross-validates — `cross_val_score`, `cross_validate`, `cross_val_predict`, `GridSearchCV`, `RandomizedSearchCV`, `learning_curve`, `validation_curve` — takes the same `cv` parameter, and it accepts three shapes: 1. **An integer**, meaning "that many folds, pick the splitter for me". 2. **A splitter object** — `KFold(...)`, `StratifiedKFold(...)`, `GroupKFold(...)`, `TimeSeriesSplit(...)`, `ShuffleSplit(...)`, `LeaveOneGroupOut()`, and so on. 3. **An iterable of `(train_indices, test_indices)` pairs**, which is the escape hatch for a scheme the library does not ship. A list of two index arrays is a perfectly valid `cv` and gives you a single custom split. A float is *not* accepted; `cv=0.2` raises. That is a common reflex from `train_test_split` and it does not carry over. ## What an integer resolves to The resolution happens in `sklearn.model_selection.check_cv(cv, y, classifier=...)`, and the callers pass `classifier=is_classifier(estimator)`. The rule: - estimator is a classifier **and** the inferred target type is binary or multiclass → `StratifiedKFold(n_splits=cv)` - anything else → `KFold(n_splits=cv)` So a `RandomForestClassifier` gets stratified folds, a `RandomForestRegressor` does not, and — the subtle case — a classifier with a *multilabel* target does not either, because stratification is only defined for a single categorical column. If you assumed your multilabel run was stratified, it was not. The default number of folds is 5. It was 3 in old versions and changed to 5 in 0.22; if you read a tutorial that says three, it predates that. ## Neither default splitter shuffles This is the actual interview point. `KFold` and `StratifiedKFold` both default to `shuffle=False`, and the integer form constructs them with defaults. Fold *k* is therefore a contiguous slice of the rows in the order they sit in your array. That is a deliberate design choice — it makes results reproducible without a seed — and it is dangerous whenever row order carries information: - A CSV sorted by label gives folds that are nearly single-class. Stratification rescues the classifier case; a regressor sorted by target value gets folds covering disjoint ranges and scores that look catastrophic for no modelling reason. - A table sorted by date makes fold 1 the oldest period and fold 5 the newest — sometimes what you want, but the *other* four folds still train on the future. - Data concatenated per source or per customer produces folds that coincide with sources. The fix is to construct the splitter yourself: ``` cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0) cross_val_score(clf, X, y, cv=cv) ``` Once you shuffle, `random_state` is part of the experiment. Two engineers comparing models with different seeds are comparing noise as well as models; pin the seed and, better, pin the whole splitter object and reuse it for every candidate so all of them see identical folds. ## What stratification buys and what it does not `StratifiedKFold` keeps each fold's class distribution close to the overall one. With imbalanced data this cuts the variance of the fold scores substantially, and it prevents the pathological fold that contains zero positives, where recall is undefined and several metrics emit warnings or fall back to zero. It still splits row by row. It does not keep an entity's rows together, and it has no concept of time. Passing an integer for grouped or temporal data produces a number, produces no warning, and produces an answer that is too good. ## Related knobs worth knowing - `cross_validate` returns a dict with `fit_time`, `score_time`, `test_score`, and optionally `train_score` (`return_train_score=True`) and the fitted estimators (`return_estimator=True`). `cross_val_score` is the thin wrapper that returns just the test scores. - `n_jobs=-1` parallelizes folds. - `error_score` controls what happens when a fit raises: it defaults to `np.nan` with a warning, so a search can quietly produce NaN rows in `cv_results_`. Set `error_score='raise'` while debugging. - The scores come back as a plain array in fold order; report the mean *and* the spread, because a mean of 0.81 over folds ranging 0.62–0.95 is not the same result as one ranging 0.79–0.83.
- Your data is sorted by target value and a regressor scores terribly under cv=5. What is happening?An integer cv gives an unshuffled `KFold` for a regressor, so each fold is a contiguous block of the sorted order — the model trains on one range of the target and is tested on a disjoint range it never saw. Pass `KFold(n_splits=5, shuffle=True, random_state=0)` so folds cover the whole range, and check whether the ordering encodes something meaningful before you shuffle.
- When you shuffle, why should you build the splitter once instead of passing cv=5 to each candidate?A shared splitter object gives every candidate the same folds, so the comparison is paired and the differences come from the models rather than from different partitions. Passing an integer repeatedly with shuffling enabled elsewhere, or varying the seed, adds partition noise that can easily exceed the true gap between two models.
- Does cv=5 stratify for a multilabel classification target?No. Stratification applies only when the target type is binary or multiclass; a multilabel indicator matrix falls through to plain `KFold`. Label balance across folds is then left to chance, which matters when some labels are rare. If you need it, supply an explicit splitter or precomputed index pairs.
- How do you cross-validate with a single, fixed validation set you already chose?Pass an iterable of index pairs: `cv=[(train_idx, valid_idx)]`. Every cross-validating helper accepts that, so `GridSearchCV` will search against exactly your split. `PredefinedSplit`, driven by a `test_fold` array, does the same thing when you want the assignment expressed per row.
saying these in an interview costs you the question
- Assumes cv=5 shuffles the data
- Thinks stratification applies to regressors too
- Passes a float like cv=0.2 expecting a hold-out fraction
- Reports only the mean fold score, never the spread
- Believes stratified folds also keep related rows together