What does stratified k-fold cross-validation preserve, and when does plain k-fold fail?
answer
- class proportions, not row counts
- think about the rarest class
- a fold left with three positives
- deal each class round-robin into folds
basics
~20 sStratified k-fold gives every fold roughly the same class proportions as the whole dataset. Plain k-fold fails when a class is rare: some folds get very few or zero minority rows, so fold scores swing or go undefined.
solid answer
~40 sPlain k-fold shuffles rows and cuts them into k equal parts, so the class balance of each fold is left to chance. Stratified k-fold instead splits each class separately and deals its rows across the folds, so every fold carries approximately the population class proportions. Take a surface-defect label at 2 percent on 5,000 rows: 100 defects, 10 folds. Stratified folds hold 10 defects each; a random split can easily leave one fold with three and another with seventeen, and recall or precision computed on three positives is nearly noise. The consequence is variance in the cross-validated estimate, not bias in the model. Stratification is essentially free and is the sensible default for classification. What it does not do is fix the imbalance itself — every fold still contains 2 percent positives.
code
python · 20 linesimport random
random.seed(0)
labels = [1] * 100 + [0] * 4900 # 2% positive
rows = list(range(len(labels)))
random.shuffle(rows)
def deal(rows, labels, k):
folds = [[] for _ in range(k)]
for cls in (0, 1): # split each class on its own
members = [r for r in rows if labels[r] == cls]
for i, r in enumerate(members):
folds[i % k].append(r) # round-robin keeps the proportion
return folds
plain = [rows[i::10] for i in range(10)]
strat = deal(rows, labels, 10)
print([sum(labels[r] for r in f) for f in plain])
print([sum(labels[r] for r in f) for f in strat])go deeper
Be ready to say in one line what is held constant — the class proportions — and to name the failure it prevents: a fold with almost no minority rows. Knowing it is the default for classification is enough at this level.
Explain the mechanics: each class is split separately and dealt across folds, so the effect is on the variance of the estimate rather than on bias. Expect to be pushed on the case where a class has fewer rows than folds.
Show that you check per-fold class counts before trusting a cross-validated comparison, and that you can say when the fold-to-fold spread reflects model instability versus split noise. Be clear that stratification is not an imbalance remedy.
Own the policy question: which splitting defaults the team encodes, when small rare classes should be merged or excluded from reported per-class metrics, and how much of a score difference between two models is real given the fold noise you actually have.
## The problem stratification solves K-fold cross-validation divides the rows into k parts, trains on k-1 of them and scores on the one left out, k times. Plain k-fold decides which row goes where by shuffling and cutting. Nothing in that procedure looks at the label, so the class composition of each fold is a random draw. When the classes are balanced and the dataset is large, that randomness is harmless: with 50,000 rows at 50/50, every fold lands within a fraction of a percent of the population rate. The trouble starts when one class is rare, when the dataset is small, or both. Concretely: 5,000 manufacturing inspection rows with a 2 percent surface-defect rate — 100 defects — split into 10 folds of 500. Under a random split the number of defects in a fold behaves like a binomial draw: it averages 10 but has a standard deviation of about 3. Folds with 4 or 5 defects, and folds with 15 or 16, are ordinary outcomes rather than freak ones. Recall measured on a validation fold containing four positives can only take the values 0, 0.25, 0.5, 0.75 or 1. The spread you then see across the ten fold scores is mostly an artefact of how the defects happened to fall, not a property of the model. ## What stratified k-fold does Stratified k-fold treats each class as its own pool. It splits the positives into k parts and the negatives into k parts, then combines one part of each into every fold. Every fold therefore holds approximately the population proportion of every class — exactly, when the class count divides evenly by k, and off by at most one row otherwise. In the example each fold receives exactly 10 defects. The effect is on the *variance* of the cross-validated estimate. The average score across folds is not systematically moved up or down; the fold-to-fold scatter shrinks, so the mean is a more precise estimate and the standard deviation across folds becomes interpretable as model instability rather than split noise. That precision matters most in exactly the setting where cross-validation is used for choosing between models: two candidates that differ by one point are indistinguishable when the split alone contributes three points of noise. ## What it is not Stratification is not a remedy for class imbalance. After stratifying, the training portion of each fold still has a 2 percent positive rate and the model still faces the same learning problem. How you handle the imbalance during training and where you set the decision threshold are separate decisions from how you cut the folds. It is also not a guarantee that fold scores will agree. Two folds with identical class counts can still score differently because they contain different, genuinely harder rows. Stratification removes one source of disagreement, not all of them. And it stratifies on the *label*, not on a feature. Balancing folds on some input column is a different, usually unmotivated, operation — the reason to stratify is that the metric is computed against the label, so a fold with an unrepresentative label mix produces an unrepresentative metric. ## The hard limit: k cannot exceed the rarest class count Stratification can only put a row of class C in a fold if such rows exist. With a 40-class product-category label where the rarest category has six rows, no 10-fold scheme can give every fold a row of that class — at least four folds will contain none, and any per-class metric for that category is undefined on those folds. The available moves are to lower k so that k is at most the smallest class count, to merge the very rare categories into an 'other' bucket before splitting, or to accept that per-class numbers for those categories carry no information and report only where support exists. Silently keeping k=10 and averaging over undefined fold values is the failure mode to avoid. ## When it barely matters With hundreds of thousands of rows and a near-balanced label, random and stratified folds are practically identical. Because stratification costs nothing, the usual advice is still to use it by default for classification — the gain grows as the data shrinks and the minority class thins out, and there is no case where it hurts.
- What do you do when a class has fewer rows than there are folds?You cannot stratify it into every fold — with six rows and k=10, at least four folds hold none of that class. Either lower k to at most the smallest class count, merge the very rare categories into an 'other' bucket before splitting, or keep k and stop reporting per-class metrics for categories with no support in a fold. What you must not do is average a metric that is undefined on several folds.
- Does stratifying the folds fix class imbalance?No. Every training portion still carries the same 2 percent positive rate, so the learning problem is unchanged. Stratification only stabilises the estimate by stopping folds from having wildly different class mixes. How the model is trained under imbalance and where the decision threshold sits are separate choices made after the folds exist.
- Is stratification worth it on a large, nearly balanced dataset?The gain is negligible — with 200,000 rows at 45/55, random folds already match the population within a rounding error. It costs essentially nothing though, and it never hurts, so keeping it as the default for classification means you do not have to re-decide when the next dataset is small or skewed.
Dealing a deck so each player gets the same mix of suits, rather than cutting the shuffled deck into piles and hoping.
saying these in an interview costs you the question
- Says stratified folds fix imbalance so no other handling is needed
- Stratifies on an input feature rather than on the label
- Keeps ten folds when the rarest class has six rows
- Believes fold scores should be identical once folds are stratified
- Thinks stratification makes the estimate unbiased rather than less variable