What do you do when 5-fold cross-validation scores swing by 0.08 between random partitions?
answer
- the number moved but the model did not
- you measured your own noise floor
- run the whole thing again, differently
- averaging repeats, not adding data
- five replications of two folds
basics
~20 sTreat that swing as a measurement of your own noise floor. Repeat the whole cross-validation with several different random partitions and average across all runs, and refuse to call any model difference smaller than the swing you just observed.
solid answer
~40 sA single 5-fold run gives one number, but it is one draw from a distribution over partitions, and on a 600-row table that distribution can be 0.08 wide. Repeated k-fold is the direct fix: run the whole k-fold procedure r times with a different random partition each time and average the `r*k` fold scores. The classic small-sample variant is 5x2 — five replications of a 2-fold split, designed for comparing two procedures on limited data. Two honest limits. Repeats do not remove the bias from training on only four fifths of the rows; that is set by k. And the `r*k` scores are not independent — they reuse the same 600 rows — so repeats sharpen your reading of this dataset, not of a fresh sample. Never pick the best repeat.
go deeper
Be ready to recognise that a score changing between runs usually means the fold partition changed, not the model, and that averaging several runs is the standard response.
Explain repeated k-fold concretely — r fresh partitions, r times k fold scores averaged — and state what it does not fix, namely the bias from each model training on a fraction of the rows.
Treat the swing as your resolution limit and act on it: refuse comparisons finer than the noise, weigh raising k or stabilising the procedure against the cost of r times more fits, and never report a hand-picked repeat.
Own how the organisation reports model results. Decide whether headline numbers must come with a measured spread, and push back when a decision hinges on a difference the evaluation cannot resolve.
## What the swing is telling you When you re-run 5-fold cross-validation with a different random partition and the reported score moves by 0.08, nothing about the model or the data has changed. The only thing that changed is which rows landed in which fold. So the 0.08 is a property of your measuring instrument, and the honest reading is: on this dataset, at this k, a single cross-validation run resolves differences of about 0.08 and no finer. That number is immediately useful. If model A scored 0.84 and model B scored 0.81 in single runs, the gap is well inside the noise and the comparison is not a result. Many wasted afternoons come from tuning against differences smaller than the instrument's resolution. The swing is large when the dataset is small (600 rows split five ways means folds of 120), when the procedure is unstable — deep unpruned trees, high-variance fits, aggressive feature selection — or when a handful of unusual rows dominate the metric, so that whichever fold they land in gets dragged. ## Repeated k-fold Repeated k-fold runs the entire k-fold procedure `r` times, drawing a fresh random partition each time, and averages over all `r*k` fold scores. Five repeats of 5-fold is 25 fits and 25 fold scores. What this buys: the average over repeats no longer depends on one lucky or unlucky partition, so the reported number is reproducible in the sense that another analyst repeating the exercise lands close to yours. The spread across repeats is itself worth reporting as a description of instrument resolution. What this does not buy: - **It does not reduce the bias set by k.** Every repeat still trains on `n*(k-1)/k` rows, so if 5-fold is pessimistic because 480 rows is not 600, ten repeats of 5-fold are pessimistic by the same amount. Bias is a function of k; partition noise is what repeats attack. - **It does not manufacture data.** All repeats reuse the same 600 rows. Averaging tells you more precisely what *this* sample says, not what a fresh 600-row sample would say. If those 600 rows are unrepresentative, every repeat inherits it. - **It does not give independent measurements.** The `r*k` scores share rows extensively — within a repeat the training sets overlap by 75%, and across repeats every score is computed on the same table. Treating them as independent observations and reasoning as if you had 25 clean samples overstates your confidence. ## The 5x2 scheme A specific variant worth knowing by name is 5x2 cross-validation: five replications of a 2-fold split, giving ten fold scores. It was proposed for the specific job of comparing two learning procedures on a limited dataset, where the usual worry is that overlapping training sets make repeated-fold comparisons look more decisive than they are. Two folds means each pair of training sets within a replication is disjoint, which is what makes the scheme attractive for a comparison; the cost is that every model trains on only half the rows, so the absolute scores are pessimistic. Use it to answer "is A better than B", not to publish A's performance. ## Other levers before you reach for repeats Repeats multiply compute by `r`, so consider the alternatives first. Raising k from 5 to 10 gives each model more training rows and often calms the swing on a small table, at twice the fits. Simplifying or regularising the procedure attacks instability at its source: if deep trees swing and shallow ones do not, the swing was telling you something about the model, not only about the split. Fixing a partition seed makes results *reproducible* but does nothing about accuracy of the estimate — it just hides the swing behind a constant, which is worse than seeing it. ## Aggregating what you get How the fold scores become one number matters. For a decomposable metric such as accuracy or mean squared error, the mean over equal-sized folds equals what you would get by pooling all out-of-fold predictions and scoring once; with unequal fold sizes you should weight each fold by its size to get the same answer. Metrics built from a ranking over the whole set are not decomposable: the mean of the per-fold values is not the value computed on the pooled out-of-fold predictions, and the two can differ noticeably on small folds. Pooling gives one stable number but assumes the models' output scores are comparable across folds, and it discards the fold-to-fold spread that told you about resolution in the first place. Report both when the audience will act on the number. ## What a weak answer looks like Re-running until a good partition appears and reporting that one. Presenting a two-decimal score from a single run on 600 rows as if it were precise. Concluding the model is broken when the split was simply noisy. All three are variations of confusing the instrument with the thing being measured.
- How should the fold scores be aggregated into one reported number?For decomposable metrics like accuracy or mean squared error, average the folds, weighting by fold size when the folds are unequal — that matches scoring all out-of-fold predictions at once. Ranking-based metrics are not decomposable, so the mean of per-fold values differs from the value on pooled out-of-fold predictions; pooling is steadier but throws away the fold-to-fold spread.
- Does repeating cross-validation remove the pessimism from training on only four fifths of the rows?No. That pessimism is set by k, because every repeat still trains on the same fraction of rows. Repeats average away the luck of a particular partition; raising k is what shrinks the training-size bias. They are different problems with different fixes.
- Why not just fix the random seed and stop worrying about the swing?A fixed seed makes the number reproducible, not accurate. The swing is still there; you have simply chosen one draw from it and hidden the rest. Worse, you may now compare two models under one arbitrary partition and read a difference that is pure partition luck.
saying these in an interview costs you the question
- Reruns until a good partition appears and reports that one
- Treats the r times k fold scores as independent samples
- Believes repeats add information the 600 rows do not contain
- Blames the model when the partition was the variable that changed
- Compares two models on a gap smaller than the observed swing