skip to content

In k-fold cross-validation, how does the choice of k trade off bias against variance?

level: middleimportance: must knowfreq 78%

answer

  1. how many rows each model gets to see
  2. the estimate is of a smaller model
  3. training folds overlap more as k grows
  4. small k reads pessimistic
  5. convention is 5 or 10, and k fits

basics

~20 s

Larger k trains each fold's model on more rows, so the error estimate is less pessimistic, but the training sets overlap more and each held-out fold is smaller, so the estimate is noisier. k=5 or 10 balances both.

solid answer

~40 s

k-fold cross-validation splits the rows into k parts, trains k times on k-1 of them and scores on the held-out part, then averages. What it estimates is the error of a model trained on `n*(k-1)/k` rows, not on all `n`. With small k that training set is much smaller than the full data, so if the learning curve has not flattened the estimate comes out pessimistic — that is the bias. Raising k shrinks that gap, but the k training sets now overlap in almost every row, so their errors are strongly correlated and averaging them cancels less noise; the estimate also becomes more sensitive to individual rows. k=5 and k=10 are the conventional compromise: training sets at 80% or 90% of full size, and only 5 or 10 fits to pay for.

go deeper

for a junior

Be ready to state the mechanics: split into k parts, train k times on all but one part, score on the held-out part, average the k scores. Know that k=5 and k=10 are the usual choices.

for a middle

Explain what the estimate is actually about — a model trained on a fraction of the rows — and why small k therefore reads pessimistic while large k costs more fits and yields near-identical, correlated training sets.

for a senior

Show you pick k from the situation: where the learning curve sits, how long one fit takes, and how many rows land in each held-out fold. Say out loud what your chosen k costs in compute.

for a principal

Own the policy question. Decide what the team's default k is, when a cheaper scheme is good enough for a screening sweep versus a headline number, and how much compute the organisation should spend to shrink an estimate nobody will act on.

## What k-fold cross-validation is k-fold cross-validation is a way to estimate how well a modelling procedure will do on rows it has not seen, using only the data you have. The `n` rows are partitioned into k roughly equal, non-overlapping folds. For each fold in turn, that fold is held out, a model is trained from scratch on the other k-1 folds, and the model is scored on the held-out fold. Each row is therefore predicted exactly once, by a model that never saw it. The k fold scores are combined into one number, usually by averaging. Two details matter before the tradeoff makes sense. First, the k models are thrown away — cross-validation evaluates a *procedure* (this algorithm with these settings on data of about this size), not a particular fitted model. The model you ship is normally refit on all `n` rows afterwards, which is a `k+1`-th fit. Second, the whole exercise costs k full training runs, so k is a compute decision as much as a statistical one. ## The bias side: how much data each model sees Each of the k models is trained on `n*(k-1)/k` rows. With k=2 that is half the data; with k=10 it is 90%; with k=n (leave-one-out) it is all but one row. Learning curves almost always slope downward: a procedure trained on fewer rows performs worse. So cross-validation with small k systematically reports a *worse* score than the model you will actually ship, which is trained on all `n`. This is the pessimistic bias of the estimate, and it shrinks as k grows. Its size depends entirely on where you sit on the learning curve. On 5,000,000 rows with a simple model the curve is flat and 2-fold is as unbiased as 10-fold in practice. On 300 rows with a flexible model the curve is steep, and the difference between training on 240 rows (k=5) and 270 rows (k=10) is visible in the number. Note the direction carefully: the bias makes the estimate pessimistic, not optimistic. Cross-validation's failure mode of being *too kind* comes from leakage, tuning against the folds, or a split that ignores structure in the rows — not from the choice of k. ## The variance side: how much the estimate moves Variance here means: if you re-ran the whole thing — a different random partition, or a fresh sample of `n` rows from the same population — how much would the reported number jump? The standard account is that variance rises with k, for one main reason. The k training sets overlap: with k=10, any two of them share about 89% of their rows, so the ten models are near-copies of each other and their errors move together. Averaging quantities that are strongly positively correlated removes much less noise than averaging independent ones, so a 10-fold average is barely more stable than a single fold's score would suggest. With k=5 the training sets share 75% of their rows — still correlated, but less so — and each held-out fold contains `n/5` rows rather than `n/10`, so each individual fold score is measured on more test rows. Be honest about the caveat: this "more folds, more variance" rule is a textbook generalisation, not a theorem. For a very stable procedure the effect is small, and empirical studies find the variance ordering can invert. What is not in dispute is that a single k-fold run on a small dataset is a noisy number, and that the noise comes largely from which partition you happened to draw. ## Why 5 and 10 k=5 and k=10 are conventions from empirical work, not mathematical optima. They sit where the training sets are 80-90% of full size (so the pessimistic bias is usually small), the number of fits is affordable, and the estimate is not dominated by any single row. Nothing breaks at k=7; you just have no reason to prefer it. Reasonable departures: - **Very large n, expensive fits.** k=3 or even a single split. When the learning curve is flat, the bias argument for a large k evaporates and you are paying compute for nothing. - **Very small n.** Push k up, all the way to leave-one-out if the fits are cheap, because every extra training row matters. - **Fold size floor.** Whatever k you pick, each fold needs enough rows for the metric to mean something. Ten folds of 12 rows each give scores that jump in steps of 1/12. ## Common confusions Raising k does not make the model better — it changes only the estimate of how good the procedure is. It also does not compensate for a partition that ignores structure in the rows, and it is not a substitute for a test set that stays untouched. And k in k-fold has nothing to do with the k in k-nearest-neighbours or the k in k-means; they are three unrelated uses of the letter.

  • Does moving from k=5 to k=10 make the model you ship more accurate?
    No. Cross-validation only estimates how well the procedure performs; it does not change the fitted model. The model you deploy is refit on all rows regardless of k. What changes is the number you report and how much compute you spent getting it.
  • When is k=2 or k=3 a defensible choice?
    When the learning curve is flat — plenty of rows relative to model complexity — so training on half or two thirds of the data scores the same as training on all of it. That removes the bias argument for large k, and a fit that costs hours makes the compute argument decisive.
  • How many model fits does a k-fold run actually cost end to end?
    k fits to produce the estimate, plus one more if you refit on all rows to get the model you ship, so k+1. Any preprocessing that learns from data — scaling, imputation, feature selection — has to be refit inside each fold too, or the estimate is optimistic.

Judging a runner by timing them on shorter practice laps: fewer, shorter laps under-rate them systematically, while many near-identical laps agree with each other so strongly that averaging tells you little new.

saying these in an interview costs you the question

  • Says a larger k always gives a better estimate
  • Thinks raising k improves the deployed model's accuracy
  • Claims small k makes cross-validation optimistic
  • Treats k=10 as a mathematical requirement rather than a convention
  • Ignores that k-fold costs k full training runs

context