skip to content

When is leave-one-out cross-validation worth its n model fits instead of 5-fold?

level: seniorimportance: should knowfreq 55%

answer

  1. k pushed all the way to n
  2. one fit per row
  3. training sets differ by one row
  4. nearly unbiased, jumpy estimate
  5. a single-row fold has no curve

basics

~20 s

Only when rows are scarce and each fit is cheap. Leave-one-out trains on all but one row, so it is nearly unbiased and deterministic, but it costs one fit per row and averages near-identical models, leaving the estimate jumpy.

solid answer

~50 s

Leave-one-out is k-fold taken to k=n: each row is held out once and the model is trained on the other `n-1`. Two things follow. It is close to unbiased for the model you ship, because each training set is one row short of the full data — that matters when rows are scarce. And it costs `n` fits. On a 120-patient biomarker cohort where 5-fold would train on 96 rows, spending 120 cheap fits to train on 119 is a good trade. On a 50,000-row trip-duration model at 40 minutes per fit, 10-fold is already a seven-hour decision and leave-one-out is years of compute. Statistically there is a cost too: the `n` training sets differ by one row, so the models are near-copies whose correlated errors barely cancel when averaged. Each test fold is one row, so per-fold ranking metrics are undefined.

go deeper

for a junior

Be ready to define it: hold out one row, train on the rest, repeat for every row, then combine the n results. Know that it means n model fits.

for a middle

Explain both sides of the trade — training sets one row short of full so the bias is tiny, against n fits and n near-identical models whose correlated errors leave the average jumpy.

for a senior

Show the arithmetic in the room. Multiply fit time by row count before recommending it, and know that single-row folds force pooling and rule out per-fold curve metrics.

for a principal

Own the standard. Decide when a team is allowed to spend hours of compute chasing a more precise estimate, and whether the resulting difference in the number could ever change the decision being made.

## What leave-one-out is Leave-one-out cross-validation (LOOCV) is the extreme case of k-fold with k set to `n`, the number of rows. Row 1 is held out, a model is trained on rows 2..n and predicts row 1; row 2 is held out, a model is trained on the rest and predicts row 2; and so on for all `n` rows. You end with one out-of-fold prediction per row and no choices to make: there is exactly one possible partition, so no shuffle, no seed, and no run-to-run variation from the split. ## The case for it **Maximum training data.** Each model sees `n-1` rows. Compared with 5-fold's 80%, that all but eliminates the pessimistic bias caused by training on a fraction of the data. On a 120-patient rare-disease biomarker cohort this is the whole argument: 5-fold trains each model on 96 patients, and a flexible model on 96 high-dimensional rows is meaningfully worse than the same model on 119. When the learning curve is steep and the dataset cannot be grown, you want every row you can get into the training set. **Determinism.** No partition to draw means no partition noise. Two people running LOOCV on the same table get the same number, which removes a whole category of argument on small datasets where 5-fold results wander. **Every row gets tested.** With `n` small, a 5-fold test fold might contain 24 rows and the rare event you care about might miss some folds entirely. LOOCV guarantees each row contributes one out-of-fold prediction. ## The case against it **Compute.** `n` fits. On a 50,000-row taxi trip-duration model at 40 minutes per fit, 10-fold costs about seven hours, which is already a real scheduling decision; LOOCV costs 50,000 fits, roughly a few years of serial compute. There is no cleverness that rescues this — the answer is simply that LOOCV is a small-`n` scheme. **Correlated fits, jumpy estimate.** Any two LOOCV training sets differ by two rows out of `n`. The `n` models are therefore near-identical and their errors move together, so averaging `n` of them cancels far less noise than averaging `n` independent measurements would. The textbook summary is "nearly unbiased but high variance". Treat it as a strong tendency rather than a law — for a very stable procedure the effect is mild, and the ordering against 10-fold is not guaranteed in every study — but the mechanism is real and worth naming. **Sensitivity to single rows.** For some procedures LOOCV behaves badly in a way that is easy to state. Consider a balanced two-class dataset and a classifier that predicts the majority class of its training set. Removing one row always tips the training majority toward the *other* class, so every held-out row is predicted wrong and LOOCV reports 0% accuracy where the true rate is 50%. **Single-row test folds.** A fold of one row can produce accuracy of 0 or 1 and nothing else, and metrics that need a ranking over multiple rows — anything built from a curve over thresholds — cannot be computed per fold at all. With LOOCV you must pool all `n` out-of-fold predictions into one set and compute the metric once over that set, which also means there are no per-fold scores to look at for spread. ## The escape hatch: closed-form LOOCV For a few procedures the `n` fits are not needed, because the leave-one-out error has an algebraic shortcut. Ordinary least squares is the standard example: fit once on all rows, and the leave-one-out residual for row i equals the ordinary residual divided by `1 - h_ii`, where `h_ii` is that row's leverage — its diagonal entry in the hat matrix. Summing the squares of those adjusted residuals gives the LOOCV squared error, known as the PRESS statistic, from a single fit. Ridge-penalised least squares has a comparable identity. When such a shortcut applies, the cost objection disappears and LOOCV becomes attractive at sizes where it would otherwise be unthinkable. Say this if it applies to the procedure being discussed; do not claim it for arbitrary models, because for a boosted tree ensemble or anything trained iteratively there is no such identity. ## How to answer the decision Estimate the two numbers out loud. One: how long does a single fit take, multiplied by `n`. Two: how far apart are `n` and `n*(k-1)/k` on the learning curve — is training on 96 rows versus 119 a real difference, or is it 40,000 versus 45,000 where nothing moves? LOOCV wins only when the second gap is material and the first product is small. In most production settings, neither holds, and 5- or 10-fold is the right default.

  • Why can leave-one-out estimates be unstable when each model trains on almost all the data?
    Because the n training sets differ by one row, so the n models are near-copies and their errors are strongly correlated. Averaging correlated quantities removes little noise, so the average stays close to the noise level of a single fit rather than shrinking toward it as independent averaging would.
  • How do you compute a ranking metric under leave-one-out when each fold holds one row?
    You cannot compute it per fold — a single row has no ranking. Pool all n out-of-fold predicted scores into one set alongside their true labels and compute the metric once over the pooled set. The price is that you get one number with no fold-to-fold spread to inspect.
  • Is there any case where leave-one-out costs about the same as one model fit?
    Yes, when the procedure has a closed-form leave-one-out identity. For ordinary least squares, the leave-one-out residual for a row equals its ordinary residual divided by one minus its leverage, so the full leave-one-out squared error follows from a single fit. No such shortcut exists for iteratively trained models.

Tasting a soup by removing one drop at a time and judging the rest: the sample is as close to the whole pot as it can be, but each tasting is nearly the same soup, so a thousand tastings tell you not much more than a few.

saying these in an interview costs you the question

  • Recommends leave-one-out without asking how long one fit takes
  • Calls leave-one-out simply the most accurate scheme
  • Thinks leave-one-out needs shuffling or a random seed
  • Tries to average a curve-based metric over one-row folds
  • Confuses being nearly unbiased with being reliably stable

context