skip to content

Five-fold CV returns accuracies 0.73, 0.78, 0.81, 0.84 and 0.89 - what do you report?

level: juniorimportance: must knowfreq 62%

answer

  1. one number hides four others
  2. mean plus a spread, not the max
  3. divide the spread by root k
  4. 0.06 over sqrt(5) is about 0.03

basics

~20 s

Report the mean, 0.81, together with the spread: the fold standard deviation is about 0.06, giving a standard error of roughly 0.03. Write it as 0.81 plus or minus 0.03 over five folds, never as the best fold's 0.89.

solid answer

~40 s

The headline is the mean across folds, 0.81 — not the 0.89 the luckiest fold produced. But a bare 0.81 hides that the folds ranged from 0.73 to 0.89. The sample standard deviation of those five numbers is about 0.06, so the standard error of the mean is `0.06 / sqrt(5)`, about 0.027. I would write `0.81 +/- 0.03 across 5 folds (fold range 0.73-0.89)`, so a reader can see the number is good to roughly the first decimal place and not to three. I would also flag that `s / sqrt(k)` is an optimistic uncertainty here: the five training sets overlap heavily, so the fold scores are positively correlated and the true spread of the CV estimate is wider than that formula suggests.

go deeper

for a junior

Know to average the folds and to attach the spread, and be able to divide a standard deviation by the square root of the fold count on the spot without a calculator.

for a middle

Explain what the fold spread mixes together: sampling noise from small held-out folds versus genuine model instability across different training subsets, and how fold size tells the two apart.

for a senior

Show that you read a wide spread as a diagnostic — check fold sizes, check for grouped or time-ordered data, repeat the split — before you either report it or act on it.

for a principal

Set the reporting convention for the team: metric, protocol, mean, spread, fold size and search size in every result, so numbers from different people can be compared without argument.

## The arithmetic first Five fold accuracies: 0.73, 0.78, 0.81, 0.84, 0.89. - **Mean** = (0.73 + 0.78 + 0.81 + 0.84 + 0.89) / 5 = 4.05 / 5 = **0.81**. This is the cross-validation estimate. - **Sample standard deviation** across folds: deviations are -0.08, -0.03, 0.00, +0.03, +0.08; squares sum to 0.0146; divide by k-1 = 4 to get 0.00365; square root gives **s ≈ 0.060**. - **Standard error of the mean** = `s / sqrt(k)` = `0.060 / sqrt(5)` ≈ **0.027**, so about 0.03. So the reportable line is something like `0.81 +/- 0.03 (5-fold CV, fold range 0.73-0.89)`. ## Why not just "0.81"? Because a single number invites false precision. A reader who sees `0.81` will happily compare it to a `0.83` from another notebook and conclude the other model is better — when the uncertainty on each is around 0.03 and the two are indistinguishable. Reporting the spread converts a comparison into an honest one: with `0.81 +/- 0.03`, that `0.83` obviously falls inside the noise. ## Why not the best fold, 0.89? Because the best fold is a maximum over five noisy estimates, and a maximum is biased upward. Quoting 0.89 is not "showing what the model can do"; it is reporting the friendliest of five draws as if it were typical. The same reasoning rules out quietly dropping the 0.73 fold as an outlier: unless you have a concrete reason (a fold that was empty of the positive class, a leakage bug in one split), the low fold is data about your model's variability, not a defect in the measurement. ## Two different things the spread is telling you Fold-to-fold variation mixes two sources, and interviewers like candidates who separate them: 1. **Test-set sampling noise.** Each fold's score is measured on only n/k examples. With 100 examples in a fold and a true accuracy near 0.81, binomial noise alone has a standard deviation of `sqrt(0.81 * 0.19 / 100)` ≈ 0.039 — larger than the 0.06 observed spread would suggest is anything but noise. With 10,000 examples per fold, the same figure drops to about 0.004, and a 0.06 spread would then be a real alarm. 2. **Training-set instability.** Each fold trains on a different 80% of the data. A high-variance learner (a deep unpruned tree, a small-sample fit) genuinely produces a different model per fold, and that shows up as spread too. So the honest reading of a wide spread is: *check the fold sizes first.* Small folds explain most wide spreads on small data. If the folds are large and the spread is still wide, you have an unstable learner or a heterogeneous dataset (groups, time periods, sites) whose slices really do behave differently — and that is worth investigating before you report anything. ## Why `s / sqrt(k)` is an optimistic uncertainty The formula `s / sqrt(k)` assumes the k fold scores are independent draws. They are not. Any two of the five training sets share three quarters of their rows, so the five fitted models are highly correlated, and so are their errors. Positive correlation between the terms you are averaging means the variance of the average is larger than the independent-case formula gives. It is a known result that no unbiased estimator of the variance of k-fold cross-validation exists from a single run of it — the fold scores simply do not contain the information. Practical consequence: treat `+/- 0.03` as a *floor* on the uncertainty and a rough scale indicator, not as a confidence interval you would defend statistically. If you need a defensible interval, repeat the cross-validation over several independent fold splits and look at the spread of the repeated means, or bootstrap at the level of examples. ## What a good report line contains - the metric and the protocol (`accuracy, stratified 5-fold`), - the mean and a spread (`0.81 +/- 0.03`, or the fold range), - the fold count and fold size, so the reader can judge the noise, - and, if the model was chosen from a search, how many candidates were tried. That is four short facts, and it turns a number that invites over-reading into a number a reader can act on.

  • Why is 0.06 divided by the square root of five an optimistic measure of uncertainty here?
    Because that formula assumes the five fold scores are independent. Each pair of training sets shares three quarters of the rows, so the fitted models and their errors are positively correlated, which widens the true variance of the average. There is in fact no unbiased estimator of a k-fold estimate's variance from one run, so treat the figure as a scale, not a defensible confidence interval.
  • The folds range from 0.73 to 0.89 - does that mean the model is unstable?
    Not on its own; check the fold size first. With a few hundred examples per fold, ordinary test-set sampling noise easily produces that range even from a perfectly stable model. If each fold holds tens of thousands of examples and the spread is still 0.06, then you have real instability or genuinely heterogeneous data slices, and that is worth digging into.
  • Would you ever quote the 0.89 fold?
    Only as part of the reported range, never as the model's score. It is the maximum of five noisy estimates and is biased upward by construction. The same discipline applies to the 0.73: report both as the range, and let the mean carry the headline.

saying these in an interview costs you the question

  • Quotes the best fold's score as the model's accuracy
  • Reports the CV mean as exact, with no spread at all
  • Confuses the fold standard deviation with the standard error of the mean
  • Drops the weakest fold as an outlier without a concrete reason
  • Treats s over root k as a rigorous confidence interval

context