Your 5-fold CV accuracy is 0.87 — what does that promise about the refit model you ship?
answer
- an average over models you deleted
- estimates a procedure, not an artifact
- trained on 80%, shipped on 100%
- conservative for size, optimistic for selection
- silent about a shifted population
basics
~20 sVery little about that specific model. 0.87 is the average score of five other models, each trained on 80% of the rows and tested on the fold it had not seen. It estimates the recipe, not the artifact.
solid answer
~50 sThe 0.87 is an average over five models that no longer exist, each trained on 80% of the training rows and scored on the 20% it was held out from. It estimates the expected accuracy of the *procedure* — this pipeline, these hyperparameters, roughly this much data, this population. The model that actually ships was trained on 100% of the rows, so it was never measured at all; because it saw more data, 0.87 is usually a mildly conservative estimate for it. Two things pull the other way: if the same folds also chose the winning setting, the reported mean carries selection optimism, and the number says nothing about a deployment population that differs from the training sample. So I quote it as an estimate of the recipe, and confirm the artifact itself on data that played no part in building it before promising a number to anyone.
go deeper
Know that a cross-validation score is an average across folds, computed from several temporary models, and is an estimate with uncertainty rather than a measurement of the model you deploy.
Explain the mechanics out loud: k models, each trained on (k-1)/k of the rows and scored on its held-out fold, then averaged — and say what training size and what population that average describes.
Name the two biases pointing in opposite directions, less data per fold and selection over many candidates, and state what you would measure on untouched data before committing to a number in production.
Own how model performance is communicated outward: which number goes into a business case or a contract, which caveats travel with it, and what evidence must exist before the organisation treats it as real.
## What the number is an average of A 5-fold cross-validated accuracy of 0.87 is produced like this: the training data is cut into five folds; five models are trained, each on four folds; each is scored on the fold it never saw; the five held-out accuracies are averaged. So 0.87 is the mean of five measurements, each taken on a different fifth of the data, by five different models — none of which is the model you deploy. That single sentence contains almost everything an interviewer is probing for. The estimate attaches to the **procedure**: this feature pipeline, this learner, this hyperparameter setting, trained on about 80% of your rows, evaluated on data drawn the same way as your training sample. It does not attach to the artifact in your registry. ## Bias one: training-set size makes it conservative Each fold model trained on `(k-1)/k` of the data. The shipped model trained on all of it. Wherever the learning curve is still rising — small samples, flexible learners, high-cardinality features — the refit model is genuinely a bit better than the average fold model. On that axis alone the CV mean is **pessimistically biased** for the model you ship, and the bias grows as k shrinks: at k=2 each fold model sees half the data, at k=10 it sees 90%, and leave-one-out is nearly unbiased for training-set size (while paying for it in variance and compute). ## Bias two: selection makes it optimistic If the same folds that produced 0.87 also chose the winning setting out of many candidates, the reported number is the **maximum** of a set of noisy estimates, and maxima of noisy estimates are biased upward. The more candidates you compared and the noisier each fold mean, the larger that optimism. The two biases point in opposite directions and there is no general rule for which dominates — it depends on k, on `n`, and on how wide your search was. The honest posture is to name both rather than to claim the estimate is neutral. ## What it assumes and therefore does not cover Even setting bias aside, the estimate is conditional on the evaluation data resembling what the model will actually meet. Standard k-fold shuffles rows, which assumes exchangeability; if your data has time order, repeated customers, sites, or any grouping, random folds leak structure across the split and 0.87 becomes an over-estimate of anything real. Covariate shift, a changed label definition, a new acquisition channel, or a seasonal effect all break the assumption too. Cross-validation measures generalisation to *more of the same data*, never to a different world. ## It is also an estimate with uncertainty 0.87 is a mean over five numbers. Report it with some sense of how much those five varied; a mean of 0.87 assembled from 0.86, 0.87, 0.87, 0.88, 0.87 is a very different object from one assembled from 0.78, 0.83, 0.87, 0.92, 0.95. The second says the recipe's performance depends heavily on which rows it gets, which is exactly the situation where the shipped model's true accuracy could sit well away from 0.87. ## Accuracy specifically One further caution when the metric is accuracy: it is only interpretable next to a base rate. At a 87% majority class, 0.87 is the score of a constant predictor. Whether 0.87 is impressive, adequate or embarrassing is a question about class balance and the cost of each error type, and cross-validation says nothing about that. ## So what can you honestly say? - "Running this pipeline with these settings on roughly this much data from this population produced a held-out accuracy averaging 0.87 across five folds, with a spread of X." - "The deployed model was trained on 25% more data than any fold model, so I expect it to be at least as good on the same population." - Not: "the model is 87% accurate", and never a contractual number. ## How to get a number about the artifact If someone needs a claim about the shipped model rather than the recipe, measure it on data that took no part in training or selection: a test set held out before any tuning and touched exactly once, a later time window, or a shadow deployment compared against production labels. That measures the artifact. It costs data and elapsed time, which is why the CV mean remains the number you tune on and the held-out measurement remains the number you promise.
- Is the cross-validation estimate biased upward or downward for the refit model?Both pulls are present. Each fold model trains on (k-1)/k of the rows, so it is weaker than the refit and the mean is pessimistic on that axis — more so for small k. Against that, if the winning setting was chosen by comparing many candidates on those same folds, the winning mean is the max of noisy estimates and is optimistic. Which dominates depends on k and on how wide the search was.
- How would you obtain an honest number for the shipped model itself?Score it on data that played no part in training or selection: a test set held out before tuning and used once, a later time period, or a shadow deployment measured against production labels. That evaluates the artifact rather than the recipe. It costs data and elapsed time, which is why the cross-validation mean stays the number you tune on.
- Does a large spread across the five fold scores change how you use the 0.87?Yes — it tells you the recipe's performance depends heavily on which rows it gets, so the shipped model's true accuracy could sit well away from the mean. Treat a wide spread as a signal to investigate: too little data, an unstable learner, or folds that are not exchangeable because of time order or grouped records.
saying these in an interview costs you the question
- Quotes the CV mean as a production guarantee or SLA
- Says the refit model will score exactly 0.87
- Thinks cross-validation measured the model that ships
- Ignores that the same folds also chose the setting
- Reads 0.87 accuracy without asking the base rate