Instead of refitting after cross-validation, should you average the k fold models into one predictor?
answer
- free ensemble, recurring costs
- members share most of their rows
- k times the serving bill
- the average itself was never scored
- one auditable artifact usually wins
basics
~20 sA defensible option, not the default. Averaging the fold models gives a variance-reduced ensemble at no extra training cost, but costs k times the inference work, leaves no single auditable artifact, and was never scored as a unit.
solid answer
~50 sYou already paid for the fold models, so averaging them gives an ensemble for free. Three things argue against the default. The members are heavily correlated — any two share most of their training rows — so the variance reduction is far weaker than genuine bagging. Inference costs k times as much, and you version, monitor and roll back k artifacts instead of one. And cross-validation scored the members individually: the average was never evaluated on data none of its members saw, so it carries no honest estimate. For an insurance pricing score, where a model-risk function wants one documented model whose behaviour can be traced, that settles it. I reach for the fold average only when a full refit is unaffordable, or when training is so unstable that a single fit is a coin flip.
go deeper
Know that the models built during cross-validation can be kept and averaged, but that the normal end of the process is a single model refit on all of the training data.
Explain why averaging k fold models is a weak form of bagging: the members are trained on heavily overlapping rows, so their errors correlate and the variance reduction is small.
Weigh the operational side — k times the inference cost, k artifacts to version and monitor, and no out-of-fold score that measures the ensemble as a unit — and name the conditions that would still justify it.
Treat this as a governance decision as much as a modelling one: settle what the organisation must be able to audit, explain and roll back, and only then whether a multi-model artifact is acceptable at all.
## The proposal Cross-validation has just produced k trained models. Rather than discarding them and paying for one more fit on the full data, keep all k and define the production predictor as their average — mean of predicted probabilities or regression outputs, or a vote for hard labels. It sounds efficient: every row was used for training by some member, you skip a training run, and ensembling is generally a good idea. It is a real technique and it is not wrong. It is simply not the default, and being able to say precisely why is the point of the question. ## What the fold average genuinely buys **No extra training.** The k fits are already sunk cost. When a single fit takes days, that matters. **Some variance reduction.** Averaging m predictors with pairwise error correlation `rho` reduces the variance component of error by roughly `rho + (1-rho)/m` of its single-model value. With independent members that is a big win. **Stability against an unlucky fit.** If training is nondeterministic or the learner is high-variance, one refit is a draw from a distribution and the average is a steadier object. ## Why it is weaker than it sounds **The members are not independent.** With 5-fold, two fold models share three of the four folds each was trained on — about 75% of their training rows. Their errors are strongly correlated, `rho` is close to 1, and the term above collapses toward no benefit. This is bagging with the resampling turned almost off. If variance reduction is what you want, deliberate bagging or a random forest over the full data does it properly. **k times the serving cost.** Latency, memory and CPU all multiply. For a batch score overnight nobody notices; for a synchronous pricing call under a latency budget it can be the whole decision. **k times the operational surface.** Five artifacts to version, ship, monitor, explain and roll back together. Any drift investigation now asks which member moved. Feature-importance and explanation outputs must be aggregated before anyone can read them. **No honest estimate for the object you are shipping.** This is the sharp one. The k fold scores measured the members *individually*, each on rows the other members had trained on. The ensemble as a unit was never evaluated on data none of its members saw, so the cross-validation mean is not an estimate of it — and it is not conservative in a predictable direction either, because the fold scores were computed under a different prediction rule. If you ship the average, you owe it a fresh evaluation: a test set held out before cross-validation began, or an outer loop in which the entire build-and-average step is what gets scored. ## The auditability argument Consider an insurance pricing score. A model-risk or compliance function typically wants one documented artifact: the coefficients or tree structure of record, a reproducible build, a written account of how each input moves the price, and a stable answer to "why was this customer quoted this premium?". An average of five models trained on overlapping subsets satisfies none of that comfortably. Every explanation becomes an explanation of an aggregate, monotonicity or business constraints must be verified five times, and the story for why the artifact contains five sets of parameters is one nobody wants to tell twice a year at a model review. In that environment the single refit model is not merely tidier; it is the thing the organisation is able to govern. This is why the question is a leadership call rather than a modelling one: the modelling difference is often small, and the operating difference is not. ## When the fold average is the right call - **A full refit is genuinely unaffordable** — the fit takes days and the deadline is tomorrow. - **Training is highly unstable** — repeated refits with different seeds differ materially, and the average is a more reliable object than any single draw. - **Serving is cheap and ungoverned** — an internal batch job where k times the compute is free and nobody needs a single artifact of record. - **You will evaluate the ensemble honestly** — you have untouched data to score the average as a unit before it goes anywhere. ## The default and how to state it Refit once on all the training data, ship one artifact, and keep the fold average in your pocket as a specific remedy for a specific problem. If asked in an interview, the strongest answer names the correlation between members, the serving and governance multiplier, and the missing estimate — then gives the narrow conditions that would flip the decision. An answer that says "ensembles are always better" misses that the ensemble here is barely an ensemble, and that the cost is paid every single time the model is called.
- If you do ship the fold average, how do you get an honest estimate for it?Score the ensemble as a single unit on data no member ever touched: a test set held aside before cross-validation began, or an outer loop where the whole build-and-average step is what the outer fold evaluates. What you cannot do is reuse the inner fold scores, because each member was tested on rows its fellow members trained on.
- Does averaging fold models help more for a high-variance or a high-bias learner?High variance. Averaging attacks variance, so deep unpruned trees or fits on small samples benefit most, while a strongly regularized linear model — whose fold copies are nearly identical — gains essentially nothing. That is also the case where a refit is cheapest, so the ensemble buys least exactly where it costs least.
- Is the fold average ever preferable to properly bagging the full training set?Rarely on statistical grounds — deliberate bagging draws more varied members and can use all the data in each draw, so it reduces variance more. The fold average wins only on cost, because the fits already exist. If you are willing to pay for training anyway, bag the full data or refit once; do not treat the fold split as a resampling scheme.
Five apprentices each trained on overlapping copies of the same recipe book will make broadly the same dish, so polling them is not five opinions. It is one opinion, delivered five times, at five times the cost.
saying these in an interview costs you the question
- Calls the fold average strictly better because it is an ensemble
- Assumes the CV mean also estimates the averaged ensemble
- Ignores that fold models share most of their training rows
- Forgets the k-fold inference latency and memory bill
- Treats an unauditable ensemble as fine in a governed model