After k-fold cross-validation picks a hyperparameter setting, which model do you actually deploy?
answer
- the folds were a measuring device
- each fold model saw part of the rows
- you selected a recipe, not a fit
- run the whole recipe once on everything
basics
~20 sRefit a single model on the entire training set with the chosen setting, and deploy that one. The k fold models were only instruments for scoring settings; each was trained on a fraction of the rows and is normally discarded.
solid answer
~50 sCross-validation evaluates a *procedure*, not a model. For each candidate setting it trains k models on k different `(k-1)/k` slices and averages their held-out scores, so that average tells you how good the recipe is — but every one of those models saw only part of the data. Once a setting wins, you discard the fold models and run the recipe once more on 100% of the training rows: same hyperparameters, and every data-dependent step (imputation, scaling, encoding, feature selection) re-estimated from scratch on the full set. That single refit model is the deliverable. More training data almost always helps a little, so the shipped model should be at least as good as the fold models were. The one thing to check before copying settings across is whether any of them depends on how many rows you trained on.
go deeper
Remember the sequence: cross-validation scores candidate settings, then one final model is trained on all the training data using the winning setting. Do not confuse the fold models with the deliverable.
Be ready to explain that each fold model trains on only (k-1)/k of the rows, and that the entire pipeline — imputation, scaling, encoding, then the estimator — is re-fitted on the full training set.
Show that you treat the refit as a controlled rerun: same pipeline, seed pinned, artifact and data snapshot versioned, and a sanity check that the refit model's predictions track the fold models' on shared rows.
Own the promotion rule your team follows: which artifact is the deliverable, what evidence ships with it, and under what narrow conditions a cheaper alternative to a full refit is acceptable.
## What cross-validation actually hands you Split the training data into k roughly equal folds. For one candidate hyperparameter setting, train k separate models: model *i* is fitted on the `k-1` folds that exclude fold *i*, then scored on fold *i*, which it never saw. Average those k held-out scores. Repeat for every candidate setting and keep the setting with the best average. Look carefully at what you now own. You own **one number per setting**, and you own `k` trained models per setting that existed only to produce those numbers. Nothing in that procedure has produced a model trained on all of your data. Cross-validation is a measuring instrument for recipes; it is not a training routine that returns a deliverable. ## Why the answer is "refit on everything" Two reasons, and they stack. **Each fold model is handicapped.** It trained on `(k-1)/k` of the rows — 80% at k=5, 90% at k=10. Learning curves for almost every learner are monotone in the useful direction: more training data lowers variance and, for flexible models, lowers achievable error. Shipping a model deliberately trained on 80% of what you have is throwing away data for no reason. **Picking one fold model is arbitrary, and picking the *best* fold model is worse than arbitrary.** The k held-out scores differ mostly because the folds differ. Selecting the highest of k noisy scores is selecting on noise; the winner's advantage is largely the luck of an easy fold, measured on only `n/k` rows. You would be choosing your production artifact by a coin flip dressed up as evidence. So: throw the fold models away, re-run the recipe once on the full training set with the winning setting, and ship that. ## What "refit" has to include The unit that was cross-validated is the whole pipeline, not just the estimator at the end. Everything fitted from data must be re-estimated on the full training set, in the same order: - imputation values (means, medians, category fallbacks), - scaling or normalisation statistics, - category encodings, including which rare levels collapse into an "other" bucket, - any data-driven feature selection or dimensionality reduction, - any target-derived transform such as a class-weight or log-odds mapping. Carrying over statistics computed during a fold is a double mistake: it wastes the extra rows, and it pairs a model trained on one sample with transforms estimated from a different, smaller one. It also quietly breaks the correspondence between the thing you measured and the thing you ship — the fold scores were produced by a pipeline whose transforms were fitted inside the fold, so a refit that skips that step is not the same recipe. ## What stays fixed, and the one caveat The hyperparameters are held at exactly the values cross-validation selected. Changing them during the refit — "the model looked a bit underfit so I raised the depth" — silently discards the whole selection exercise and leaves you with a setting nothing measured. The caveat is that a few settings encode an assumption about how much data there is: the number of boosting rounds, the neighbour count in a nearest-neighbour model, penalty strength, and any constraint written as an absolute row count. Those were tuned at `(k-1)/k` of the final size, so they may need rescaling rather than copying. That is a real refinement of the rule, not an exception to it. ## The cost is negligible Selecting over `C` candidate settings with k folds already cost `C * k` fits. The refit adds exactly one more. If the refit is the fit you cannot afford, the tuning run was already unaffordable, and the honest response is a cheaper search — not shipping a fold model. ## Operational hygiene around the refit Because this single artifact is now the product, treat the refit as a controlled rerun rather than an afterthought: - pin and record the random seed, so the artifact is rebuildable; - record the winning setting, the data snapshot, and the code version alongside the model; - run a smoke comparison — predictions from the refit model against predictions from a fold model on a sample of rows. They should correlate strongly. A large systematic divergence usually means a preprocessing step behaved differently on the full set (a category that no longer collapses, a scaler with a different range) rather than a genuine improvement. ## What the CV number now means The average you selected on describes the procedure at fold-sized training sets. It is a reasonable and slightly conservative estimate for the refit model, not a measurement of it. Quote it as an estimate of the recipe, and keep whatever untouched data you have for confirming the artifact itself.
- During the final refit, do you re-estimate preprocessing such as scaling, imputation and encoding?Yes. Every data-dependent transform is part of the recipe and must be re-fitted on the full training set, in the same sequence, before the model is trained. Reusing means, medians or category maps computed inside a fold both wastes the extra rows and pairs the model with statistics from a different, smaller sample than the one it was trained on.
- How do you keep the single refit model reproducible when training is nondeterministic?Pin and record the seed, the exact hyperparameter setting, the data snapshot and the code version, so the artifact can be rebuilt. If repeated refits with different seeds disagree materially, that instability is itself a finding: fix it with more regularization, more rounds or a more stable learner, or report it — do not quietly ship whichever seed looked best.
Cross-validation is a taste test on small batches: you cook five test portions to decide the recipe, then cook the real dish once, at full size, using the recipe that won. You do not serve the tasting spoons.
saying these in an interview costs you the question
- Deploys the best-scoring fold model as the final model
- Believes cross-validation returns a trained model to ship
- Reuses scaler or imputer statistics fitted inside the folds
- Changes hyperparameters during the refit by feel
- Assumes the refit model inherits the CV score exactly