Should you refit on train plus validation once the hyperparameters are chosen?
answer
- validation rows have done their job
- more rows, same frozen configuration
- do not re-run the search afterwards
- the score must describe the shipped model
- watch procedures needing a held-out block
basics
~20 sUsually yes. Once the configuration is frozen, the validation rows are just labelled data, and training on more rows generally gives a better model. Report the test score for the refit model, and never redo selection after folding validation in.
solid answer
~50 sThe validation rows were spent buying a decision; once that decision is frozen they are ordinary labelled data, and on anything short of a huge dataset the extra 20% of rows is worth real accuracy. So the normal sequence is: select on validation, freeze the configuration, refit that exact configuration on train plus validation, then score the untouched test set once. Two disciplines matter. First, the number you report has to belong to the model you actually shipped - pairing a validation score from the pre-refit model with a test score from the refit one is a common muddle. Second, you do not re-run the search after folding validation in; that would use those rows for selection again and you would need a fresh validation partition. The exception is a training procedure that needs a held-out block to run at all, such as a stopping rule - then either carve a smaller one out of the enlarged training rows or fix the stopping point you already found.
go deeper
Be ready to say that after the hyperparameters are chosen, the validation rows can usually be added back into training, and that the test set is still scored once at the end.
Explain why adding those rows does not touch the test set, and why the reported number has to come from the model that was actually refit rather than the earlier one.
Demonstrate the exceptions and the discipline: procedures that need a held-out block to train at all, when the refit is not worth the compute, and why you never choose between pre- and post-refit models on the test score.
Own the end-of-project convention: whether shipped models are retrained on everything, what is disclosed when they are, and how the team keeps an untouched partition available for the next round of work.
## Why refit at all A three-way split spends rows on two different things. The training rows buy model quality; the validation rows buy a decision. Once the decision is made and frozen, the validation rows have finished their job, and they are simply labelled examples the final model has not been allowed to learn from. On the steep part of the learning curve that matters. Going from 60% of the rows to 80% is a third more training data, and for a small or medium table that is a visible improvement. On a very large dataset it is invisible, and skipping the refit to keep the pipeline simple is a perfectly defensible call. ## Does it contaminate the test set? No, and being able to say why cleanly is most of the value of this question. Contamination means the reported estimate was influenced by the rows it is computed on. The test rows were not fitted on and drove no decision, and folding validation rows into training does not change that. The validation rows only ever informed a selection, and that selection is already frozen. What *would* contaminate the test set is re-running the search after the refit and re-scoring test to see whether you prefer the answer. That is selection on the test partition, and it converts your one honest estimate into another validation score. ## What the test set is for afterwards After the refit, the test set has exactly one job: produce a single estimate of how the shipped artefact performs on data like this. Some consequences people miss: - **The number describes the refit model.** If you refit, you score the refit model. Reporting the earlier validation number alongside it as though both describe the same object is a muddle that surfaces immediately when the two disagree. - **It is still one evaluation.** The refit does not reset the budget. If the test set had already been scored, refitting does not launder it. - **A small change is not a signal.** If the test score drops slightly after the refit, compare the gap to the standard error at that test size before concluding anything. And whichever way it moved, do not pick between the two models on the basis of the test score - that choice is selection on the test set and it forfeits the estimate. ## When not to refit - **The procedure needs a held-out block.** Some training runs use a held-out block as part of the fitting procedure itself, for example to decide when to stop adding boosting rounds. Folding the validation set into training removes the block the procedure needs. The options are to carve a smaller held-out block out of the enlarged training rows, or to fix the stopping point discovered during selection and train for that fixed amount. - **Retraining is expensive and the gain is negligible.** On a very large dataset the extra rows do not move the metric, and a second full fit costs real time and money. - **The artefact is already validated in some heavier sense.** If the frozen model has been through review, approval or downstream integration testing, replacing it with a differently trained object may require repeating that work, and the accuracy gain has to justify it. ## Should the final model also train on the test rows? This is a real judgment call rather than a rule. Training the shipped model on train, validation and test gives it every row you have, which is tempting when data is scarce. The cost is that you no longer have an evaluated artefact: the number you publish describes *the procedure* applied to that much data, not the specific object you deployed, and you have no untouched rows left for the next verdict. Teams with scarce data do this and disclose it; teams with plenty of data should not bother, because the marginal rows buy nothing and the cost is the loss of a clean hold-out for the next iteration. ## The sequence, stated once 1. Split into train, validation, test before anything else. 2. Select the configuration using validation. 3. Freeze everything - hyperparameters, features, threshold, preprocessing choices. 4. Refit that frozen configuration on train plus validation. 5. Score the test set once, on the refit model, and report the number with its uncertainty. 6. Do not go back to step 2 without saying so and without fresh rows.
- Does refitting on train plus validation contaminate the test set?No. The test rows were never fitted on and drove no decision, and folding validation rows into training does not change either fact - those rows only ever informed a selection that is now frozen. What would contaminate it is re-running the search after the refit and re-scoring test to see which answer you prefer.
- Should the final production model also be trained on the test rows?It is defensible when data is scarce, but it costs the artefact-level guarantee: the published number then describes the procedure trained on that much data rather than the object you shipped, and you have no untouched rows for the next verdict. Teams with plenty of data should not bother.
- The test score drops slightly after refitting on more rows. What do you do?Compare the drop with the standard error of the estimate at that test size - a change inside the noise band is not evidence of anything. Verify nothing else changed in the pipeline. Do not choose between the pre-refit and post-refit models by their test scores; that is selection on the test set.
saying these in an interview costs you the question
- Reports the validation score as if it described the refit model
- Re-runs the hyperparameter search after folding validation into training
- Picks whichever of the two models scores higher on test
- Believes training on validation rows leaks into the test set
- Refits blindly when the procedure needs a held-out block