Why does oversampling the minority class before the train/test split inflate the validation score?
answer
- check both sides of the split
- a copied row is not an unseen row
- validation must keep production's class balance
- resample inside the training fold only
basics
~20 sCopies or near-copies of minority rows land on both sides of the split, so the model is scored on rows it trained on, and validation loses the real class balance. Split first, resample the training part only.
solid answer
~40 sTwo separate things break. First, duplication crosses the boundary: if a minority row is copied and one copy is drawn into training while another lands in validation, the model can memorise it and the memory is scored as generalisation, so minority recall looks far better than it is. Synthetic oversampling such as SMOTE is no safer — a synthetic point built from neighbouring minority rows can be almost the same record as a validation row. Second, the validation set's class prevalence has been changed, so precision and any decision threshold chosen on it will not hold at the real base rate. The safe order is: make a stratified split first, then resample only the training portion, refitting that resampling inside each cross-validation fold, and evaluate on untouched, naturally imbalanced data.
go deeper
Remember the order: split first, then balance only the training data. Know that duplicating rows before splitting can put the same record in both training and validation.
Explain both failure modes — memorised duplicates crossing the boundary and a held-out set whose class balance no longer matches production — and say which metrics each one distorts.
Diagnose it in someone else's results: look for duplicate rows across the split, a suspiciously balanced validation set, and a precision figure that cannot survive the real base rate.
Decide the evaluation contract for the team: which prevalence metrics are reported at, where thresholds are chosen, and how resampling is kept inside the training path by construction.
## The setup A fraud or churn dataset arrives at 2% positives. A common reflex is to balance it — duplicate minority rows, or generate synthetic minority points — and then split into train and test. Cross-validated recall comes back at 0.95, everyone is delighted, and production recall is a third of that. The mistake is one line of ordering. ## Failure one: the same row on both sides Random oversampling literally copies minority rows. If the copies are made before the split, the shuffle scatters them: copy A goes to training, copy B goes to validation. The model sees copy A during fitting, and then is asked to classify copy B, which is byte-for-byte identical. Any model with enough capacity gets it right by memory. Because the minority class is exactly the class you duplicated, the inflation lands squarely on minority recall — the metric everyone is watching. Synthetic oversampling does not escape this. A synthetic minority point is constructed from real minority rows; if one of its parents ends up in validation, the training set contains a point sitting essentially on top of a validation row. The generated point was not in the original data, but the information in it was. Undersampling the majority class before splitting behaves differently: it creates no duplicates, so it does not produce this memorisation leak. It still distorts the split in the second way below, and it discards rows that should have been available for evaluation. ## Failure two: the validation set no longer looks like production Even with no duplicate crossing the boundary, balancing before the split changes the class prevalence of the held-out data. Prevalence-sensitive metrics then report numbers that cannot transfer: - **Precision** depends directly on the base rate. At 50% positives a model looks precise; at 2% the same model's positive predictions are mostly false alarms. - **Accuracy** becomes meaningless in the other direction — the trivial majority baseline moves from 98% to 50%. - **A chosen threshold** tuned on balanced validation data will fire far too often on real traffic. Recall and ROC-AUC are insensitive to prevalence, which is one reason a balanced-validation setup can look internally consistent while precision collapses on deployment. ## The leakage-safe order 1. Split first, stratified on the label so both sides keep the natural prevalence. 2. Inside the cross-validation loop, after the fold split, resample **only** that fold's training portion. 3. Leave the validation fold and the final test set exactly as they came — imbalanced. 4. Report metrics at the real prevalence, and pick the operating threshold on data at that prevalence. Resampling is a fitted step like any other: it is part of the training procedure, not part of the data preparation that precedes evaluation. ## What the honest numbers look like When the order is fixed, the minority recall usually falls and precision becomes interpretable. That drop is not a regression — the earlier number was measuring memorised duplicates. A quick diagnostic when you inherit a suspicious result: check whether any validation row is an exact duplicate of a training row, and check the class balance of the validation set against the raw table. A balanced validation split in an imbalanced problem is a strong smell all by itself. ## Spotting it in a review Read top to bottom and find the line that changes the number of rows. If it appears before the split, or before the fold loop, the evaluation is compromised. The same rule catches other row-changing steps in the wrong place: de-duplicating, dropping outliers, or filtering rows by a data-derived rule. ## What to say in an interview Name both failures — duplicated information crossing the split, and a validation set whose class balance no longer matches deployment — then give the order: stratified split, resample the training fold only, evaluate on natural prevalence. Mentioning that undersampling avoids the first failure but not the second shows you understand the mechanism rather than reciting a rule.
- Where exactly does resampling belong inside a cross-validation loop?Inside the loop, after the fold split, applied to that fold's training portion alone. The validation portion keeps its natural class balance and is never resampled. Each fold rebuilds the resampling from scratch, so no generated or duplicated row can reach the data used to score that fold.
- Why not balance the validation fold too, so the metrics are computed on an even split?Because deployment runs at the imbalanced prevalence. Precision, accuracy and any threshold picked on a balanced validation set will not hold on real traffic. Report at the natural base rate; if you want a prevalence-independent view, use recall or ROC-AUC rather than rebalancing the evaluation data.
- Does undersampling the majority class before splitting cause the same leak?Not the duplication leak — it copies nothing, so no row appears on both sides. It still distorts the held-out set's class prevalence, making precision and thresholds untransferable, and it throws away majority rows that should have been available for evaluation. Do it inside the training fold.
saying these in an interview costs you the question
- Balances the classes first, then splits, calling it preprocessing
- Reports recall from an artificially balanced validation set as production performance
- Thinks duplicated rows across the split are harmless because they are real
- Assumes synthetic points cannot leak since they were not original rows
- Tunes the decision threshold on a rebalanced validation set