Why is it wrong to fit a scaler on the full dataset before splitting into train and test?
answer
- some transforms learn numbers first
- who was allowed to see the test rows
- mean and SD carry test-set information
- refit inside every cross-validation fold
- split first, then fit, then apply
basics
~20 sFitting a scaler on all rows lets the test rows' mean and standard deviation shape the transform, so the test set is no longer unseen and the score is optimistic. Fit on training rows only, then apply it to test.
solid answer
~50 sStandardisation is a fitted step: it learns a mean and a standard deviation from data and then applies `(x - mean) / sd`. If those two numbers are computed over every row, the test rows have contributed to the transform used on the training rows and to their own scaled values, so the test set is no longer data the procedure has never seen and the score it produces is optimistic. The safe order is: split first, fit every stateful step on the training portion only, then apply the already-fitted numbers unchanged to validation and test. Under k-fold cross-validation the same steps must be refit inside every fold on that fold's training part, not once on the whole training table. For plain standardisation with many rows the inflation is small, but the same mistake with a median imputer, a target encoder or a selector is far worse.
go deeper
Be ready to state the rule and the order out loud: split first, fit the scaler on training rows, apply those same numbers to test. Know that a scaler learns a mean and a standard deviation.
Explain the mechanics — which parameters each step learns, why refitting must happen inside every cross-validation fold, and why the inflation is small for a scaler but large for a target-dependent step.
Show that you can audit an existing workflow: point at the lines above the split that compute something across rows, and describe how you would re-measure to size the damage.
Own the structural fix. Argue for encapsulating the whole fitted sequence so the correct order is the default rather than a review checklist item, and for one final test set touched once.
## What "fitting" a preprocessing step means Preprocessing splits into two kinds of operation. A **stateless** operation is computed from one row and nothing else: taking the log of a positive column, dividing two fields in the same row, parsing a date, converting a unit. A **fitted** (stateful) operation first estimates parameters from a sample of rows, then applies them: - standardisation learns a column mean and standard deviation, then applies `(x - mean) / sd`; - min-max scaling learns a minimum and a maximum; - median imputation learns the column median used to fill blanks; - one-hot encoding learns the vocabulary of categories; - a target encoder learns a per-category statistic of the label; - feature selection learns which columns survive; - PCA learns the component directions. Everything in the second list has a *fit* phase and a *transform* phase, and it is the fit phase that can leak. ## Why fitting before the split leaks The point of a held-out split is to estimate how the **whole procedure** behaves on rows it has never seen. If the mean and standard deviation were computed over all rows, then the held-out rows contributed to those numbers. Two things go wrong at once: the training features were built with knowledge of the held-out distribution, and each held-out row was scaled partly by itself. Neither situation exists at deployment time — tomorrow's request cannot contribute to statistics that were frozen when you shipped the model. The evaluation is therefore measuring an easier problem than the real one. ## How big is the leak? It scales with two things: how much the fitted step depends on the *target*, and how few rows each learned parameter is estimated from. - **Standardisation on a large sample:** small. The mean over all rows and the mean over the training rows differ little, and the inflation is typically a fraction of a point. Min-max is more fragile because it is driven by extremes: one outlier that happens to live in the test set moves the transform. - **Median imputation of a missing-income column computed on the whole table, then folded:** the same shape of mistake, still mild, but a number derived from the validation rows is now baked into every imputed training row. - **A merchant-category target encoder fitted on all rows before folding:** severe, because the encoded feature contains the validation rows' *labels*. A category with a single occurrence gets encoded with that row's own outcome, and the model can read the answer straight out of the feature. - **A supervised feature selector run before folding:** worst of all, because it searches many columns using every label. Small sample sizes amplify all of these. ## The leakage-safe order 1. Split first — train/test, and inside cross-validation, into folds. 2. Fit every stateful step on the training rows only, in sequence. 3. Apply the fitted parameters, unchanged, to validation and test. 4. In k-fold, repeat steps 2 and 3 **inside every fold**: refit, never reuse a fit made on other rows. 5. Touch the final test set once, at the end. The useful mental model is that the preprocessing sequence and the model are one object. Anything you would refit when new training data arrives belongs inside the fold. ## What is genuinely safe before the split Row-wise deterministic work that consults no other row: logs, ratios, date parts, unit conversion, dropping or renaming a column by name. Also transforms whose constants come from outside the sample — a published currency table, a domain-fixed category list — because nothing was estimated from this data. Be suspicious of "cleaning" that is a fitted step in disguise: removing outliers beyond three standard deviations, dropping near-constant columns, or de-duplicating on a similarity threshold all consult the whole table. ## Spotting it in a workflow Read the code from the top and, for every line above the split, ask: *does this line compute a number from more than one row?* If yes, it is a fit and it is in the wrong place. Warning signs downstream include a validation score that beats the final test score, a score that quietly drops when someone reorganises the code, and an offline-to-live gap nobody can explain. The direct check is to redo the run with the correct order and compare: a leaky scaler moves the number slightly; a leaky target-dependent step moves it a lot.
- Which preprocessing steps are safe to apply before the split?Row-wise deterministic ones that consult no other row: a log or square root, a ratio of two fields in the same record, parsing a date, converting units, dropping a column by name. Anything that estimates a parameter from the sample — a mean, a standard deviation, a median, a category statistic, a surviving feature list, PCA directions — must be fitted on training rows only.
- How does this change when you use k-fold cross-validation instead of a single split?Every fitted step must be refit inside each fold on that fold's training portion. Fitting once on the whole training table before folding still leaks each validation fold's rows into the transform that is used to score them. Treat preprocessing plus model as a single object that is rebuilt from scratch k times.
- After fixing the order the score dropped. Which number do you report?The lower one. The earlier score was measured on rows the transform had already consumed, so it estimated performance on an easier problem than the real one. The drop is not a regression; it is the first honest measurement, and it is the number that should be compared against production.
It is like letting students see the exam paper while you write the study guide: their marks no longer tell you how well the guide teaches anyone else.
saying these in an interview costs you the question
- Says scaling ignores the target, so it cannot leak
- Calls fitting on all rows fine because the transform is unsupervised
- Refits the scaler on the test set to scale it properly
- Treats preprocessing as outside the model, so outside the split
- Assumes the leak is always negligible, so order never matters