Why does picking features by their correlation with the target before cross-validating inflate the score?
answer
- the filter already read the labels
- twenty thousand candidates, all of them noise
- cross-validation protects only what it contains
- selection is a model step, not preparation
- move the ranking inside the fold
basics
~20 sThe selection read every row's label, including validation rows, so the surviving columns were chosen partly because they fit those rows. Selection is part of the model: redo it inside each fold on training rows only.
solid answer
~50 sSelection by correlation with the target is a supervised, high-search-space fit, which makes it the most damaging thing you can run before a split. Take 20,000 pure-noise columns and keep the top 50 by correlation with the label across all rows: some columns will correlate by chance, and because the label of every row — including the validation rows — helped choose them, cross-validating a model on those 50 columns can report around 3% error on data with no signal at all, where the honest error is 50%. The fix is to treat selection as a model step: inside every fold, run the selection on that fold's training rows only, then fit and score. Do that on the same noise and the error goes back to chance. The number you report must be the score of the whole procedure — selection plus model — on rows neither step has seen.
go deeper
Know the headline rule: choosing features using the target column counts as training, so it must happen after the split, on training rows only.
Explain why the optimism grows with the number of candidate columns screened, and describe the corrected loop where the ranking is recomputed inside each fold.
Demonstrate that you can size and prove the damage — re-run with selection in-fold, or shuffle the labels and show the workflow still claims a signal that cannot exist.
Own the definition of a reportable number: what counts as part of the model, which held-out data may be touched and how often, and how that is enforced across teams rather than remembered.
## The textbook demonstration Build a dataset with 20,000 feature columns of pure random noise and a balanced binary label that is also random — by construction there is nothing to learn, and the best achievable error rate is 50%. Now run the sequence that looks perfectly reasonable: 1. Compute each column's correlation with the label, over all rows. 2. Keep the 50 strongest. 3. Cross-validate a classifier on those 50 columns. The reported error comes back in the low single digits — around 3%. Nothing was learned; the number is manufactured entirely by step 1 having been run in the wrong place. ## Why it is so much worse than a leaky scaler Two factors multiply. **It is supervised.** A scaler consumes only the feature distribution; a correlation filter consumes the *labels*, which is precisely the information the evaluation is supposed to withhold. The surviving columns encode a summary of every row's answer. **It searches.** With 20,000 candidates, the maximum sample correlation over pure noise is large by chance alone. Selection keeps exactly the columns whose accidents fit *this* dataset, validation rows included. When the fold is then held out, those accidents are still present in the retained columns, so the model reproduces them and calls it prediction. The more candidates you screen, the bigger the illusion — the optimism grows with the size of the search, not with the size of the signal. ## Why cross-validation does not save you Cross-validation only protects the steps that happen inside it. Freezing the feature set before the loop means every fold's model is built on columns chosen with that fold's labels, so all k estimates are contaminated in the same direction, and averaging them makes the wrong number look stable. A tight standard deviation across folds is not evidence of honesty; it is k repetitions of the same leak. ## The correct placement Selection belongs inside the fold: 1. Split into folds. 2. For each fold, compute the selection rule on the **training portion only** — its correlations, its rankings, its surviving column list. 3. Fit the model on those columns of the training portion. 4. Transform the validation portion with the column list from step 2 and score it. 5. Average across folds. Different folds will now keep different columns, and that instability is itself informative: if the selected set changes wildly, the selection was fitting noise. On the pure-noise dataset above, this procedure returns roughly 50% error, which is the truth. The same rule holds for any selection rule that consults the label — a univariate score, a ranking from a model's coefficients or importances, a wrapper search. Unsupervised filters (dropping constant or near-duplicate columns) leak only the held-out rows' feature distribution and are far milder, but the safe habit is to refit them in-fold too. ## Sizing the damage on a real project When you inherit a suspicious result, two checks settle it quickly. Re-run the whole procedure with selection moved inside the fold and compare the two numbers; the gap is the leak. Or run the original procedure against a randomly shuffled label — an honest pipeline collapses to chance, while a leaky one still returns a flattering score. That second check is the most convincing demonstration to a sceptical colleague, because it needs no theory: the data provably contains no signal, and the workflow still claims one. ## The rule to carry away Anything that consults the target and produces a decision reused at scoring time — which columns to keep, how a category maps to a number, which rows to drop — is part of the model. It is refit whenever the model is refit, and it is measured on data that neither it nor the model has seen. The honest score is always the score of the entire procedure end to end.
- How would you prove to a sceptical colleague that the 3% figure is fake?Randomly shuffle the labels and re-run the identical workflow. The shuffled data provably contains no signal, so an honest procedure must return chance error. If the same workflow still reports single-digit error, the number is manufactured by the step order, and no argument about the model is needed.
- Does the problem disappear if the selection rule never looks at the target?It shrinks a lot but does not vanish. Dropping constant or near-duplicate columns uses only the held-out rows' feature distribution, which is the mild kind of leak, comparable to a scaler fitted too early. It is still safest to refit the rule inside each fold so the reported score covers the whole procedure.
- Different folds now keep different feature sets. Is that a problem?No, it is information. Cross-validation is estimating the performance of the *procedure*, not of one frozen feature list, so folds are allowed to disagree. Wildly unstable selections tell you the rule is fitting noise; if you need one final feature set, choose it by running the selection once on the full training data after the estimate is in hand.
It is like choosing the exam questions after seeing which ones the class happens to get right, then reporting the class average as evidence of teaching.
saying these in an interview costs you the question
- Calls feature selection preprocessing rather than part of the model
- Says selecting on all rows is fine because no model was fitted yet
- Trusts a cross-validated score computed on a pre-frozen feature set
- Reads low variance across folds as proof the estimate is honest
- Blames the optimism on the classifier instead of the step order