What is target leakage, and why does a leaky feature still score well on a held-out test set?
answer
- Great offline, useless live
- The split is not the culprit
- Ask when each value is written
- Exists only because the outcome happened
basics
~20 sTarget leakage means a feature carries information recorded only after the label event. The held-out rows were built from the same post-outcome data and carry the same column, so offline scores look excellent and production collapses.
solid answer
~50 sTarget leakage is when a feature carries information that only came into existence after the outcome it is supposed to predict. A classic case: an insurance table keeps a `settled_amount` column while the model predicts whether a claim will be filed at all. The settlement only exists because the claim exists, so the model learns "amount is present" and scores near-perfectly. A held-out split does not catch it, because the split is over rows while the defect is in a column: train rows and test rows were both built from the same post-outcome snapshot, so the shortcut works on both sides. The failure appears only at serving time, when the field is empty or holds its pre-outcome value. The check is temporal, not statistical: for each feature ask when its value is physically written, relative to the moment the prediction is needed.
go deeper
Be ready to define target leakage in one sentence and name a concrete leaky column. Interviewers want to hear that the value was created after the outcome, not that it correlated strongly with the label.
Expect to explain why a random split cannot detect it: the leaked column sits on both sides of the split, so the defect is in the data rather than the resampling. Walk through a per-feature availability check.
Demonstrate the habit of auditing a feature list against the decision timestamp before trusting any score, and say how you would prove the leak empirically by retraining without the suspect column.
Own the framing that leakage is a data-contract problem: which fields are guaranteed at decision time, who records their write semantics, and what mechanism stops a post-outcome column reaching a training table again.
## The definition **Target leakage** is a defect in the *dataset*, not in the model. One or more features carry information that did not exist at the moment the prediction has to be made, and they only carry it because the outcome has already happened by the time the training table was assembled. The decisive test is temporal, not statistical. For every feature, ask one question: **at the instant this prediction is needed in production, is this exact value already written — and is it this value?** If the value is written later, or is later overwritten, the feature is leaked. ## Three shapes it takes **Post-outcome fields.** An insurance claims table carries a settled amount alongside a flag for whether a claim was ever filed. The settlement is created *by* the claim. A model predicting whether a claim will be filed learns that a non-null settlement means yes, and separates the classes almost perfectly. **Proxies and near-copies.** The leaked field does not have to be the label. A hospital record's discharge disposition code may be written or rewritten only after a readmission has occurred. It is not the readmission flag, but it is close enough that a tree can reconstruct it in one split. **Operational artefacts.** A collections dataset contains a payment-plan start date. Payment plans are only ever opened for accounts that have already defaulted, so the field is non-null on almost exactly the positive class. Nobody put it there maliciously; it is simply what the operational system records, and the modelling table inherited every column. ## Why the held-out set does not save you The intuition behind a held-out set is that the model has never seen those rows, so a high score means it generalises. That intuition assumes the *only* way a model can cheat is by memorising rows. Leakage is a different cheat: the shortcut lives in a column, and every row in the table — held out or not — was generated by the same process, after the outcome was known. The test rows carry the leak just as faithfully as the training rows, so the score is high everywhere and nothing looks wrong. This is exactly what separates leakage from **overfitting**. Overfitting is the model memorising noise in the training rows; you detect it because held-out performance falls below training performance. Leakage keeps *both* numbers high. That is why "the train and test scores agree, so we are fine" is a dangerous sentence: agreement between them rules out overfitting and says nothing at all about leakage. ## What it looks like in practice Symptoms cluster: - A score near the ceiling for a problem the business has never solved — 0.99-plus area under the ROC curve, or accuracy far above what a domain expert would guess. - One feature dominating every importance ranking, with the model collapsing to near-random when it is removed. - A field whose missingness alone is almost a perfect classifier. - A model whose live performance, once deployed, falls to roughly the base rate — because in production the leaked field is null, defaulted, or not yet written. ## Finding it The cheapest detection is not statistical. It is reading the column list against the decision moment and, where you cannot tell, asking whoever owns the source table when each field is written. A short structured pass works well: 1. Write down the **decision timestamp** — the exact event that triggers a prediction (a quote request, an admission, an account entering arrears). 2. For each feature, record when the value is first written and whether it is ever updated afterwards. 3. Flag anything written after the decision timestamp, anything updated afterwards, and anything whose existence is conditional on the outcome. 4. Confirm suspects empirically: fit a model on each suspect feature alone, and retrain the full model with the suspects dropped. A collapse from near-perfect to a plausible number confirms the diagnosis. ## Fixing it Three options, in order of preference. **Remove** the feature outright, if nothing legitimate is left in it. **Re-derive** it as of the decision moment — sometimes the underlying system does hold the earlier value, and the modelling table simply took the latest one. **Keep it** if the audit shows the value genuinely is available and unchanged at prediction time; a strong legitimate predictor is not leakage. ## The nuance that trips candidates Correlation strength is not the criterion. A credit score correlates powerfully with default and is available before the decision — that is a good feature, not a leak. A settlement amount correlates powerfully and is unavailable — that is a leak. The distinction is availability in time. Conversely, a leaked feature is not always suspiciously named; the most damaging ones are innocuous operational fields that happen to be created by the outcome.
- Is a feature that correlates very strongly with the target automatically leakage?No. A strong legitimate predictor is exactly what you want. Leakage is about availability in time, not correlation strength: the question is whether the value exists, with that same value, at the moment the prediction must be made. A credit score correlates strongly with default and is available beforehand; a settlement amount correlates strongly and is not.
- How is target leakage different from overfitting?Overfitting is the model memorising noise in the training rows, and you catch it because held-out performance drops below training performance. Leakage is a defect in the data, so held-out rows carry the same shortcut and both numbers stay high. More data or stronger regularisation mitigates overfitting; only removing or re-deriving the feature fixes leakage.
- A post-outcome column must stay in the warehouse for reporting. How do you keep it out of the model?Do not rely on remembering to drop it. Mark it in the data dictionary as written after the decision, and build the training table from an explicit allowlist of approved features rather than by taking every column and removing the ones that look wrong. Allowlists fail closed when someone adds a new column upstream.
It is like grading an exam where the answer key is stapled to every paper, in both the practice pile and the marking pile. Every score is high, and none of them predicts how students do without the key.
saying these in an interview costs you the question
- Says a random train/test split prevents leakage
- Treats any strong correlation with the target as leakage
- Calls it overfitting and adds regularisation
- Trusts the offline score because the test set was untouched
- Only suspects columns with suspicious-looking names