A claim-prediction model scores 0.997 held-out AUC. What is your first hypothesis, and how do you test it?
answer
- Too good is a bug report
- Suspect the data before the model
- Which single column reproduces the score?
- When is that value written?
- Drop it, retrain, watch it fall
basics
~20 sAssume target leakage before assuming success. On a real business problem, 0.997 almost always means a feature encodes the outcome. Test it by scoring each feature alone, auditing when the top one is written, and retraining without it.
solid answer
~50 sMy first hypothesis is target leakage, not a great model: claim, default and readmission problems rarely separate that cleanly, so I treat 0.997 as a bug report. I test it in cost order. First, fit a model on each candidate feature alone and see which single feature reproduces most of the score, because leakage usually concentrates in one or two columns. Second, trace those columns to the source table and establish when the value is physically written relative to the decision moment; a field like a payment-plan start date only ever exists for accounts that already defaulted. Third, retrain with the suspects dropped and see where the score lands. A fall to a plausible 0.72 confirms it. If no single feature explains the score, I check whether the label definition itself was derived from a field that also entered the feature set.
go deeper
Know the reflex: a near-perfect score on a business problem is a reason to investigate, not to celebrate. Be able to say you would check where the strongest feature comes from before reporting anything.
Be ready to actually run the diagnosis — rank features by their standalone predictive power, retrain without the leading one, and quantify how much of the score that single column was carrying.
Show the full loop including the human step: going to the owner of the source table to learn when each field is written, then reporting a corrected, lower number with confidence rather than defending the original.
Frame the cost. A leaked model that ships makes confident wrong decisions and burns credibility for the next one. Decide what evidence a model owes before it touches a decision, and who is accountable for signing that off.
## Why 0.997 is a bug report Area under the ROC curve is the probability that a randomly chosen positive case is ranked above a randomly chosen negative case; 0.5 is coin-flipping and 1.0 is perfect separation. A value of 0.997 says that for essentially every positive-negative pair, the model puts them in the right order. Real business outcomes — whether a claim gets filed, whether an account defaults, whether a patient returns within thirty days — are driven substantially by events that have not happened yet at prediction time. Nothing you can measure beforehand orders them that well. So the prior matters. Across the population of real projects, near-perfect held-out scores are produced by data defects far more often than by exceptional models. The correct professional reflex is to treat the number as evidence of a bug and go looking for it, rather than to write it in a slide. ## The diagnostic ladder Run the cheap checks first. **1. Sanity-check the plausible range.** Ask a domain expert, or look at what published or internal baselines achieve on the same decision. If the best anyone has managed is 0.78, your 0.997 needs an extraordinary explanation. **2. Localise the signal.** Fit a model on each candidate feature in isolation and rank the resulting scores. Leakage is usually concentrated: one column alone reaches 0.99, and everything else sits between 0.55 and 0.65. This is more informative than a global importance ranking, because it tells you what a feature can do on its own rather than how a fitted model happened to split its attention. **3. Audit the suspect's write time.** This is the step that actually decides the case, and it is not a statistical step. Go to the source system or its owner and establish: when is this value first written, is it ever updated afterwards, and does its existence depend on the outcome? A payment-plan start date is populated only after an account defaults. A settlement amount is created by the claim. A discharge disposition may be rewritten after a readmission. None of these are visible from the data alone. **4. Ablate and re-measure.** Drop the suspects, retrain, and compare. A collapse from 0.997 to something in the 0.65-0.80 range is the confirmation, and the lower number is your real starting point. If the score barely moves, you have not found the leak yet — keep going. **5. Audit the label.** If no single feature explains the separation, look at how the target was computed. If the label was derived from a column that also entered the feature table, or from anything downstream of it, the model can effectively invert the derivation. Read the label definition and exclude every source field it touches, plus anything derived from them. ## Ruling out the benign explanation Sometimes a problem really is easy — you are detecting a condition that a sensor almost directly measures, or the positive class is defined by a rule involving observable inputs. Two checks separate that from leakage. - **Availability.** Confirm that the deciding feature is present, populated and holding the same value at the moment the prediction must be served, not just in the historical table. - **Base rate and operating point.** With a 0.2% positive rate, area under the ROC curve flatters a model; precision at the threshold you would actually deploy can still be poor. A high AUC with disappointing precision is a different problem from leakage, and it is worth distinguishing before you blame the data. ## What you do once it is confirmed Remove or re-derive the offending features, retrain, and report the corrected number as the real result. Then close the loop on process: record the leaked columns in the data dictionary as post-outcome, and build training tables from an approved feature list rather than by dropping columns that look wrong. Handle the communication deliberately. If a stakeholder has already seen 0.997, explain that it was unachievable rather than lost — the model was reading the answer — and give them the honest figure with what it means for the decision they wanted to automate. Shipping a leaked model is worse than shipping no model: it degrades silently, produces confident wrong decisions, and destroys trust in the next model you build.
- The strongest feature turns out to be legitimate and available at prediction time. What now?Then the problem may genuinely be easy, but verify before celebrating. Confirm the stored value is the one available at the decision moment rather than a later overwrite, and check the class balance: with a 0.2% positive rate the ranking metric flatters the model while precision at your deployment threshold may be poor. Report the operating-point numbers, not just the area under the curve.
- How do you rule out leakage hiding in the label definition itself?Read how the target was computed and list every source field it touches. If any of those fields, or anything derived from them, also entered the feature table, the model can invert the derivation and the separation is meaningless. Exclude the whole derivation chain, retrain, and see whether the score survives.
- What do you tell a stakeholder who is already excited about the 0.997?Say plainly that the number was never achievable: the model was reading a field created by the outcome. Give the corrected figure, explain what it supports and what it does not, and name the cost you avoided — a leaked model fails silently in production and every decision made on it would have been unvalidated.
A student who scores full marks on every practice paper is either exceptional or holding the answer key. You check for the key first, because that explanation is far more common.
saying these in an interview costs you the question
- Ships it because the test set was held out
- Blames luck and re-splits with a different random seed
- Deletes the top feature without checking when it is written
- Assumes the problem is simply an easy one
- Adds regularisation to bring the score down to something believable