A readable pneumonia-risk rule list learned that asthma lowers risk — how do you respond?
answer
- the pattern is real in the data
- who got treated, and how fast
- the label carries the care policy
- held-out accuracy will not flag it
- a black box learns it invisibly too
basics
~20 sThe rule is a true pattern in the data and a lethal policy. Asthmatic pneumonia patients were routed straight to intensive care, so they survived more often. The label reflects the treatment they received, not their underlying risk.
solid answer
~50 sFirst, the model is not broken: in the historical records asthmatic pneumonia patients really did die less often, because they were fast-tracked into intensive care. The model learned the hospital's treatment policy, which is baked into the outcome, rather than the biology of risk. That makes it catastrophic for the intended use, a triage tool that would send home the group which survived only because it was admitted. The fix is not a better fit: remove or hand-edit that term, restrict the model to information available before any triage effect, or redefine the target as risk under a fixed care policy, and have a clinician review every rule for the same pattern. The interpretability lesson matters most. A black box on the same records learns the same relationship and hides it; the readable model is the one an expert can spot it in, and edit.
go deeper
Recall that a learned association can be genuine in the data and still wrong to act on. Be able to say that the asthmatic patients survived more often because they were treated more aggressively.
Explain the mechanism by which the treatment policy enters the label, and why standard validation endorses rather than flags the rule. Name the general test: could this feature have influenced how the case was handled?
Lay out a concrete remediation plan: constrain or edit the term, restrict inputs to the pre-decision information set, redefine the target, and run a systematic domain review of every rule rather than only the one that looked odd.
Own the policy: which high-stakes decisions require a model whose every term a domain expert has signed off, when a model may rank but never withhold, and how you institutionalise the review so the next inverted rule is caught before deployment rather than after.
## What happened A model was trained to predict mortality risk for pneumonia patients so that low-risk patients could be treated as outpatients. A readable rule-based model produced, among its rules, one saying in effect: *having asthma lowers this patient's risk of dying*. That is clinically absurd — asthma is a respiratory condition and pneumonia is a respiratory illness — and it is also, in the training data, **true**. Clinicians treated asthmatic pneumonia patients as high risk and routed them directly into intensive care. The aggressive care worked. So in the historical records, the asthma group shows *lower* observed mortality. The model was right about the data and would have been lethal in deployment: used for triage, it would have recommended sending home exactly the patients whose survival depended on being admitted. ## Name the mechanism This is not noise, not overfitting, and not a bug in the learner. It is **the outcome being contaminated by the intervention** — often called confounding by indication, or treatment leakage. The recorded outcome is not "risk of dying from pneumonia" but "risk of dying given the care this patient actually received", and the care depended on the very features the model uses. Whenever a historical label was produced under a policy that reacted to the inputs, the model learns the policy along with the phenomenon. The general test: **would this feature have influenced the treatment or the process that generated the label?** If yes, the learned relationship may be inverted relative to the causal one. Look for it in every domain where a decision intervened between the features and the outcome: past fraud reviews, past loan approvals, past manual interventions on flagged cases. ## What you actually do about it Several responses, usually in combination: 1. **Do not ship the rule.** In a rule list or a sparse model you can delete or overwrite the offending term and refit the rest around the constraint. That is a real advantage of readable models: they can be hand-edited under domain review. 2. **Constrain the direction.** If the domain says the effect of a risk indicator must not be protective, encode that as a constraint on the term's direction rather than hoping the data behaves. 3. **Fix the information set.** Restrict the model's inputs to what is known at the true decision point, before any triage response has occurred. Anything recorded after the decision, or that proxies for the decision, is leakage. 4. **Redefine the target.** Model risk under a fixed care policy, or model the untreated course of illness, rather than the observed outcome under a responsive policy. This is harder and may require a subgroup where the policy did not apply. 5. **Review every rule with a clinician, not just the suspicious one.** If one rule inverted, others did too and were merely less obviously wrong. The audit has to be systematic. 6. **Change the deployed decision.** Sometimes the honest answer is that the model may rank patients for attention but must not be permitted to *withhold* care — a decision-policy change, not a modelling change. ## Why this is an argument about interpretability, not about medicine The critical point for an interview: **a black box trained on the same records learns the same relationship**. It is in the data; a flexible model will find it. The difference is that in the ensemble it exists as a fragment of a decision surface nobody inspects, contributing quietly to millions of predictions. In the rule list it exists as a line of English that a clinician read and objected to within seconds. So the readable model did not create the problem; it *surfaced* one that would otherwise have shipped silently. That is the strongest available argument for interpretable-by-design models in high-stakes settings, and it is stronger than the usual appeal to trust or comfort. It also explains why aggregate validation metrics do not protect you here: the dangerous rule *improves* held-out accuracy, because the held-out data was generated under the same care policy. Nothing in the score curve flags it. Only a human reading the model does. ## Interview framing Resist the reflex answer of "it is spurious correlation, drop the feature". Say instead that the correlation is real, explain the mechanism that produced it, note that validation cannot catch it because the test data carries the same policy, then give the fixes and the deployment restriction. Close with the interpretability point — that the readable model is the one where the failure is visible and editable. That sequence is what a senior answer looks like here.
- Would cross-validation or a held-out test set have caught this rule?No. The held-out data was generated under the same care policy, so the asthma rule improves held-out accuracy too. Every aggregate metric endorses it. The failure only appears when the model is deployed to change the policy that created the pattern, or when a human reads the rule and objects.
- Is simply dropping the asthma feature a sufficient fix?Rarely. Other features correlated with fast-tracking will pick up the same signal, so the inverted effect reappears under another name. You need to address the mechanism: restrict inputs to the true decision point, constrain effect directions from domain knowledge, and audit the remaining terms with a clinician rather than removing one column and declaring victory.
- Where else does this pattern show up outside healthcare?Anywhere a past decision intervened between features and outcome. Fraud models trained on transactions where flagged cases were already blocked, credit models trained only on approved applicants, and churn models where at-risk customers were given retention offers all learn the intervention as if it were the phenomenon.
It is like concluding that carrying an umbrella keeps you dry in a storm and therefore banning umbrellas from the forecast: the protective association exists only because of the response it triggered.
saying these in an interview costs you the question
- Calls the rule a spurious correlation with no mechanism
- Says the model overfitted and needs more regularisation
- Believes a held-out test set would have flagged it
- Drops the one feature and declares the problem fixed
- Claims a black box would not have learned the same relationship