An inpatient deterioration model's offline AUC rose while escalations that changed care fell after launch - what divergences explain that?
answer
- two numbers, two measurements
- the ward reads twenty rows
- cohort is not the served population
- changed care needs a clinician to act
- recompute at k, then restrict, then novelty
basics
~20 sThree divergences do it: the offline score integrates over a ranking the ward never reads past row twenty, the retrospective cohort is not the population the live path scores, and changed care counts action taken, not ranking accuracy.
solid answer
~40 sStart by naming what each number measured. `ROC AUC` rising from 0.81 to 0.86 is an average over every ordered pair in the ranking, so a candidate can gain across ranks 200-900 while making the top twenty - the only rows a shift consumes - worse. Second, the frozen cohort includes admissions the deployed path never scores, so the two numbers were read on different populations. Third, *escalations that changed care* counts clinician action, so a candidate that promotes patients the team was already watching looks identical offline and adds nothing online. The diagnosis is mechanical: recompute `precision@20` on the cohort, then recompute the offline metric on live-scored patients only, then check how much of the new top twenty the team already knew about.
code
pseudocode · 13 linesscores = rank_desc(score(candidate, todays_patient_days))
// the sweep read: every ordered pair in the ranking counts
auc = mean over (pos, neg) pairs of [ rank(pos) < rank(neg) ]
// the read the shift actually produces
worklist = first(scores, 20)
acted = [ p in worklist if escalated(p) and care_changed(p) ]
online = count(acted)
// so a candidate can raise `auc` while shrinking `online`:
// improve ordering at ranks 200..900 -> many pairs fixed, auc up
// worsen ordering at ranks 1..20 -> few pairs, but the whole worklistgo deeper
Hold on to the shape of the problem: an average over a whole ranking and a count of actions taken on twenty rows can move in opposite directions without either being miscalculated.
Be able to explain the pair mechanics - which pairs ROC AUC averages over, why the deployed list length changes the answer, and what recomputing at k=20 would show.
Run the diagnosis in order and say what each step rules out: metric slice, served population, redundancy with what the team already knew, worklist overlap, then per-segment.
Treat the gap as a measurement asset: decide how much live-shift time the organisation will spend to learn what the cheap proxy cannot see, and what the proxy has to change to be worth trusting next time.
## The setting A rapid-response team works a twenty-row deterioration worklist each shift. A candidate model raised `ROC AUC` on the frozen retrospective cohort from 0.81 to 0.86 against the incumbent. After launch, with the list length unchanged at twenty, *escalations that changed care* fell from 3.1 to 2.6 per shift. Both numbers are correct. They measure different things, and the gap between them is the ordinary condition of an ML system, not a bug report. ## Divergence 1: the offline metric integrates over a ranking nobody reads `ROC AUC` is the probability that a randomly chosen deteriorating patient-day is ranked above a randomly chosen non-deteriorating one. Every pair counts equally, including the millions of pairs formed from ranks the ward will never reach. A candidate that fixes ordering across ranks 200-900 and worsens ranks 1-20 raises the average while shrinking the only slice the shift consumes. - This alone can produce the observed fall, and it is the first thing to test: recompute `precision@20` - the deployed list length - on the same cohort for both systems. - It also explains why the offline metric should have been `precision@20` from the start. Choosing `ROC AUC` bought stability across candidates at the cost of measuring rows nobody sees. ## Divergence 2: the cohort is not the served population The frozen cohort is whatever the data warehouse could assemble. The live path scores whoever it can score, when it can score them. - Admissions discharged before the first score exist in the cohort and never appear on a worklist. - Patients whose inputs arrive late or not at all are scored in the retrospective join and skipped live. - Wards not yet onboarded are in the cohort and not in the deployment. If the candidate's gain is concentrated in patients the live path never scores, the offline improvement is real and unreachable. Recomputing the offline metric restricted to live-scored patient-days separates this from divergence 1. ## Divergence 3: the online read counts changed care, not correct ranking *Escalations that changed care* requires a clinician to act and the action to alter management. Two candidates with identical ranking quality can differ by a lot here: - A candidate that promotes patients the team was already watching scores well against the label and adds no action. - A candidate that surfaces patients late in their trajectory is correct and useless; the team arrives after the decision point. - A candidate whose top rows are dominated by one ward can exhaust that team's attention while others see nothing. Offline, none of these are visible, because the label does not record what the team already knew or when acting was still possible. ## What does not explain a day-one fall A deterioration service that works **prevents the event it predicts**, so a deployed model erodes the very outcome the offline label rewards. This is real, but it is not the explanation here: the incumbent was already preventing events, so both systems' cohorts carry the same effect. What it does mean is that the frozen cohort is not a neutral referee over long horizons - it is a record of care under the incumbent's influence, which is one reason the offline number drifts away from the online one the longer it is reused. ## Diagnosing it in order 1. **Re-read the offline metric at the deployed slice.** Compute `precision@20` for both systems. If the candidate loses, divergence 1 is the whole story and the sweep metric was wrong. 2. **Restrict to the served population.** Recompute on live-scored patient-days only. A gain that vanishes here is divergence 2. 3. **Measure novelty, not just correctness.** For the new top twenty, count how many patients the team had already flagged. A high share is divergence 3. 4. **Compare the two worklists directly.** Overlap between incumbent and candidate top-twenty lists tells you how much of the online change the model is even responsible for. 5. **Split by segment before concluding.** An aggregate fall can be one ward's collapse, and the fix is different from a uniform regression. ## The framing lesson The pairing is not chosen once and forgotten. When the two numbers disagree, the honest response is to treat the gap as information about the proxy - it tells you which property of the online read your offline metric is blind to - and then to narrow it deliberately: fix `k` at the deployed list length, evaluate on the served population, and add a novelty measure if redundancy is what the service is actually losing to.
- Which recomputation do you run first, and why that one?`precision@20` on the same frozen cohort, because it is free and it tests the cheapest explanation: that the offline metric measured ranks the ward never reads. If the candidate also loses at k=20, you have found the whole story without touching the live population, and the fix is to change the sweep metric rather than to investigate the deployment.
- How do you measure the redundancy the offline metric cannot see?Count, for each new top-twenty list, how many patients the team had already escalated or flagged by other means before the alert. That share is the fraction of correct predictions that cannot change care. It needs no new modelling - only that the worklist and the team's existing flags are logged with timestamps so the order of the two is recoverable.
- Does a narrower offline-online gap mean a better model?No - it means a better proxy. Narrowing the gap is a measurement improvement: the offline metric now predicts the online read more reliably, so a sweep is worth more. A model with a wide gap can still be the better one; you just cannot tell from the sweep, and each candidate costs live-shift time to judge.
A city bus network can cut the average journey time across every route while making the one route you ride slower. The twenty-row worklist is that one route.
saying these in an interview costs you the question
- Concludes the online read is broken because the offline number rose.
- Assumes a higher ROC AUC implies a better top-twenty worklist.
- Compares offline and online numbers without checking they cover the same patients.
- Counts a correct prediction the team already acted on as a win.
- Blames the model's own preventive effect for a day-one fall against an incumbent.