Two deterioration rankers: the candidate wins on pooled weekly precision but loses in both wards - which reading do you trust?
answer
- a reversal, not a gap
- the mix moved, not only the rate
- segment sizes changed between systems
- pooled and per-ward answer different questions
- report coverage beside each segment precision
basics
~20 sNeither reading is wrong; they answer different questions. The reversal comes from the candidate moving worklist rows toward the higher-prevalence ward, so the honest read is whichever matches how worklist capacity is actually allocated - and both get reported.
solid answer
~40 sThis is a mix shift, not a paradox to be resolved by picking a favourite number. The candidate spent more of the shared 200-row weekly budget on the ward with the higher deterioration base rate, so it can lose within each ward and still win pooled. Which reading gates the launch depends on how capacity works: if one hospital-wide team works a single pooled list, the pooled number is reachable and real; if each ward has its own team and its own fixed twenty rows per shift, no system can move rows between wards and only the per-ward numbers are attainable. Report both, and report the **coverage shift** - surgical alerts falling from 160 to 40 is a clinical decision, not a rounding detail in a launch metric.
code
pseudocode · 15 linesbudget = 200 // pooled worklist rows for the week
for each system in [incumbent, candidate]:
rows = first(rank_desc(score(system, week_patient_days)), budget)
pooled = precision(rows)
for each ward in [medical, surgical]:
seg = [ r in rows if r.ward == ward ]
report(system, ward, count(seg), precision(seg)) // count FIRST
report(system, "pooled", count(rows), pooled)
// incumbent -> medical 40/50%, surgical 160/15%, pooled 200/22%
// candidate -> medical 160/45%, surgical 40/10%, pooled 200/38%
// the reversal is only readable because count(seg) sits beside precision(seg)go deeper
Remember that combining groups can reverse a comparison when the group sizes also change - a rate table without counts cannot show you that.
Explain the mechanism: the candidate moved worklist rows toward the higher base-rate ward, and that reallocation outweighs its within-ward ranking loss.
Tie the choice of reading to the capacity model - pooled list against per-ward teams - and insist on coverage counts and a pre-agreed segment list in the launch report.
Own the allocation question the numbers are really about: a ranker that shifts attention between wards is making a clinical priority decision, and you decide who signs it off.
## The numbers One pooled worklist of 200 rows per week, drawn across a medical and a surgical ward, scored by precision - the share of worklisted patient-days that really deteriorated. | ward | incumbent rows / true | incumbent precision | candidate rows / true | candidate precision | |---|---|---|---|---| | medical | 40 / 20 | 50% | 160 / 72 | 45% | | surgical | 160 / 24 | 15% | 40 / 4 | 10% | | **pooled** | **200 / 44** | **22%** | **200 / 76** | **38%** | The candidate is worse in the medical ward (45% against 50%), worse in the surgical ward (10% against 15%), and better pooled (38% against 22%). No arithmetic is wrong. This is **Simpson's paradox**: a comparison that reverses when segments are combined, because the segment *sizes* moved at the same time as the rates. ## Why the reversal happens The medical ward has a far higher base rate of deterioration - half its worklisted rows are true, against 15% on the surgical side. The pooled precision of a system is therefore driven mostly by **how it splits the 200-row budget**, not by how well it ranks inside each ward. - The incumbent put 160 of 200 rows on the low-prevalence surgical ward. - The candidate put 160 of 200 on the high-prevalence medical ward. - That reallocation is worth more than the candidate's within-ward ranking loss, so the pooled number rises while both per-ward numbers fall. The lesson is not that the pooled figure lies. For the same 200 rows of team attention, the candidate surfaces 76 true deteriorations against 44. If the constraint really is 200 rows of attention anywhere in the hospital, that is a genuine improvement. ## Which reading is the honest one depends on how capacity is allocated | capacity model | the reachable read | why | |---|---|---| | one hospital-wide rapid-response team working a pooled list | pooled precision | the system genuinely controls the split, so the reallocation is a real choice it made | | a team per ward with a fixed 20-row shift list each | per-ward precision | no ranker can move rows between wards, so the pooled figure describes an allocation nobody can execute | | a pooled list with a per-ward floor | both, with the floor stated | the floor caps the reallocation, so the pooled gain is only partly available | This is a framing question, not a statistics question: it is answered by asking who works the list, not by choosing an estimator. ## What per-segment reporting has to carry A per-segment table with rates alone hides the mechanism that produced the reversal. Every segment row needs: 1. **The row count for that segment**, under each system - the coverage. Rate without count cannot show a reallocation. 2. **The segment's base rate**, so a low precision can be read as hard rather than broken. 3. **The change in coverage between systems**, called out where it is large. Surgical coverage dropping from 160 rows to 40 is the headline, not a footnote. 4. **The segments chosen in advance**, by ward and by admission route, so this is a standing report and not a slice discovered after a disappointing result. ## The framing commitment The practical output of this discussion is a rule written down before any candidate exists: which segments are reported every time, which reading the launch turns on given the capacity model, and whether any segment carries a floor below which an aggregate win does not authorise a launch. Without that rule, the two tables become an argument settled by whoever prefers their own number. With it, the reversal is a routine event with a defined response - and a real one: a candidate that concentrates attention on the medical ward is making a clinical allocation decision, and that decision belongs to the ward owners rather than to whoever tuned the ranker.
- Does reporting per-segment numbers mean the pooled number should be dropped?No. The pooled number answers a real question - what the hospital gets for 200 rows of attention - and the per-ward numbers answer a different one. Dropping either loses information. What you drop is the habit of letting one of them stand alone: the reporting contract names both, plus the coverage counts that explain any disagreement between them.
- How do you choose which segments are reported?Pick them in framing, from how care is organised and where the model's inputs differ: ward, admission route, and any group where data availability changes. Fixing the list in advance keeps a disappointing pooled result from triggering a hunt for a flattering slice, and keeps a genuine regression in a small segment from being discovered only after launch.
saying these in an interview costs you the question
- Treats the pooled win as evidence the candidate is better in each ward.
- Calls the pooled number a lie whenever the segments disagree with it.
- Explains the reversal without checking how row counts moved per ward.
- Reports per-segment precision with no coverage counts behind it.
- Picks the reporting segments after seeing which reading is flattering.