A treatment beats its rival on small kidney stones and on large ones but loses overall, so which result do you report?
answer
- assignment to treatment was not random
- severity drives both assignment and outcome
- stone size is fixed before treatment
- stratify, then standardise to one case mix
- randomisation would balance the strata
basics
~20 sReport the stone-size-specific results. Stone size influenced both which treatment a patient received and how likely success was, so the pooled comparison blends the treatment effect with differences in case severity. Only the size-specific comparison supports a causal reading.
solid answer
~50 sReport the stratified result, and say why. In the well-known kidney-stone comparison, the more invasive treatment won among small stones and won again among large stones, yet lost on the pooled success rate — because surgeons preferentially gave it to the harder large-stone cases, which fail more often whatever you do. Stone size is a genuine confounder: it is fixed before treatment, it drove assignment, and it drives the outcome independently. The pooled number therefore answers "how did patients who happened to receive each treatment do?", which folds in case severity, rather than "what would happen if we assigned this treatment?". If a single headline figure is required, standardise: apply each treatment's size-specific success rates to one common distribution of stone sizes and report that, naming the reference distribution. And say what would have prevented the problem: randomised assignment, which breaks the link between severity and treatment.
go deeper
Recognise the shape and say the safe thing: the two treatments were given to different kinds of patient, so the overall rates are not comparable. Naming the reversal is enough at this level.
Explain why stone size qualifies as a confounder — fixed before treatment, influences who receives which treatment, influences success on its own — and show how that produces unequal weights in the pooled rate.
Give the decision and the caveats: report stratum-specific results, standardise if one figure is demanded, name the reference case mix, and state what residual confounding you have not handled.
Frame it as a design problem. Argue for randomised or protocol-driven assignment so severity cannot track treatment, and set the standard for how observational comparisons may be published inside the organisation.
## The scenario A study compares two treatments for kidney stones. Treatment A has a higher success rate among patients with **small** stones. Treatment A also has a higher success rate among patients with **large** stones. Pooled across all patients, treatment **B** has the higher success rate. Every one of those statements is true of the same data set at the same time. ## Why the reversal appeared Assignment was not random. The more invasive treatment was preferentially given to the harder cases — patients with large stones — while the less invasive one was used mostly on small stones. Two facts then collide: - **Large stones fail more often under either treatment.** Severity drives the outcome directly. - **Large-stone patients are concentrated in treatment A's pool.** Severity drives the assignment. A variable that drives both the treatment received and the outcome is a **confounder**: a common cause sitting behind both. Treatment A's pooled success rate is therefore a weighted average dominated by its hard cases, and treatment B's is dominated by its easy ones. The pooled comparison is contaminated by the difference in case mix, and that contamination is large enough to reverse the sign. ## Which number answers which question This is the part worth being precise about, because it is the whole answer. - The **pooled rate** answers a descriptive question: *among patients who actually received each treatment under this hospital's referral habits, what fraction succeeded?* That is a real, useful number for planning — but it describes a population, not a treatment. - The **size-specific rates** answer the causal question: *for a patient with a stone of this size, which treatment is more likely to succeed?* Because the comparison is now made among patients alike on the variable that determined assignment, the difference is attributable to the treatment rather than to who got it. Since the interview question is which treatment to *choose*, the stratified table is the one to report — and if you are choosing for a specific patient, you report the row that matches that patient's stone size, not any aggregate at all. ## Producing one honest headline number Clinicians and executives will still ask for a single figure. Do not go back to the raw pooled rate. Instead, **standardise**: take each treatment's success rate within each stone-size stratum, and weight those rates by one common stone-size distribution applied to both treatments — the overall patient population, say, or the population you intend to treat next year. Both treatments are then evaluated on the same hypothetical case mix, and the standardised figures agree in direction with the stratum-specific results. Two disciplines make this credible: state the reference distribution explicitly, and publish the stratum-specific rates next to the standardised number so a reader can audit the weighting. ## Assumptions the stratified answer still rests on The stratified table is the better answer, not an unconditionally correct one. Conditioning on stone size removes the confounding *by stone size only*. It leaves: - **Residual confounding.** If surgeons also chose on comorbidity, age or stone position, and those affect success, the within-stratum comparison remains biased. The honest phrasing is that stratification handles the confounders you measured and split on, and nothing else. - **Coarse strata.** "Small" and "large" are a crude split of a continuous variable. If severity varies substantially inside a stratum and is unevenly distributed between arms within it, some bias survives. - **Thin cells.** Splitting finely enough to remove bias eventually leaves cells too small to estimate anything precisely. This is the real tension in practice, and naming it is a senior-level signal. ## What randomisation would have changed Had treatment been assigned at random, stone size would be balanced between arms in expectation. Balanced composition means the pooled comparison uses effectively the same weights for both arms, so the pooled and stratified answers agree up to sampling noise, and the reversal cannot appear systematically. This is precisely what randomisation buys, and saying so demonstrates that you understand the paradox as a property of *assignment*, not of arithmetic. ## Common wrong answers Reporting the pooled figure "because it uses all the data" misses that both tables use all the data. Claiming the pooled number is more reliable because its sample is larger confuses precision with bias — the pooled estimate is a precise estimate of a confounded quantity. Dismissing the stratified split as subgroup fishing is wrong here too: stone size was specified in advance as a severity marker and determined assignment, which is the opposite of a post-hoc slice hunted for significance. ## The rule to carry away When a comparison reverses on stratifying, ask one question about the stratifier: was it fixed before the treatment, and does it influence both who got the treatment and the outcome? If yes, it is a confounder and the stratified table is your answer. If not, the reversal is telling you something else, and the pooled table may well be the number you want.
- What would have changed if treatment had been randomly assigned?Randomisation balances stone size across arms in expectation, so both arms carry effectively the same case mix. With equal weights, the pooled comparison agrees with the stratum-specific ones up to sampling noise, and a systematic reversal cannot arise. That is exactly the property randomisation is bought for.
- You are pushed for one headline success rate per treatment. What do you publish?A standardised rate: apply each treatment's stratum-specific success rates to one common stone-size distribution, used identically for both treatments, and name that reference distribution in the caption. Publish the stratum-specific rates beside it. Never resolve the demand by falling back to the raw pooled rate.
- What would make you distrust the stratified result as well?Residual confounding. Stratifying on stone size handles stone size and nothing else, so if surgeons also selected on comorbidity or age and those affect success, bias survives within strata. Coarse strata hide severity variation too. Say what you adjusted for, and treat everything unmeasured as an open assumption.
saying these in an interview costs you the question
- Reports the pooled rate because it uses all the data
- Says the pooled table is more reliable because its sample is larger
- Dismisses the size split as post-hoc subgroup fishing
- Averages the two size-specific rates without weighting them
- Claims stratifying always fixes confounding completely