Is the stratified table always the right answer when a rate reverses after stratifying?
answer
- more granular is not automatically safer
- when was the variable measured?
- did the treatment itself change it?
- common cause before treatment means stratify
- the data cannot pick between the tables
basics
~20 sNo. Stratify only when the stratifier is a common cause of group and outcome, fixed before treatment. If treatment itself changes that variable, splitting on it hides part of the effect and the pooled table is right.
solid answer
~50 sNo — "more granular is always safer" is the single most common mistake here. Stratifying is correct when the stratifier is a pre-treatment common cause: it influenced who ended up in each group and it influences the outcome on its own. Then the pooled comparison is confounded and the strata are the causal answer. But if the variable is measured after the treatment and the treatment changes it, splitting on it holds constant part of the very thing the treatment did, and the pooled comparison is the one that answers the causal question. A third case: if the stratifier has no association with group membership, the weights match and no systematic reversal is possible either way. The decisive point is that the table cannot tell you which case you are in — the causal ordering comes from domain knowledge, not from the data.
go deeper
Know that stratifying is not automatically the safe choice, and that whether a variable was recorded before or after the treatment matters. You are not expected to adjudicate the harder cases yet.
Explain the two contrasting cases concretely: a pre-treatment common cause that must be split on, and a post-treatment variable the treatment changed, which must not be. Use the timing test to tell them apart.
Demonstrate the workflow: state the estimand first, fix the adjustment set before looking at results, and justify the choice from how the data were generated rather than from which table looks better.
Own the guardrails. Decide how adjustment sets are agreed and recorded before analysis, how descriptive and causal figures are labelled differently in reporting, and how you prevent post-hoc slicing from becoming the house style.
## The instinct and why it fails Once a candidate has met Simpson's paradox, the lesson they usually take away is "always look at the subgroups". That is a rule about *looking*, and looking is free. As a rule about *which estimate to trust* it is wrong, because adjusting for the wrong variable creates bias just as reliably as failing to adjust for the right one. The data are identical in both cases; only the causal story differs, so only the causal story can decide. ## Case 1: the stratifier is a pre-treatment common cause — stratify This is the classic confounded comparison. The variable satisfies two conditions: 1. It is **fixed before** the treatment or grouping is applied — nothing about the treatment could have changed it. 2. It **influences both** which group a unit ends up in and the outcome, through separate paths. Severity of illness in an observational treatment comparison is the standard case: sicker patients get the aggressive option, and sicker patients do worse regardless. Here the pooled comparison mixes the treatment difference with a composition difference, and the within-stratum comparisons are the ones with a causal reading. If a single number is wanted, reweight the stratum-specific rates for both groups to one common distribution and label it. ## Case 2: the treatment changes the stratifier — pool Now suppose the variable being split on was measured **after** the treatment and is affected by it. Conditioning on it compares units held equal on something the treatment moved — which means the comparison has already discarded whatever the treatment accomplished through that route. The stratified numbers then understate or distort the effect, and the pooled comparison is the one that answers "what happens if we apply this treatment?" The tell is temporal: could this variable have been recorded before assignment? If it could not, be extremely reluctant to condition on it. In practice, a great deal of accidental damage in analytics comes from splitting a treatment comparison by a variable the treatment plainly influenced. ## Case 3: the stratifier is unrelated to group membership — either table If the split variable has no association with which group a unit is in — as happens by construction in a randomised experiment for any pre-treatment variable — then both groups are spread across the strata in the same proportions. Equal spread means equal weights, and equal weights mean the pooled ordering follows the stratified ordering. A reversal can still show up as small-sample noise, and it will appear now and then if you slice enough ways, but it is not systematic and it is not evidence of confounding. This is why chasing a sign flip through arbitrary cuts of an experiment is a way to fool yourself rather than a diagnostic. ## Case 4: the question you were actually asked is marginal Even in a confounded comparison, the pooled number is sometimes the right deliverable — because someone asked a descriptive question rather than a causal one. How many patients did we successfully treat last year? What is the realised rate across our current population? Those are questions about the population as it actually is, composition included. The pooled figure answers them correctly. What is illegitimate is answering a causal question with that number, or letting a reader assume the descriptive figure compares interventions. Name the estimand out loud and the ambiguity dissolves. ## The decision procedure 1. Ask what question the number must answer: causal ("if we did this") or descriptive ("what happened"). 2. If causal, place the stratifier in time relative to the treatment. Fixed before it? Candidate confounder. Determined after it and plausibly changed by it? Do not condition on it. 3. If it is a pre-treatment variable, ask whether it plausibly influences the outcome on its own as well as the group assignment. Both arrows present means stratify. 4. Decide the adjustment set **before** looking at how the numbers move. Choosing the split that produces the sign you like is not analysis. 5. Report the stratified table, and if a headline number is demanded, standardise to a stated reference composition. ## The point interviewers are testing The candidate who says "always stratify" has memorised a story. The candidate who asks when the variable was measured, whether the treatment could have moved it, and what question the number is meant to answer has understood that the reversal is a fact about the data-generating process, and that the arithmetic never adjudicates between two causal stories. Being able to say cleanly "these two tables are both correct, and which one I report depends on assumptions I hold outside this data set" is the answer.
- What information outside the table settles which analysis is correct?The causal ordering: when each variable was determined relative to the treatment, and which variables plausibly influence which. That comes from domain knowledge, protocol documents and how the data were collected — never from the numbers. Two identical tables with different backstories warrant different analyses, which is the whole lesson.
- In a randomised experiment, can splitting by an arbitrary pre-treatment variable still reverse the result?Only as noise. Randomisation makes pre-treatment variables independent of assignment in expectation, so both arms have the same spread across strata and the weights match. Slice enough ways and some cut will flip by chance, which is why unplanned subgroup slicing after seeing the headline result is unreliable rather than informative.
- When is the pooled number the number the business actually wants?When the question is descriptive rather than causal: the realised rate over the population as it currently is, for capacity planning, cost or reporting. Composition is part of the answer there. Say explicitly that it describes the current population and does not compare interventions, so nobody reads a policy recommendation into it.
saying these in an interview costs you the question
- Says always stratify because more granular is always better
- Conditions on a variable measured after the treatment
- Believes the data alone can identify the correct table
- Slices by many variables until the sign flips, then reports that
- Cannot say what question the reported number answers