In causal inference, what does conditional ignorability require of treatment assignment?
answer
- what adjustment is actually buying you
- as good as randomized, but only within strata
- condition on X, then independence
- potential outcomes independent of T given X
- untestable; argue the assignment mechanism
basics
~10 sConditional ignorability requires that within every level of the measured covariates X, which units received treatment is independent of their potential outcomes. Treatment must be as good as randomly assigned inside each X stratum.
solid answer
~40 sFormally the assumption is `(Y(1), Y(0)) independent of T given X`: once you condition on the covariates you actually measured and adjust for, assignment carries no information about how a unit would have responded under either condition. That is exactly what licenses reading a within-stratum difference in observed outcomes as a causal effect rather than a selection artefact. The real work is arguing what X must contain. For a marketing email that people choose to open, X would have to hold everything driving both opening and buying: prior purchase history, engagement, tenure, and purchase intent at that moment, and the last of those is almost never recorded. The assumption is untestable from the data itself, so you defend it with an argument about the assignment mechanism, not with a statistic.
go deeper
Be able to say in plain words that adjustment assumes the treated and untreated look alike once you account for the measured variables, and that this is an assumption rather than something the data proves.
Expect to write the statement, explain that treatment may depend on X but not on the potential outcomes given X, and explain why balance tables cannot verify it.
Show that you start from the assignment mechanism: who decides treatment, on what information, and which of those drivers you actually recorded. Name the missing ones and the direction they bias.
Own the call on whether an untestable assumption is defensible enough to publish a causal claim, and set the team norm that the adjustment set is argued and written down before the analysis runs.
## What the assumption says Every unit has two potential outcomes: `Y(1)`, the outcome it would show under treatment, and `Y(0)`, the outcome it would show under control. Conditional ignorability, also called unconfoundedness, conditional exchangeability, or selection on observables, is the statement `(Y(1), Y(0)) independent of T | X` Read it as: among units that share the same values of the measured covariates X, the ones who ended up treated are not systematically the ones who would have done better (or worse) anyway. Assignment inside an X stratum is as good as a coin flip with respect to potential outcomes. Note what it does **not** say. It does not require T to be independent of X. Treatment probability is allowed to depend on X as strongly as you like, which is the whole point of an observational study: older users get the offer more often, sicker patients get the drug more often. What must vanish is any *residual* dependence between assignment and the outcomes a unit would have had. ## Why it is the load-bearing assumption Without it, the observed difference in a stratum is a mixture of two things: the causal effect, and the difference in how the two groups would have done regardless. Conditional ignorability is what lets you write `E[Y | T=1, X=x] = E[Y(1) | X=x]` and the same for control, so the within-stratum contrast is a genuine causal contrast for units with `X = x`. Every adjustment method in the observational toolbox, whatever its machinery, rests on this same claim; they differ in how they average over X, not in whether they need it. ## Arguing it in practice Because the assumption concerns unobserved potential outcomes, no diagnostic can confirm it. Balance on X after adjustment shows only that you succeeded at what you attempted; it is silent about anything you never measured. The defensible workflow is: 1. Describe the assignment mechanism in words. Who decides who gets treated, and what do they know when they decide? 2. Enumerate the drivers of that decision that also move the outcome. Those are the variables X must contain. 3. Check which of them you actually have, at the right time and with acceptable measurement quality. 4. Say plainly which drivers are missing and in which direction they would bias the estimate. The marketing-email case is a clean illustration. If treatment is defined as opening the email, the people who open are the people already leaning toward buying. Writing `(Y(1), Y(0)) independent of T | X` out loud forces the question of what X would need: prior purchases, recency, browsing in the last hour, lifecycle stage, and, ideally, the intent that made them open in the first place. That last one is latent, so the honest conclusion is that opening is a poor treatment definition and the assignment (being sent the email) is the better one. ## Choosing X badly More covariates is not automatically safer. - **Post-treatment variables.** A variable measured after treatment and affected by it, such as a mediator, absorbs part of the effect you are trying to estimate and can introduce bias where none existed. X should be pre-treatment. - **Proxies with error.** A noisy proxy for a confounder removes only part of the confounding; residual bias survives adjustment and is easy to mistake for a clean result. - **Timing.** A covariate recorded at the wrong moment (after exposure began) silently becomes a post-treatment variable. ## Weaker variants worth knowing Only mean independence, `E[Y(t) | T, X] = E[Y(t) | X]`, is needed for estimating average contrasts; full distributional independence is more than most estimators use. This matters when someone objects that full independence is implausible: the estimator may need less. ## Sensible interview register Strong answers name the assumption, write it, then immediately move to the mechanism argument for the specific setting at hand. Weak answers treat it as a box ticked by throwing every available column into a model, or claim a balance table proved it.
- Should you condition on every variable you have, to make ignorability more plausible?No. Covariates must be pre-treatment. A variable measured after treatment and affected by it, such as a mediator, absorbs part of the causal effect and can introduce bias rather than remove it. The adjustment set should come from reasoning about the assignment mechanism, not from whatever columns happen to be in the table.
- What is the difference between exchangeability and conditional exchangeability?Marginal exchangeability says the treated and control groups are comparable overall, which is what randomization delivers by design. Conditional exchangeability claims comparability only within levels of X, so an observational analysis must adjust before it compares anything. The first is bought by the design; the second has to be argued from subject knowledge.
- Two analysts adjust for different covariate sets and get different estimates. What does that tell you?That at least one adjustment set fails to make treatment ignorable, though not which. Specification disagreement is evidence about fragility, not a licence to average the answers. The useful response is to work out which variables drive the gap and whether the mechanism argument supports including them, rather than reporting the specification you prefer.
It is like claiming that within each classroom the teacher handed out the extra tutoring by lottery. Across the whole school the tutored students look different, but inside one classroom the pick was blind to who would have done well anyway.
saying these in an interview costs you the question
- Claims covariate balance after adjustment proves ignorability
- Says adjusting for more variables always reduces bias
- Confuses independence of T and X with independence of T and potential outcomes
- Believes a large enough sample fixes unmeasured confounding
- Adjusts for variables measured after treatment