What is a backdoor path in a causal DAG, and what does the backdoor criterion require?
answer
- arrow pointing into the treatment
- non-causal association leaks backwards
- two conditions, one about descendants
- block every such path, then average
- several valid sets, one identified effect
basics
~20 sA backdoor path from treatment T to outcome Y is any path starting with an arrow into T; it carries non-causal association. A valid adjustment set contains no descendant of T and blocks every such path.
solid answer
~40 sA **backdoor path** is a path between treatment T and outcome Y whose first edge points *into* T, for example `T <- Z -> Y`. It transmits association without transmitting the effect of T, so leaving it open biases any comparison of treated and untreated units. The backdoor criterion says Z is a valid adjustment set when two conditions hold: (1) no member of Z is a descendant of T, and (2) Z blocks every backdoor path from T to Y. Then the causal effect is identified by the adjustment formula: `P(Y | do(T=t)) = sum over z of P(Y | T=t, Z=z) * P(Z=z)`, averaging over the population distribution of Z. Crucially, you never adjust for variables on the directed paths from T to Y — those carry the effect you are measuring.
go deeper
Recall the shape: a backdoor path starts with an arrow pointing into the treatment, and the classic one is a common cause of treatment and outcome. Know that you control for it and never for things the treatment causes.
State both conditions precisely and apply them to a small drawn graph, enumerating paths one by one. Be able to write the adjustment formula and say why the outer average uses the population distribution of the adjustment variables.
Show judgment when several valid sets exist: pick on measurement quality, sample size in each stratum and variance, and say out loud when no valid set exists so the effect is simply not identified.
Own the framing that identification is an assumption problem, not an estimator problem. Decide when a study should not be run on observational data at all, and how the team documents the graph the conclusion rests on.
## The problem being solved You want the causal effect of a treatment T on an outcome Y, but all you have is observational data. The raw comparison `P(Y | T=1) - P(Y | T=0)` mixes two things: association that flows because T actually changes Y, and association that flows for other structural reasons. The backdoor criterion is a graphical rule that tells you exactly which variables to condition on so that only the first kind survives. ## Front doors and back doors Think of T as a room. Association can leave T through the **front door** — along arrows pointing away from T, such as `T -> M -> Y`. That is the causal effect, and you want it intact. Association can also leave through the **back door** — along a path whose first edge points *into* T. `T <- Z -> Y` is the canonical case: some earlier variable Z influences both whether a unit is treated and what its outcome is. That path creates association between T and Y even if T does nothing at all. Any path from T to Y beginning with an arrow into T is called a **backdoor path**, however long it is; `T <- Z1 -> Z2 -> Y` is a backdoor path too. ## Blocking a path A path is **blocked** by a set Z when at least one of the following holds along it: - The path contains a chain `A -> M -> B` and M is in Z. - The path contains a fork `A <- M -> B` and M is in Z. - The path contains a collider `A -> C <- B` and **neither** C nor any descendant of C is in Z. Colliders block by default; conditioning on one opens the path. If every path of a given kind is blocked, no association flows along it. ## The criterion A set of measured variables Z satisfies the **backdoor criterion** relative to the ordered pair (T, Y) when: 1. **No node in Z is a descendant of T.** 2. **Z blocks every backdoor path from T to Y.** When both hold, Z is a valid adjustment set and the causal effect is identified from observational data by the adjustment formula: `P(Y | do(T=t)) = sum over z of P(Y | T=t, Z=z) * P(Z=z)` Read that carefully. Inside the sum you use the *conditional* distribution of Y given both T and Z, but you average over the *marginal* distribution of Z in the whole population — not its distribution among the treated. That reweighting is what turns a within-stratum comparison into a population-level causal contrast. ## Why each condition is there Condition 2 is the obvious one: open backdoor paths are exactly the routes carrying spurious association, so shut them. Condition 1 is the one candidates forget. Descendants of T sit on or downstream of the causal pathway. Adjusting for a variable on a directed path from T to Y removes the part of the effect that travels through it, so you end up estimating something smaller than the total effect — and sometimes something with no causal interpretation at all. The criterion therefore bans the entire descendant set of T from the adjustment set, which is a simple, checkable rule you can apply node by node. ## Worked enumeration Take a five-node graph with edges `Z1 -> T`, `Z1 -> Z2`, `Z2 -> Y`, `Z3 -> T`, `Z3 -> Y`, plus the causal edge `T -> Y`. Enumerate the backdoor paths from T to Y — every path whose first edge points into T: - `T <- Z1 -> Z2 -> Y` - `T <- Z3 -> Y` The second is closed only by Z3. The first is a route through two non-colliders, so putting **either** Z1 or Z2 into the adjustment set blocks it. Hence `{Z1, Z3}` is a minimal sufficient adjustment set, `{Z2, Z3}` is another minimal one, and `{Z1, Z2, Z3}` is a valid but non-minimal set. All three identify the same causal effect; they differ in statistical efficiency and in how much data you need, not in bias. ## When no valid set exists If a backdoor path runs entirely through unmeasured nodes — say an unobserved U with `U -> T` and `U -> Y` — then no set of measured variables blocks it, and the effect is **not identified** by adjustment. No estimator repairs this. The honest answers are to measure the confounder, to find a different identification strategy, or to report the effect as unidentified with a sensitivity analysis. Reaching for a more elaborate model instead is the classic mistake: the obstacle is structural, not statistical. ## Minimal versus maximal sets "Control for everything you have" is not the criterion and is not safe: it can pull in descendants of T, and it can open paths that were closed. Conversely, among the sets that *do* satisfy the criterion, a smaller one is usually preferable — fewer strata, less variance, less extrapolation — but any valid set gives an unbiased target. The choice among valid sets is an efficiency and data-availability decision; the choice of whether a set is valid at all is a graph decision.
- For the graph with edges Z1 -> T, Z1 -> Z2, Z2 -> Y, Z3 -> T, Z3 -> Y and T -> Y, give a minimal adjustment set.The backdoor paths are `T <- Z1 -> Z2 -> Y` and `T <- Z3 -> Y`. Z3 is the only node that closes the second, and either Z1 or Z2 closes the first. So `{Z1, Z3}` is minimal, `{Z2, Z3}` is equally minimal and equally valid, and `{Z1, Z2, Z3}` is valid but larger. All three identify the same effect; they differ only in variance and data requirements.
- Is adjusting for more variables always safer?No. The criterion is about which paths are open, not about how many variables you have. Extra variables can violate the no-descendant condition, and some variables open a path that was closed rather than closing one. Beyond validity, larger sets mean thinner strata, more variance and more extrapolation. Choose a set the graph justifies, then prefer a smaller valid one.
- What do you do when the only variable blocking a backdoor path is unmeasured?Then no measured set satisfies the criterion and the effect is not identified by adjustment; reporting an adjusted estimate as causal would be false. The real options are to obtain a proxy or direct measurement of that variable, adopt a different identification strategy that does not need it, or state the effect as unidentified and quantify how strong the unmeasured confounding would have to be to overturn your conclusion.
- Why does the adjustment formula average over the marginal distribution of Z rather than its distribution among the treated?Because the target is what would happen if the whole population were treated. Within each stratum of Z the comparison is confounder-free, but strata are unequally represented among the treated. Reweighting by the population distribution of Z removes that imbalance. Averaging by the treated distribution answers a different question, about the treated subpopulation rather than everyone.
Association leaving the treatment by the front door is the effect you want. A back door is a route around the outside of the house that lets association arrive at the outcome without ever passing through the treatment's own doing.
saying these in an interview costs you the question
- Calls any path from treatment to outcome a backdoor path
- Adjusts for variables on the causal pathway
- Says controlling for everything available is always safe
- Forgets the condition banning descendants of the treatment
- Claims a fancier estimator can fix an unmeasured confounder