skip to content

What is the fundamental problem of causal inference in the potential outcomes framework?

level: middleimportance: must knowfreq 78%

answer

  1. count how many outcomes each unit has
  2. you only ever see one of them
  3. the unobserved one is a counterfactual
  4. half the table is blank: missing data
  5. so only averages are recoverable

basics

~20 s

Each unit has two potential outcomes, one under treatment and one under control, but only the one matching the arm it actually received is ever observed. The other stays missing, so an individual causal effect is never measured directly.

solid answer

~50 s

Every unit has two potential outcomes: `Y(1)`, the outcome it would show under treatment, and `Y(0)`, the outcome it would show under control. The causal effect for that unit is `Y(1) - Y(0)`. But the unit is either treated or not, so exactly one of the two is realised and the other is a counterfactual that is never observed: take a headache pill and you see your pain score an hour later with the pill, never the pain score you would have had that same hour without it. That is the fundamental problem. The individual effect is not merely hard to measure, it is unobservable in principle. Causal inference works around it by giving up on individuals and targeting averages, and by relying on an assignment mechanism (randomisation, ideally) that lets one group's observed outcomes stand in for the other group's missing ones.

go deeper

for a junior

Be able to state that each unit has an outcome under treatment and one under control, and that only one is ever seen. Knowing the phrase 'you cannot observe the counterfactual' and why before-after is not it is enough here.

for a middle

Expect to write the observation rule Y = D*Y(1) + (1-D)*Y(0), explain why the unit-level effect is undefined in data, and connect that to why the framework targets averages instead.

for a senior

Show that you use the framing in practice: state the missing-data view, and say which feature of a study design makes the missing averages recoverable before you discuss any model or adjustment.

for a principal

Be ready to argue why teams that skip this framing ship confident but wrong readouts, and to set a standard where every causal claim names the counterfactual it is asserting rather than the model that produced the number.

## The setup The potential outcomes framework starts by attaching two numbers to every unit (a person, a session, a store) instead of one. For unit `i` and a binary treatment: - `Y_i(1)` is the outcome that unit would show **if treated**. - `Y_i(0)` is the outcome that same unit would show **if not treated**. Both numbers are treated as fixed properties of the unit that exist before anyone decides who gets what. The **individual treatment effect** is their difference: ``` tau_i = Y_i(1) - Y_i(0) ``` This is the quantity a causal question is really about: what did the treatment *do to this unit*. ## The observation rule Let `D_i` be 1 if unit `i` was treated and 0 otherwise. What the data records is ``` Y_i = D_i * Y_i(1) + (1 - D_i) * Y_i(0) ``` One of the two potential outcomes is switched on by the assignment, and the other is switched off forever. You never see both terms for the same unit, so `tau_i` can never be computed for any single unit — not with a bigger sample, not with better instruments, not with more precise measurement. That impossibility is the **fundamental problem of causal inference**. ## A headache pill, one person You have a headache, take a pill, and an hour later rate your pain 3 out of 10. That number is `Y(1)`. The question "did the pill help?" asks about `Y(1) - Y(0)`, where `Y(0)` is the pain you would have reported at that same hour, with that same headache, had you not taken the pill. That hour happened once. `Y(0)` is not an unrecorded number sitting somewhere waiting to be found; there is no world in which it was ever written down. Repeating the experiment tomorrow gives you a different headache, a different unit. Notice which comparisons do **not** rescue you. Pain before the pill is not `Y(0)`: it is the outcome at a different time, and headaches fade on their own. Your friend who skipped the pill is not `Y(0)` either: that is a different person, with a different headache. ## Why "missing data" is more than a metaphor Write the ideal table: one row per unit, one column for `Y(1)`, one column for `Y(0)`. Causal questions are trivial arithmetic on that table. Real data is that same table with exactly one of the two cells blanked out in every row — a missing-data pattern, and a severe one, because half the cells are missing and each row is missing at least one. The mechanism that blanks the cells is the **assignment mechanism**: whatever process decided who got treated. This is the pivot of the whole framework. What you are willing to assume about the assignment mechanism is exactly what you are willing to assume about the missingness, and it is what determines whether the blanks can be filled in on average. If treatment was assigned by a coin flip, the blanked cells are missing in a way unrelated to their values, and the untreated group's observed outcomes are a fair stand-in for the treated group's missing `Y(0)`. If treatment was chosen by the units themselves, or by someone who could guess the outcomes, the blanks are missing precisely because of what they contain, and a naive comparison measures the selection as much as the treatment. ## What survives Individual effects are gone, but averages are not. Population quantities like `E[Y(1) - Y(0)]` are differences of two averages, and an average over a group can be estimated from a *different* set of units than the one that supplied the other average — as long as the two groups are comparable. That is the trade the field makes: abandon the unit-level question, keep the population-level one, and spend all the effort on the comparability argument. Every design you have heard of — randomised experiments, matching, adjustment, natural experiments — is a strategy for making the missing averages recoverable, not for recovering individual counterfactuals. ## Common traps - **"With enough data we could see the individual effect."** Sample size buys precision on averages. It never gives a unit a second history. - **Before-versus-after as the counterfactual.** This substitutes the unit's earlier outcome for `Y(0)`, which is only valid if nothing else changed over the interval — an assumption, not an observation. - **Treating the observed `Y` as if it were both potential outcomes.** Once you write `Y = Y(1)` for a treated unit, remember `Y(0)` for that unit still exists as a quantity and is simply unknown. ## What to say in an interview Define the two potential outcomes, write the observation rule, state that exactly one is realised so the unit-level effect is undefined in the data, then pivot: this makes causal inference a missing-data problem whose missingness mechanism is the assignment mechanism, which is why we estimate averages and why how units came to be treated matters more than which model we fit.

  • If unit-level effects are never observable, how can any average effect be estimated?
    Because an average over a group can be assembled from different units than the ones supplying the other average. If assignment is a coin flip, the untreated group's mean outcome is a fair estimate of the mean `Y(0)` the treated group would have had, so the difference in group means estimates the average effect even though no single unit's effect is known.
  • Why do you call this a missing-data problem rather than just a measurement limitation?
    Because the ideal data is a complete table of `Y(1)` and `Y(0)` per unit, and the observed data is that table with one cell per row blanked. The process doing the blanking is the assignment mechanism, so assumptions about assignment are literally assumptions about the missingness pattern — which is what makes the blanks fillable on average or not.
  • Does a crossover design, where the same person is treated in one period and untreated in another, solve the problem?
    No, it relabels it. The unit becomes person-period, and each person-period still has only one realised outcome. It works only under added assumptions: no carryover from the first period, and no drift in the person or the environment between periods. You are still substituting an assumption for the missing outcome.

It is like judging a fork in the road by driving one branch. You can time the route you took, but the trip you did not take was never driven, so its time was never recorded anywhere.

saying these in an interview costs you the question

  • Claims a large enough sample reveals individual treatment effects
  • Says the outcome before treatment is that unit's control potential outcome
  • Treats the observed outcome as if both potential outcomes were seen
  • Confuses a unit-level effect with a group average effect
  • Thinks better measurement can recover the counterfactual

context