skip to content

Why is a naive treated-minus-untreated difference biased when doctors assign each patient the arm they expect to help most?

level: seniorimportance: should knowfreq 55%

answer

  1. who gets treated is not accidental here
  2. assignment depends on the outcomes themselves
  3. split the observed gap into two terms
  4. one term is the effect, one is baseline difference
  5. the sign of the observed gap can flip

basics

~20 s

Because assignment depends on the potential outcomes themselves. The treated are the patients who would do worst untreated, so the observed gap mixes the true effect with a baseline difference between the groups and can even flip its sign.

solid answer

~50 s

In Rubin's 'perfect doctor' setup each patient receives the arm that would give them the better outcome, so the treated are exactly those who fare badly under control and the untreated are those already doing well. The observed comparison decomposes as `E[Y | D=1] - E[Y | D=0] = ATT + (E[Y(0) | D=1] - E[Y(0) | D=0])`: the effect on the treated plus a baseline gap between the two groups had neither been treated. With a perfect doctor that second term is strongly negative, so a genuinely useful treatment can look lethal. More data shrinks the standard error of a biased quantity without touching the bias. What fixes it is an assignment mechanism that does not depend on the potential outcomes — randomisation — or a defensible argument that the groups are comparable once measured characteristics are accounted for.

go deeper

for a junior

Be able to say that people who receive a treatment are often different from those who do not, so a raw comparison of the two groups can reflect that difference rather than the treatment.

for a middle

Expect to derive the split of the observed gap into an effect term and a baseline term, and to explain why the baseline term is unobservable.

for a senior

Demonstrate that you predict the direction of the bias from how units were selected, and that you reach for design changes rather than more covariates when the mechanism is outcome-dependent.

for a principal

Be prepared to set policy on when observational readouts may inform a decision at all, and to defend spending on holdouts and randomised rollouts against the cheaper biased alternative.

## Where the bias comes from Write the two group means the data actually gives you. For treated units you observe their treated potential outcome, and for untreated units their control potential outcome: ``` E[Y | D=1] = E[Y(1) | D=1] E[Y | D=0] = E[Y(0) | D=0] ``` Subtract, then add and subtract the unobservable `E[Y(0) | D=1]` — what the treated group would have shown untreated: ``` E[Y | D=1] - E[Y | D=0] = ( E[Y(1) | D=1] - E[Y(0) | D=1] ) <- ATT + ( E[Y(0) | D=1] - E[Y(0) | D=0] ) <- baseline (selection) gap ``` The first bracket is the average effect on the treated. The second compares the two groups in the same, untreated state: it is zero only if the treated and untreated groups were alike to begin with. The naive difference equals the ATT plus that gap, and nothing in the data reveals how big the gap is. ## The perfect doctor Rubin's illustration makes the gap as adversarial as it can be. Suppose a doctor knows both potential outcomes for every patient and always prescribes the arm that will serve that patient better. Take two patients, outcome measured in years of survival: - Patient A: `Y(1) = 2`, `Y(0) = 1`. Effect `+1`, so the doctor treats A. - Patient B: `Y(1) = 8`, `Y(0) = 9`. Effect `-1`, so the doctor leaves B untreated. What you observe is 2 years for the treated patient and 9 years for the untreated one. The naive difference is `2 - 9 = -7`: the treatment appears to cost seven years of life. The truth is that the ATE is `(+1 + -1) / 2 = 0`, the ATT is `+1`, and the ATU is `-1`. Check the decomposition: the baseline gap is `E[Y(0)|D=1] - E[Y(0)|D=0] = 1 - 9 = -8`, and `ATT + gap = 1 - 8 = -7`, exactly the observed number. Everything about this doctor is benign — every patient got the better arm — and yet the resulting data is maximally misleading. This is the point of the example: bias is not caused by carelessness or by bad actors. It is caused by the *dependence between the assignment and the potential outcomes*. ## Real assignment mechanisms behave the same way The perfect doctor is an extreme, but the shape recurs everywhere: - **Confounding by indication:** sicker patients are the ones who get the aggressive therapy, so treated cohorts look worse. - **Targeted retention or win-back programmes:** offers go to accounts already sliding, so the treated churn more. - **Voluntary adoption:** users who opt into a new feature are the engaged ones, so the feature looks miraculous — the bias in the other direction. In each case the sign of the bias is predictable from *why* units were selected, which is a useful diagnostic: before looking at the numbers, ask which way the selection should push, and see whether the naive estimate is suspiciously consistent with that push. ## Why more data does not help The decomposition contains no sample size. Adding units drives `E[Y|D=1] - E[Y|D=0]` toward its true population value, and that value is `ATT + gap`. Precision improves; the target being estimated is still the wrong one. A tight confidence interval around a biased quantity is a confident wrong answer, and large observational datasets make this failure worse, not better, because narrow intervals read as authority. ## What actually fixes it - **Randomised assignment.** A coin flip cannot depend on potential outcomes, so the baseline gap is zero in expectation and the naive difference becomes a valid estimate of the average effect. - **A comparability argument on observed characteristics.** If you can argue that within groups sharing the measured covariates the assignment carries no further information about the outcomes, an adjusted comparison can be defended. The strength of that argument, not the sophistication of the model, is what the claim rests on. - **Designs that borrow variation from elsewhere** — a policy change, a discontinuity, an encouragement — each replacing the missing baseline with a different, explicitly stated assumption. What never fixes it: adding covariates until the coefficient looks reasonable, or reporting the naive gap with a caveat sentence attached. ## Answering well Say the decomposition out loud, name the second term as a baseline difference rather than 'noise', give the perfect-doctor case to show the sign can flip, then state plainly that bias is a property of the assignment mechanism, so the remedy is design or an identification argument, not sample size.

  • Under what condition does the naive difference in means equal the ATT?
    When the baseline gap `E[Y(0) | D=1] - E[Y(0) | D=0]` is zero, meaning the treated and untreated groups would have shown the same average outcome had neither been treated. That is an assumption about the assignment mechanism, not something the data can confirm, though pre-treatment outcomes and covariates can make it more or less plausible.
  • Which way would you expect the bias to run for a win-back offer sent to lapsing accounts?
    Against the offer. Recipients were selected because they were already sliding, so their untreated outcome would have been worse than the comparison group's, making the baseline gap negative. A naive comparison then understates the effect and can show a helpful offer as harmful.
  • Does adding more covariates to a regression fix this?
    Only to the extent that the covariates capture the whole reason units were selected. If assignment depended on something unmeasured, such as a clinician's judgement or a customer's private intent, no amount of adjustment removes the gap. The number moves, which is easy to mistake for progress.

It is like judging a hospital by average patient outcomes when the sickest patients are all sent there. The transfer rule, not the care, drives most of the difference.

saying these in an interview costs you the question

  • Says a larger sample removes selection bias
  • Assumes treated and untreated groups started out comparable
  • Reads the raw group gap as the ATT with no argument
  • Calls the baseline difference random noise
  • Adds covariates until the estimate looks plausible

context