skip to content

Why can a treatment with an average effect of zero still help some users and harm others?

level: juniorimportance: must knowfreq 62%

answer

  1. an average hides its spread
  2. effects can point in opposite directions
  3. condition on covariates, not everyone
  4. conditional effects average up to the overall one
  5. under-50s gain, over-70s lose

basics

~20 s

An average pools opposite effects. A drug that raises recovery for patients under 50 and lowers it by a similar amount for patients over 70 posts an overall average near zero while both real effects remain untouched underneath it.

solid answer

~50 s

The average treatment effect is a mean over the whole population, so effects that point in opposite directions cancel. The quantity that survives the cancellation is the conditional average treatment effect, `CATE(x) = E[Y(1) - Y(0) | X = x]` — the mean effect among units whose covariates equal `x` — and the overall average is just `CATE(x)` averaged over the population. A drug that raises recovery rates for under-50s and lowers them for over-70s can report an average indistinguishable from zero while both group effects are large and real. So a null headline result is evidence of no effect *on average*, not of no effect. The practical consequence is that the decision is often not ship or do not ship but ship to whom. The catch is that heterogeneity must be demonstrated rather than assumed — noise manufactures apparent group differences too.

go deeper

for a junior

Be ready to say plainly that an average can cancel opposite effects, and that a null headline result means no effect on average rather than no effect at all.

for a middle

Explain the conditional average treatment effect as the mean effect among units sharing covariate values, and that averaging it over the population returns the overall effect.

for a senior

Show how you would establish heterogeneity: estimate the difference between group effects with an interval, respect the roughly fourfold power cost, and confirm on fresh data before acting.

for a principal

Own the framing that the decision is often ship-to-whom rather than ship-or-not, and that a null average concealing harm to an identifiable group can block a launch outright.

## The average is a summary, not a promise Each unit `i` has two potential outcomes: `Y_i(1)`, what it would show if treated, and `Y_i(0)`, what it would show if untreated. The individual treatment effect is `Y_i(1) - Y_i(0)`, and the average treatment effect is the mean of that difference over the population, `ATE = E[Y(1) - Y(0)]`. Averaging is lossy by construction. A mean of zero is compatible with every unit having exactly zero effect, and equally compatible with half the population gaining eight points and the other half losing eight. Nothing in the number itself distinguishes those worlds, and they call for opposite decisions. ## The conditional average treatment effect The finer-grained quantity is the conditional average treatment effect: `CATE(x) = E[Y(1) - Y(0) | X = x]` Read it as: among units whose measured characteristics equal `x`, this is the average difference the treatment makes. It is still an average — over everyone sharing those characteristics — but a much narrower one. The two quantities are tied together by iterated expectation: averaging `CATE(X)` over the distribution of `X` in the population returns the overall average effect. That identity is the whole story in one line. The overall number is a weighted blend of the conditional ones, and a blend can be zero when its ingredients are not. ## A worked case Take a trial with equal numbers of patients under 50 and over 70. Suppose the drug raises the recovery rate by 8 percentage points in the younger group and lowers it by 8 points in the older group. The overall effect is `0.5 x (+8) + 0.5 x (-8) = 0`. The trial reports nothing. Yet one group should receive the drug and the other should be protected from it, and a report that stops at the headline number gets both decisions wrong. This sign-flipping pattern — sometimes called a qualitative interaction — is rarer than the milder kind, where the effect is positive everywhere but three times larger in one group than another. Both matter, but the sign flip is the one that makes a null result actively dangerous rather than merely uninformative. ## Why individual effects stay out of reach A unit is either treated or not; you never observe both of its potential outcomes, so `Y_i(1) - Y_i(0)` is never seen for any single unit. This is the fundamental problem of causal inference. It means the conditional average is the finest thing you can actually estimate: you can compare groups of similar units, one treated and one not, but you can never compute a person's own effect and check it. Every method for heterogeneous effects is therefore a method for estimating conditional averages that are progressively more conditional, not a method for recovering individual effects. ## Establishing heterogeneity is harder than spotting it A difference between two group estimates is itself an estimate, and a noisy one. If the overall effect is measured with standard error `SE`, then each half-sample group effect has roughly `SE x sqrt(2)`, and their difference has roughly `2 x SE`. Detecting heterogeneity of the same magnitude as the main effect therefore needs on the order of four times the sample that detected the main effect. This is why apparent subgroup differences are so common and so often false. Two habits follow. First, test the difference between the group effects directly and put an interval on that difference — never conclude heterogeneity from one group being statistically significant and another not, because that comparison is between a significance verdict and a significance verdict, not between two effects. Second, treat a difference found by inspecting the data as a hypothesis, and confirm it on data that had no part in generating it. ## What this changes in practice A null average result should trigger three questions rather than a shutdown. Is there a group with a credible positive effect large enough to act on? Is there a group being harmed, and can it be identified from data available at decision time? And is the population you measured the population you will deploy to — because if the mix of groups shifts, the average effect shifts with it, even though no conditional effect changed at all. That last point is why an average effect transported to a new population so often fails to replicate: the effects did not move, the weights did. ## What an interviewer is listening for The crisp version: an average cancels opposing effects; the conditional average effect is the object of interest; a null is a statement about the average only; and heterogeneity is expensive to establish, so a claimed subgroup effect must be defended, not just displayed.

  • How would you check whether the age difference in that drug trial is real rather than noise?
    Estimate the difference between the two age-group effects directly and put a confidence interval on that difference. Comparing 'significant in one group, not in the other' is not evidence of heterogeneity. Power is the constraint: each group uses half the data and you then take a difference, so the standard error roughly doubles and detecting heterogeneity of the same size needs about four times the sample. Replication on fresh data is the strongest confirmation available.
  • If the overall effect is null but the heterogeneity is credible, what do you actually recommend?
    A targeted decision rather than a blanket one: treat the group with a positive estimated effect, withhold from the group being harmed, and verify the targeted policy prospectively. Before that, check the split is operational — the group has to be identifiable from data available at decision time. If it is not, a null average that conceals real harm to an identifiable group is a reason not to ship at all.
  • Does a large positive average effect guarantee the treatment helps most units?
    No. A large average can come from a huge benefit to a small minority while the majority is unaffected or slightly hurt. The average is a sum over effects, not a description of the typical unit. Looking at the distribution of estimated conditional effects, or at a few splits chosen in advance, tells you whether the average describes many units or a few.

A river can average waist-deep and still drown you. The mean says nothing about the channel in the middle.

saying these in an interview costs you the question

  • Treats a null average as proof the treatment does nothing
  • Reads two subgroup p-values instead of testing their difference
  • Assumes every unit receives the average effect
  • Calls any subgroup gap heterogeneity without checking noise
  • Forgets individual effects are never observed

context