skip to content

Heterogeneity and Credibility

Whether one average effect is the whole story and whether the number survives scrutiny: subgroup variation, unmeasured confounding, placebo checks. Interviewers want an estimate defended.

on this pageshow

explore

questions

10

Why can a treatment with an average effect of zero still help some users and harm others?

level: juniorimportance: must knowfreq 62%

answer

  1. an average hides its spread
  2. effects can point in opposite directions
  3. condition on covariates, not everyone
  4. conditional effects average up to the overall one
  5. under-50s gain, over-70s lose

basics

~20 s

An average pools opposite effects. A drug that raises recovery for patients under 50 and lowers it by a similar amount for patients over 70 posts an overall average near zero while both real effects remain untouched underneath it.

solid answer

~50 s

The average treatment effect is a mean over the whole population, so effects that point in opposite directions cancel. The quantity that survives the cancellation is the conditional average treatment effect, `CATE(x) = E[Y(1) - Y(0) | X = x]` — the mean effect among units whose covariates equal `x` — and the overall average is just `CATE(x)` averaged over the population. A drug that raises recovery rates for under-50s and lowers them for over-70s can report an average indistinguishable from zero while both group effects are large and real. So a null headline result is evidence of no effect *on average*, not of no effect. The practical consequence is that the decision is often not ship or do not ship but ship to whom. The catch is that heterogeneity must be demonstrated rather than assumed — noise manufactures apparent group differences too.

go deeper

for a junior

Be ready to say plainly that an average can cancel opposite effects, and that a null headline result means no effect on average rather than no effect at all.

for a middle

Explain the conditional average treatment effect as the mean effect among units sharing covariate values, and that averaging it over the population returns the overall effect.

for a senior

Show how you would establish heterogeneity: estimate the difference between group effects with an interval, respect the roughly fourfold power cost, and confirm on fresh data before acting.

for a principal

Own the framing that the decision is often ship-to-whom rather than ship-or-not, and that a null average concealing harm to an identifiable group can block a launch outright.

## The average is a summary, not a promise Each unit `i` has two potential outcomes: `Y_i(1)`, what it would show if treated, and `Y_i(0)`, what it would show if untreated. The individual treatment effect is `Y_i(1) - Y_i(0)`, and the average treatment effect is the mean of that difference over the population, `ATE = E[Y(1) - Y(0)]`. Averaging is lossy by construction. A mean of zero is compatible with every unit having exactly zero effect, and equally compatible with half the population gaining eight points and the other half losing eight. Nothing in the number itself distinguishes those worlds, and they call for opposite decisions. ## The conditional average treatment effect The finer-grained quantity is the conditional average treatment effect: `CATE(x) = E[Y(1) - Y(0) | X = x]` Read it as: among units whose measured characteristics equal `x`, this is the average difference the treatment makes. It is still an average — over everyone sharing those characteristics — but a much narrower one. The two quantities are tied together by iterated expectation: averaging `CATE(X)` over the distribution of `X` in the population returns the overall average effect. That identity is the whole story in one line. The overall number is a weighted blend of the conditional ones, and a blend can be zero when its ingredients are not. ## A worked case Take a trial with equal numbers of patients under 50 and over 70. Suppose the drug raises the recovery rate by 8 percentage points in the younger group and lowers it by 8 points in the older group. The overall effect is `0.5 x (+8) + 0.5 x (-8) = 0`. The trial reports nothing. Yet one group should receive the drug and the other should be protected from it, and a report that stops at the headline number gets both decisions wrong. This sign-flipping pattern — sometimes called a qualitative interaction — is rarer than the milder kind, where the effect is positive everywhere but three times larger in one group than another. Both matter, but the sign flip is the one that makes a null result actively dangerous rather than merely uninformative. ## Why individual effects stay out of reach A unit is either treated or not; you never observe both of its potential outcomes, so `Y_i(1) - Y_i(0)` is never seen for any single unit. This is the fundamental problem of causal inference. It means the conditional average is the finest thing you can actually estimate: you can compare groups of similar units, one treated and one not, but you can never compute a person's own effect and check it. Every method for heterogeneous effects is therefore a method for estimating conditional averages that are progressively more conditional, not a method for recovering individual effects. ## Establishing heterogeneity is harder than spotting it A difference between two group estimates is itself an estimate, and a noisy one. If the overall effect is measured with standard error `SE`, then each half-sample group effect has roughly `SE x sqrt(2)`, and their difference has roughly `2 x SE`. Detecting heterogeneity of the same magnitude as the main effect therefore needs on the order of four times the sample that detected the main effect. This is why apparent subgroup differences are so common and so often false. Two habits follow. First, test the difference between the group effects directly and put an interval on that difference — never conclude heterogeneity from one group being statistically significant and another not, because that comparison is between a significance verdict and a significance verdict, not between two effects. Second, treat a difference found by inspecting the data as a hypothesis, and confirm it on data that had no part in generating it. ## What this changes in practice A null average result should trigger three questions rather than a shutdown. Is there a group with a credible positive effect large enough to act on? Is there a group being harmed, and can it be identified from data available at decision time? And is the population you measured the population you will deploy to — because if the mix of groups shifts, the average effect shifts with it, even though no conditional effect changed at all. That last point is why an average effect transported to a new population so often fails to replicate: the effects did not move, the weights did. ## What an interviewer is listening for The crisp version: an average cancels opposing effects; the conditional average effect is the object of interest; a null is a statement about the average only; and heterogeneity is expensive to establish, so a claimed subgroup effect must be defended, not just displayed.

  • How would you check whether the age difference in that drug trial is real rather than noise?
    Estimate the difference between the two age-group effects directly and put a confidence interval on that difference. Comparing 'significant in one group, not in the other' is not evidence of heterogeneity. Power is the constraint: each group uses half the data and you then take a difference, so the standard error roughly doubles and detecting heterogeneity of the same size needs about four times the sample. Replication on fresh data is the strongest confirmation available.
  • If the overall effect is null but the heterogeneity is credible, what do you actually recommend?
    A targeted decision rather than a blanket one: treat the group with a positive estimated effect, withhold from the group being harmed, and verify the targeted policy prospectively. Before that, check the split is operational — the group has to be identifiable from data available at decision time. If it is not, a null average that conceals real harm to an identifiable group is a reason not to ship at all.
  • Does a large positive average effect guarantee the treatment helps most units?
    No. A large average can come from a huge benefit to a small minority while the majority is unaffected or slightly hurt. The average is a sum over effects, not a description of the typical unit. Looking at the distribution of estimated conditional effects, or at a few splits chosen in advance, tells you whether the average describes many units or a few.

A river can average waist-deep and still drown you. The mean says nothing about the channel in the middle.

saying these in an interview costs you the question

  • Treats a null average as proof the treatment does nothing
  • Reads two subgroup p-values instead of testing their difference
  • Assumes every unit receives the average effect
  • Calls any subgroup gap heterogeneity without checking noise
  • Forgets individual effects are never observed

context

open as a page

What does a sensitivity analysis for unmeasured confounding tell you about a causal estimate?

level: juniorimportance: must knowfreq 55%

basics

~10 s

A sensitivity analysis says how strong an unmeasured confounder would have to be to overturn the estimate. It never shows confounding is absent; it prices how much hidden bias the finding can tolerate.

open as a page

How do negative-control outcomes and exposures expose residual confounding in an observational study?

level: seniorimportance: must knowfreq 48%

basics

~20 s

You rerun the analysis on a relationship that must be null: an outcome the treatment cannot cause, or an exposure that cannot cause the outcome, both sharing the study's confounding. Finding an effect where none can exist proves bias remains.

open as a page

How do the S-, T- and X-learner meta-learners differ when estimating conditional treatment effects?

level: middleimportance: should knowfreq 48%

basics

~20 s

The S-learner fits one outcome model with treatment as a feature and differences its predictions. The T-learner fits separate treated and control models and subtracts them. The X-learner adds a stage that models imputed per-unit effects from each side and blends them.

open as a page

Why does targeting a retention discount by churn risk differ from targeting by uplift?

level: middleimportance: should knowfreq 55%

basics

~20 s

A churn-risk model ranks who will leave; an uplift model ranks whose behaviour the discount changes. Those are different people: the highest-risk customers often leave regardless, and some contented ones cancel only because the offer reminded them the subscription exists.

open as a page

What does an E-value reported alongside an observational risk ratio actually mean?

level: middleimportance: should knowfreq 42%

basics

~20 s

The E-value is the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need with both the treatment and the outcome, beyond measured covariates, to fully explain away the observed association.

open as a page

A causal forest says a discount lifts only price-sensitive new users while the overall effect is flat — do you roll out targeted?

level: principalimportance: should knowfreq 38%

basics

~20 s

Not on the model's word alone. Confirm the segment on randomized data the forest never saw, checking that high-ranked users really show a bigger treated-versus-control gap, then weigh whether the rule is computable at decision time and whether the margin justifies the machinery.

open as a page

A lift measured on US desktop power users is proposed for a mobile-first market abroad. How do you judge whether it transports?

level: principalimportance: should knowfreq 45%

basics

~20 s

Ask which variables modify the effect and how their distribution differs in the target population. An estimate transports only if the effect modifiers are measured, overlap between populations, and the mechanism and baseline rate carry over. Otherwise reweight or re-test.

open as a page

How do you evaluate an uplift model when no customer's individual treatment effect is observed?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Score the ranking, not the units. Sort a held-out randomized sample by predicted uplift and compare treated with untreated outcomes inside each slice: a good model shows a large incremental gap at the top that shrinks further down. The Qini curve summarises that.

open as a page

In Rosenbaum bounds for a matched study, what does the sensitivity parameter Gamma represent?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Gamma is the largest factor by which two units with identical measured covariates may differ in their odds of receiving treatment. Gamma equals 1 means assignment within matched pairs is effectively random; larger values allow more hidden bias.

open as a page