What is the difference between ATE, ATT and ATU as treatment-effect estimands?
answer
- all three average the same difference
- what changes is who is averaged over
- one conditions on being treated
- one on being untreated
- the population version is the weighted blend
basics
~20 sATE averages the effect Y(1) - Y(0) over the whole population, ATT over only the units that actually got treated, and ATU over the untreated ones. ATE is their size-weighted mix, and the three can differ in sign.
solid answer
~50 sThese are estimands: the quantities you are after, named before any estimator. The ATE is `E[Y(1) - Y(0)]` over the whole population — what you would get treating everyone versus no one. The ATT is `E[Y(1) - Y(0) | D = 1]`, the same contrast restricted to units that actually received treatment. The ATU is that contrast on the untreated. They are tied together: with a treated share `p`, `ATE = p * ATT + (1 - p) * ATU`. Take a retention save-offer given only to callers flagged as churn-risk. The ATT is what the offer bought among flagged callers, the ATU is what it would buy on everyone else, and only the ATE speaks to giving it to the whole base. They can even differ in sign if the offer rescues genuinely at-risk callers but merely discounts customers who were never going to leave.
go deeper
Know that these are three averages of the same difference over different groups, and be able to say which group each one covers without hesitating.
Be ready to write each estimand in potential-outcome notation, derive the size-weighted relation between them, and give a case where they differ.
Show you pick the estimand from the decision at hand, and that you can say which one a given study design actually identifies before quoting any number.
Own the standard that every readout names its population and estimand, and be able to argue when extrapolating a targeted effect to a full rollout is defensible and when it is a business risk.
## Estimand first An **estimand** is the quantity you want to know, stated in terms of potential outcomes and a population. An **estimator** is the recipe applied to data to approximate it. Interviewers care about this order because a technically flawless estimator aimed at the wrong estimand answers a question nobody asked. With `Y(1)` the outcome under treatment, `Y(0)` the outcome under control, and `D` the treatment indicator: - **ATE — average treatment effect:** `E[Y(1) - Y(0)]`, averaged over the entire population of interest. - **ATT — average treatment effect on the treated:** `E[Y(1) - Y(0) | D = 1]`, averaged only over the units that were actually treated. - **ATU — average treatment effect on the untreated:** `E[Y(1) - Y(0) | D = 0]`, averaged over units that were not treated. Also written ATC, effect on the controls. - **CATE — conditional average treatment effect:** `E[Y(1) - Y(0) | X = x]`, the effect within a slice defined by covariates. Naming it is enough here; estimating it is a separate craft. Note what is being conditioned on. ATT and ATU condition on the *treatment status*, not on any characteristic. Both still involve unobservable terms: ATT needs `E[Y(0) | D = 1]`, the outcome the treated would have had untreated, which is never observed. ## The decomposition Let `p = P(D = 1)` be the treated share of the population. Then ``` ATE = p * ATT + (1 - p) * ATU ``` because the population splits into treated and untreated, and the ATE is the size-weighted average of the effect within each part. Two consequences follow. First, if `p` is small, the ATE is dominated by the ATU — the effect on people who never got the treatment. Second, ATT and ATE coincide when the effect is the same in both parts, which is guaranteed when the effect is constant across units and holds on average when treatment was assigned independently of the potential outcomes, as in a randomised experiment. Outside those cases, quoting one as if it were the other is a substantive error, not a labelling nicety. ## A worked case: the save-offer A retention team gives a save-offer — a discount plus a service credit — only to inbound callers whose account is flagged as churn-risk. Three different numbers answer three different business questions: - **ATT:** among flagged callers who got the offer, how much churn did the offer prevent? This is what the current programme is worth. It is the number that supports "keep running this". - **ATU:** among the unflagged callers who never received it, what would it have done? Probably much less: many of them were staying anyway, so the offer mostly hands out margin. It can be near zero or negative on a revenue outcome. - **ATE:** what would the offer do across the whole base? By the decomposition it is a blend, dragged toward the ATU because the flagged group is a small slice. It is the number that supports "give it to everyone". If the ATT is a healthy save rate and someone quotes it in a proposal to extend the offer to every caller, the proposal is built on the wrong estimand. The extension decision is about people currently untreated, so the ATU (or the ATE for a full launch versus no programme) is what matters. ## Which one your data hands you The estimand you can credibly reach depends on how units came to be treated and on who is in the study. - A randomised experiment run **within** the flagged population estimates the average effect in that population; for the programme as run, that is the ATT. - The same experiment says nothing directly about unflagged callers. Extending its number to them is an extrapolation, and should be stated as one. - Observational comparisons of treated and untreated units, once you have argued the groups are comparable, often deliver the ATT more naturally than the ATE, because the treated group is the one you have and reweighting to the full population makes stronger demands. ## Interview tells Strong answers say the population out loud: "the average effect on flagged callers who were offered the save, on 90-day churn." Weak answers say "the treatment effect" and stop. A good follow-up to expect is a policy question — you are asked to justify a rollout, and the honest reply is that the pilot identified the ATT, so a full launch needs either an assumption of a homogeneous effect or a test on the untreated segment. ## Common traps - Treating ATT and ATE as interchangeable outside randomisation. - Averaging ATT and ATU without weights: the ATE weights by group size, not equally. - Reporting an effect estimated on a targeted slice as if it applied to the whole base. - Forgetting that ATT still contains an unobservable term and therefore still needs an identification argument.
- When do the ATE and the ATT coincide?When the average effect is the same among the treated and the untreated. That is guaranteed if the effect is constant across units, and it holds on average when treatment was assigned independently of the potential outcomes, as in a randomised experiment. Under targeting or self-selection it generally fails, because the people who got treated are chosen for traits that also change how much the treatment helps.
- A pilot ran only on the targeted segment, but leadership wants to launch to everyone. What do you tell them?That the pilot identifies the effect on the targeted segment, and the launch decision hinges on the untargeted majority, whose effect was never measured. Options are to state the extrapolation explicitly, run a holdout arm on the untargeted segment, or launch with a randomised holdout so the broader effect is measured as it rolls out.
- If 10% of the population is treated, ATT is 8 and ATU is 0, what is the ATE and what does that imply?ATE = 0.1 * 8 + 0.9 * 0 = 0.8. The programme is strongly effective where it is aimed and nearly worthless elsewhere, so the population-average number understates the value of the targeted programme and correctly warns against a blanket rollout. Reporting either number alone misleads a different decision.
saying these in an interview costs you the question
- Uses ATE and ATT interchangeably outside a randomised design
- Averages ATT and ATU without weighting by group size
- Quotes a targeted pilot's effect as the effect for everyone
- Thinks ATT is directly observed because those units were treated
- Names an effect without naming the population it averages over