Whose treatment effect does an instrumental variable actually estimate?
answer
- not everyone responds to the instrument
- four groups by response pattern
- the unmoved contribute nothing
- one assumption exists to rule out backwards responders
- the effect is local to a subgroup you did not pick
basics
~20 sOnly the compliers, the units whose treatment status the instrument actually changed. That is the local average treatment effect. Units who would take the treatment regardless, or refuse it regardless, supply no identifying variation and contribute nothing to the estimate.
solid answer
~50 sSplit the population by how each unit responds to the instrument. **Compliers** take the treatment when the instrument pushes them and not otherwise. **Always-takers** take it either way, **never-takers** take it neither way, and **defiers** do the opposite of the push. Always-takers and never-takers never change treatment status, so they contribute no variation the instrument can exploit, and they drop out of the estimate entirely. Defiers would contribute with the wrong sign, which is why **monotonicity**, the assumption that nobody responds backwards, is required. What survives is the average effect among compliers, the local average treatment effect. This matters because compliers are defined by the instrument, not by anything you chose. Draft-lottery compliers are men who served because their number came up and would not have volunteered. A different instrument for the same treatment identifies a different subgroup and can legitimately give a different number.
go deeper
Know that the estimate does not apply to everyone, only to the people the instrument actually moved, and that this subgroup has a name worth remembering.
Be able to define all four response groups, explain why the unmoved ones drop out of both the numerator and the denominator, and connect monotonicity to the exclusion of defiers.
Show you would check whether the complier population matches the population the decision targets, and be able to describe compliers through observed covariates rather than implying a population-wide claim.
Own the estimand question before the analysis starts: decide which population the business decision is about, and be prepared to reject an otherwise clean design whose compliers are the wrong people.
## The four principal strata In a design with a binary instrument and a binary treatment, imagine knowing, for each unit, what it would do under both instrument values. There are exactly four possibilities. - **Compliers** take the treatment when the instrument pushes them toward it, and do not otherwise. Their treatment status is a function of the instrument. - **Always-takers** take the treatment either way. The instrument is irrelevant to them. - **Never-takers** decline the treatment either way. The instrument is equally irrelevant. - **Defiers** do the opposite of the push: treated when not encouraged, untreated when encouraged. Nobody is labelled in the data, because each unit is observed under only one instrument value. The strata are counterfactual constructs, and yet they are what determines what your number means. ## Why only compliers count The instrumental estimate is, in its simplest form, the instrument's effect on the outcome divided by its effect on treatment take-up. Always-takers and never-takers contribute zero to the denominator, because the instrument does not move them. They also contribute zero to the numerator, because if the instrument reaches the outcome only through treatment, and the instrument does not change their treatment, it cannot change their outcome either. They cancel out entirely. Defiers do move, but backwards. Their contribution enters both numerator and denominator with the opposite sign, so a mixture of compliers and defiers yields a difference of subgroup effects that need not equal any real effect for anybody, and can even fall outside the range of individual effects in the population. **Monotonicity** rules them out by assumption: the instrument pushes every unit in the same direction or leaves it alone. Under relevance, independence, exclusion and monotonicity, the estimate equals the average treatment effect among compliers. That is the local average treatment effect, local in the sense of local to a subpopulation you did not choose. ## Why this is not a technicality Suppose a randomly sent invitation to try a feature raises usage from 10 percent to 40 percent. The compliers are the 30 percent of users who tried the feature because they were asked and would not have found it otherwise. That group is systematically different from the 10 percent of always-takers who sought the feature out, who are typically the most engaged users, and different again from the 60 percent of never-takers who ignore invitations entirely. If the feature helps enthusiasts more than it helps the marginally interested, the complier effect will understate the effect on the enthusiasts, and it says nothing at all about the never-takers. The consequence for decision-making is direct. If you plan to ship the feature to everybody by default, the population you care about includes the never-takers, and your instrument told you nothing about them. If instead you plan to keep running invitation campaigns, the compliers are precisely the population your campaign will move, and the local effect is the relevant number. The estimand you need is a function of the decision, and the estimand the instrument gives is a function of the instrument. ## Characterising the compliers You cannot label individuals as compliers, but you can learn about the group. The complier share is the first stage itself, the shift in take-up the instrument produced. You can also compare the distribution of observed covariates for compliers against the full sample, which lets you say things like the compliers skew toward newer users or toward one region. Reporting that profile alongside the estimate is what separates a careful analysis from one that quietly implies a population-wide effect. ## Different instruments, different numbers Because each instrument defines its own compliers, two valid instruments for the same treatment can produce different estimates without contradiction. The draft lottery identifies the effect of military service for men who served because their number was low and would not have enlisted voluntarily; that is not the effect for career volunteers, and it was never meant to be. Reading disagreement between two instruments as proof that one is invalid is a common misreading, though systematic disagreement is a reason to look harder at the exclusion restriction of each. ## Extrapolating beyond compliers Going from the complier effect to a broader population requires an extra assumption you must state out loud, typically that treatment effects do not vary in a way correlated with compliance behaviour. Sometimes that is plausible and sometimes it is obviously false. What is never acceptable is making the leap silently by describing a local estimate as if it were the population effect. ## What interviewers listen for They want the four strata named, monotonicity connected to defiers rather than to some vague notion of consistency, and, above all, a candidate who spontaneously asks whether the compliers are the people the decision is about. Getting the algebra right and the population wrong is the characteristic failure mode here.
- Why is monotonicity needed, and what breaks without it?Monotonicity says the instrument never pushes any unit in the opposite direction, ruling out defiers. Without it, defiers enter the estimate with reversed signs, so the result becomes a difference between complier and defier effects rather than an average of anything. That quantity can lie outside the range of every individual effect in the population, which makes it uninterpretable regardless of how precise it looks.
- Can you identify which individuals in your data are compliers?No, because each unit is observed under only one instrument value, so its counterfactual behaviour is unknown. You can, however, estimate the complier share, which is exactly the first-stage shift in take-up, and compare the covariate distribution of compliers against the whole sample. Reporting that profile is how you communicate who the estimate is about without pretending to label individuals.
- Two valid instruments for the same treatment give different estimates. What is your first interpretation?That they identify effects for different complier groups, which is expected when effects vary across units rather than a contradiction. The first thing to do is describe how the two complier populations differ. Only if the gap is large and the complier profiles look similar should you start suspecting that one instrument's exclusion restriction fails.
A megaphone announcement changes only the behaviour of people who were undecided. It tells you nothing about those already committed either way, and its measured effect is theirs alone.
saying these in an interview costs you the question
- Describes the estimate as the effect for the whole population
- Cannot name the four response strata
- Thinks compliers can be identified individually in the data
- Treats monotonicity as an untestable technicality with no consequence
- Reads two instruments disagreeing as proof one is invalid