How does two-stage least squares turn an instrument into a causal estimate?
answer
- two regressions, one idea
- predict treatment first, then use the prediction
- keep only instrument-driven variation
- covariates belong in both stages
- with binary variables it is a simple ratio
basics
~20 sTwo-stage least squares first predicts the treatment from the instrument, then relates the outcome to that prediction. Only the instrument-driven part of the treatment's variation is used, and that part is assumed free of the confounding.
solid answer
~50 sStage one predicts the treatment from the instrument and any covariates, producing a fitted treatment value for each unit. Stage two relates the outcome to those fitted values rather than to the observed treatment. The point is that the fitted values contain only the variation in treatment that the instrument created, and under the exclusion and independence assumptions that variation is uncorrelated with the confounders, so the resulting coefficient is a causal effect rather than a mixture. Any covariate you use must appear in **both** stages. With a single binary instrument, a binary treatment and no covariates, the whole procedure collapses to a ratio: the difference in mean outcome between instrument groups divided by the difference in treatment take-up between them, often called the Wald estimator. That ratio makes the logic obvious. The instrument's effect on the outcome gets scaled up by how much treatment it managed to shift.
go deeper
Be ready to describe the two stages in order and say what each regression's dependent variable is. Knowing that the second stage uses predicted treatment rather than observed treatment is the key recall.
Expect to explain why using fitted values removes the confounded variation, why covariates must appear in both stages, and how the binary case reduces to a ratio you can compute from four group means.
Demonstrate that you know what precision you are buying and paying: be able to say how a weak first stage inflates uncertainty and how you would sanity-check the estimate against the raw group differences before trusting it.
Own the framing decision for stakeholders: whether the organisation should act on the scaled-up treatment effect or on the assignment effect, and be clear about which one the decision at hand actually needs.
## The idea in one sentence Treatment `D` contains two kinds of variation: the part driven by hidden confounders, which you cannot trust, and the part driven by the instrument `Z`, which you can. Two-stage least squares isolates the second part and uses only that to estimate the effect on `Y`. ## Stage one Predict the treatment from the instrument and the covariates. The output is a fitted treatment value for every unit, sometimes written `D_hat`. Two properties matter. First, `D_hat` is a function of the instrument and the covariates only, so anything correlated with the instrument's exogenous variation carries through and anything else does not. Second, the strength of this stage is the whole foundation of the estimate; if the instrument barely moves treatment, `D_hat` barely varies and the second stage has almost nothing to work with. ## Stage two Relate the outcome to the fitted treatment values, keeping the same covariates. The coefficient on `D_hat` is the instrumental-variables estimate of the treatment effect. Because `D_hat` was built only from the instrument, it is uncorrelated with the confounders that contaminate the observed `D`, which is exactly what makes the coefficient interpretable causally. The assumptions do all the work. Relevance keeps the denominator away from zero. Independence makes the instrument's variation exogenous. The exclusion restriction is what lets you attribute the instrument's association with the outcome entirely to the treatment channel, so scaling it up by the first stage yields the treatment effect rather than a jumble of channels. ## Covariates go in both stages If you condition on a covariate, it belongs in the first stage and the second stage alike. Including it in only one breaks the correspondence between the fitted values and the model you claim to be estimating. Exogenous covariates act as their own instruments, which is why the count of instruments must be at least the count of endogenous treatments for the estimate to exist at all. ## The just-identified case: a ratio you can compute by hand With one binary instrument, one binary treatment and no covariates, two-stage least squares reduces to `effect = [ E(Y | Z=1) - E(Y | Z=0) ] / [ E(D | Z=1) - E(D | Z=0) ]` The numerator is the instrument's raw effect on the outcome, the effect of being nudged whether or not you complied. The denominator is the share of units whose treatment status the instrument moved. Dividing scales a diluted effect back up to a per-treated-unit effect. A concrete version. A randomly sent invitation raises use of a feature from 10 percent to 40 percent, and raises the retention rate from 20.0 percent to 23.0 percent. The numerator is 3.0 percentage points, the denominator is 30 percentage points, and the estimate is 0.03 / 0.30 = 0.10, that is, 10 percentage points of retention per unit actually moved into using the feature. Notice how small the numerator is compared with the estimate: the ratio is doing a lot of amplification, which is why a small first stage is so dangerous. ## Two mental traps The first is thinking that the instrument replaces the treatment as the regressor of interest. Relating the outcome to the instrument directly gives the diluted effect of assignment, not the effect of treatment; you must divide by the first stage to get the treatment effect. The second is thinking the fitted values are a better measurement of treatment. They are not; they are a much worse one. `D_hat` is a deliberately impoverished version of `D` that has thrown away most of its variation, keeping only the slice with a clean provenance. Throwing information away on purpose is precisely the point, and it is also why instrumental-variables estimates are far less precise than the naive comparison. You pay for credibility in variance. ## What the estimate actually is If treatment effects vary across units, the coefficient is not the effect for everyone. It is the average effect among the units whose treatment status the instrument changed, and that subgroup is defined by the instrument you happened to pick. A different instrument for the same treatment can legitimately produce a different number without either being wrong.
- What does the coefficient from relating the outcome directly to the instrument represent?It is the effect of assignment rather than of treatment, diluted by everyone who was assigned but did not change behaviour. It answers what happens if you send the nudge, which is a legitimate quantity in its own right, but it understates the effect of the treatment itself. Dividing it by the first-stage shift in take-up rescales it into the per-unit treatment effect.
- Why is an instrumental-variables estimate so much less precise than the naive comparison?Because it uses only the sliver of treatment variation the instrument generated and discards the rest. Precision scales with how much treatment the instrument moved, so a design shifting take-up by 5 points is dramatically noisier than one shifting it by 40 points on the same sample. Credibility is bought with variance, and the exchange rate is set by the first stage.
- What happens if you include a covariate in the second stage but not the first?The fitted treatment values no longer correspond to the model you are estimating, and the estimate loses its interpretation. Any covariate you condition on must enter both stages, where it effectively serves as its own instrument. The same discipline applies to functional form: a transformation or interaction used in one stage must be reflected in the other.
saying these in an interview costs you the question
- Relates the outcome directly to the instrument and calls it the treatment effect
- Says the fitted values are a cleaner measurement of treatment
- Includes covariates in only one of the two stages
- Ignores that precision collapses when the first stage is small
- Assumes the estimate equals the effect for everyone in the sample