What assumptions must an instrumental variable satisfy to identify a causal effect?
answer
- three conditions, only one testable
- the instrument has to move the treatment
- no path to the outcome except through treatment
- as good as randomly assigned
- exclusion is an argument, not a test
basics
~20 sA valid instrument needs relevance, meaning it actually shifts the treatment, and the exclusion restriction, meaning it affects the outcome only through that treatment. It must also be as good as randomly assigned with respect to unmeasured confounders.
solid answer
~50 sThree conditions. **Relevance**: the instrument genuinely moves the treatment, so `Cov(Z, D)` is not zero once covariates are held fixed. This is the only one you can check in the data, through the first-stage regression. **Independence**: the instrument is as good as randomly assigned, unrelated to the unmeasured confounders that made the naive comparison untrustworthy. **Exclusion restriction**: the instrument affects the outcome only through the treatment. Independence and exclusion are not testable in a just-identified design; you defend them with substantive argument about how the instrument came to be. The Vietnam draft lottery is the classic illustration: the number was literally drawn at random, it raised the probability of military service, and the debated part is whether a low number could change later earnings by any route other than serving. A fourth condition, monotonicity, is needed to interpret the estimate as a complier effect.
go deeper
Be ready to name the two headline conditions in plain words: the instrument must move the treatment, and it must not touch the outcome any other way. Knowing one concrete example, such as a lottery, is enough at this level.
Expect to state all three conditions precisely, say which one the first-stage regression checks, and explain why the other two are defended by argument rather than by a test.
You should be able to take a proposed instrument from a real project and argue both sides of its exclusion restriction, naming the specific alternative pathway that would break it and what evidence would make you abandon the design.
Own the call about whether an instrumental-variables design is worth running at all against simpler alternatives, and be explicit about what the team should conclude if the exclusion restriction is merely plausible rather than defensible.
## What an instrument is for Suppose you want the causal effect of a treatment `D` on an outcome `Y`, but the people who take `D` differ from the people who do not in ways you cannot measure. Motivation, health, prior skill, whatever it is, it sits in the error term and is correlated with `D`. A straight comparison of treated and untreated units then mixes the effect you want with the effect of that hidden difference, and no amount of adding the covariates you happen to have observed fixes it. An instrumental variable `Z` is a third variable that shoves `D` around for reasons that have nothing to do with the hidden difference. If you can find one, you can throw away all the variation in `D` that you do not trust and keep only the slice of it that `Z` created. The causal effect is then read off that trustworthy slice. That plan only works if `Z` has three properties. ## 1. Relevance `Z` must actually move `D`. Formally `Cov(Z, D)` is non-zero, and conditional on whatever covariates you include, `Z` must retain real partial explanatory power for `D`. This is the one assumption the data can speak to: regress the treatment on the instrument and the covariates, and look at how strongly the instrument predicts. This regression is called the first stage. Relevance is not a yes-or-no property. An instrument can be statistically significant in the first stage and still be far too weak to support a usable estimate, because the estimate divides by the first-stage strength. A near-zero denominator makes everything unstable. That is the weak-instrument problem, and it is why the first-stage F statistic gets reported alongside every instrumental-variables estimate. ## 2. Independence, or as-good-as-random assignment `Z` must be unrelated to the unmeasured confounders and to potential outcomes. If the instrument is itself assigned by the same hidden forces that drive treatment, it inherits exactly the problem you were trying to escape. The cleanest way to earn this assumption is for `Z` to be physically randomized: a lottery draw, a randomly sent invitation, a coin flip in an administrative process. When the instrument is merely a feature of the world rather than a designed randomization, independence becomes an argument you have to make and an interviewer will press. ## 3. The exclusion restriction `Z` must affect `Y` **only** through `D`. No direct effect, no effect through a side channel. This is the assumption that fails most often and the one that cannot be checked. In a just-identified design, with one instrument for one treatment, the data are consistent with any amount of exclusion violation; nothing in the numbers reveals it. With more instruments than treatments you can run an overidentification test, but that only asks whether the instruments agree with each other, and it is uninformative if they are all invalid in the same direction. The exclusion restriction is a statement about mechanism, and it is defended with mechanism talk. Draw the path you claim does not exist and say why it does not exist. ## 4. Monotonicity, if you want to interpret the number The three conditions above buy you identification. A fourth, monotonicity, says the instrument pushes every unit in the same direction or not at all, never the opposite way. Without it the estimate is a difference of subgroup effects that need not correspond to anything. With it, the estimate is the average effect for the units whose treatment the instrument actually changed. ## A worked example The Vietnam-era draft lottery assigned each birth date a random sequence number, and low numbers meant a high chance of being drafted. Take military service as `D`, later civilian earnings as `Y`, and the lottery number as `Z`. - Relevance holds and is visible in the data: men with low numbers served at a much higher rate. - Independence holds by construction: the numbers were drawn at random, so they cannot be correlated with anyone's underlying earning capacity. - Exclusion is the arguable one. If a low number caused a man to stay in school longer to avoid the draft, then the instrument reached earnings by a route other than service, and the restriction is violated. Notice that this critique is about mechanism, not about the data. ## What interviewers are listening for Candidates who name relevance and exclusion but forget independence, or who claim exclusion is testable, get marked down. So does anyone who proposes an instrument on the grounds that it correlates strongly with the outcome. Correlation with the outcome is not a qualification; it is a warning sign, because a variable that predicts `Y` directly is exactly the variable that violates exclusion. The qualification is a credible story about why the instrument moved the treatment for reasons unrelated to everything else.
- Which of these assumptions can you actually test, and how?Only relevance. Regress the treatment on the instrument plus your covariates and inspect how strongly the instrument predicts, usually through the first-stage F statistic. Independence and the exclusion restriction are untestable in a just-identified design; they are defended by explaining how the instrument was generated and what pathways it could plausibly travel. An overidentification test with several instruments checks agreement among them, not the validity of any single one.
- Give an example where the exclusion restriction plausibly fails.Draft-lottery numbers as an instrument for military service. A low number raised the chance of serving, but it also gave men a reason to change schooling or occupational choices to avoid being drafted. Those choices affect later earnings directly, without running through service, so the instrument reaches the outcome by a second path and the estimate is contaminated.
- Why is a variable that strongly predicts the outcome a poor instrument candidate?Direct predictive power for the outcome is evidence of a pathway that does not go through the treatment, which is precisely what the exclusion restriction forbids. A good instrument should look almost irrelevant to the outcome except insofar as it changes treatment status. The screening question is not how well it predicts the outcome but whether the only way it could touch the outcome is by moving treatment.
Think of the instrument as a randomly assigned nudge on a door you cannot open yourself. It must actually move the door, and it must not reach into the room by any other opening.
saying these in an interview costs you the question
- Claims the exclusion restriction can be verified from the data
- Picks an instrument because it strongly predicts the outcome
- Names relevance and exclusion but forgets random assignment
- Treats a statistically significant first stage as sufficient relevance
- Assumes an overidentification test proves the instruments are valid