skip to content

What does a first-stage F statistic of 4 imply about a 2SLS estimate?

level: seniorimportance: should knowfreq 48%

answer

  1. the estimate divides by the first stage
  2. small denominator, unstable ratio
  3. biased back toward the naive comparison
  4. the rule of thumb threshold is ten
  5. more weak instruments makes it worse

basics

~20 s

A first-stage F of 4 signals a weak instrument. The two-stage least squares estimate is biased back toward the confounded naive comparison, standard errors inflate, and conventional confidence intervals undercover. Around 10 is the classic minimum.

solid answer

~50 s

The first-stage F measures how strongly the instrument moves the treatment once covariates are accounted for. A value of 4 is weak. Three consequences follow. First, in finite samples the estimate is biased toward the very association you were trying to escape, so a weak-instrument result that looks reassuringly close to the naive comparison is evidence of failure rather than confirmation. Second, the estimate divides by a small, noisy first stage, so the sampling distribution has heavy tails and standard errors blow up. Third, conventional confidence intervals undercover, meaning the true 95 percent interval is wider than the one you printed. The rule of thumb is a first-stage F above 10, and more recent work argues even that is far too lenient for honest inference. Remedies: find a stronger instrument, drop weak surplus instruments, and report weak-instrument-robust inference such as an Anderson-Rubin confidence set.

go deeper

for a junior

Know that the strength of the first stage must be reported and that a low F means the instrument barely moves the treatment, so the resulting number cannot be trusted.

for a middle

Be able to explain mechanically why a small first stage destabilises the estimate, given that the estimate divides by it, and to recall the threshold of roughly 10 as a rule of thumb.

for a senior

Show you would diagnose strength before debating validity, recognise that agreement with the naive estimate under weak identification is a symptom rather than a check, and know that robust inference exists for this case.

for a principal

Own the decision of whether to publish a weakly identified result at all, what caveat framing goes to decision-makers, and when to redirect the investment toward a design that can actually answer the question.

## What the first-stage F measures The first stage of an instrumental-variables design predicts the treatment from the instrument and the covariates. The first-stage F statistic tests whether the excluded instruments add explanatory power for the treatment beyond the covariates alone. It is a measure of how hard the instrument pushes, and because the instrumental-variables estimate divides by that push, it controls everything about the estimate's behaviour. An F of 4 says the push is small relative to the noise. An F of 40 says it is solid. That gap changes the estimate qualitatively, not just cosmetically. ## Consequence one: bias toward the confounded estimate The headline claim for instrumental variables is consistency: with a valid instrument and enough data, the estimate converges to the causal effect while the naive comparison does not. In finite samples that promise degrades smoothly with instrument strength. When the instrument is weak, the estimate is pulled back toward the plain association between treatment and outcome, confounding and all. A common approximation, used to motivate the rule of thumb, is that the finite-sample bias of the instrumental estimate relative to the naive one is on the order of 1/F, which is where the threshold of about 10 and its promise of roughly ten percent of the naive bias come from. The practical implication is uncomfortable. If your instrument is weak and your instrumental estimate lands close to the naive estimate, you cannot read that as two methods agreeing. It is exactly what a broken design produces. ## Consequence two: heavy tails and inflated uncertainty The estimate is a ratio whose denominator is an estimated first-stage effect. When that denominator is small and noisy, occasional draws land near zero and the ratio explodes. The sampling distribution stops resembling a normal curve, and in the just-identified case it does not even have a finite mean. Standard errors computed as if everything were well behaved understate the true spread, so nominal 95 percent intervals cover the truth substantially less than 95 percent of the time. This is the part candidates most often miss: weak instruments do not merely make you imprecise, they make your stated precision dishonest. ## Consequence three: adding instruments makes it worse The intuitive fix, throwing in more instruments to strengthen identification, backfires. Each additional weak instrument adds a little more spurious first-stage fit, and the bias toward the naive estimate grows with the number of instruments. The classic cautionary case is the use of quarter of birth, interacted with compulsory-schooling laws and birth cohort, as an instrument for years of education. The interaction set runs to well over a hundred instruments, each individually feeble. Critics showed that with that many weak instruments the estimates drifted toward the ordinary association, and that deliberately fabricated instruments with no real relationship to schooling reproduced similar-looking results. The lesson was not that quarter of birth is a stupid idea; it was that instrument count and instrument strength interact badly. ## What to do about it 1. **Report the first stage, always.** An instrumental-variables result presented without its first-stage strength is not reviewable, and any competent interviewer will ask for it first. 2. **Find a stronger instrument.** No statistical technique substitutes for an instrument that actually moves treatment. In a design you control, this means engineering a bigger shift: a more persuasive nudge, a larger eligibility change, a longer exposure. 3. **Prune, do not pile on.** Prefer a single strong instrument to a bag of weak ones. Just-identified estimation has better bias behaviour than heavily over-identified estimation with weak instruments. 4. **Use weak-instrument-robust inference.** The Anderson-Rubin approach constructs a confidence set whose coverage is valid regardless of instrument strength, at the cost of being wide, and sometimes unbounded, when the instrument is weak. An unbounded interval is not a failure of the method; it is an honest report that the data cannot pin the effect down. 5. **Be willing to abandon the design.** Sometimes the correct answer is that this instrument cannot answer this question, and the effort belongs in getting a real experiment or a different design instead. ## The F of 4 versus F of 40 contrast With an F of 40 you have a defensible estimate whose remaining risk lives in the exclusion restriction, an argument about mechanism. With an F of 4, the argument about mechanism is beside the point, because even a perfectly valid instrument would produce an unreliable number. Diagnose strength before you spend energy debating validity.

  • Your weak-instrument estimate happens to match the naive comparison closely. Is that reassuring?
    No, it is a warning. Weak instruments bias the estimate toward the confounded association in finite samples, so agreement is the expected symptom of failure rather than independent corroboration. Convergence between the two is only meaningful when the first stage is strong. With an F of 4 you should report the weakness prominently and treat the apparent agreement as uninformative.
  • Why does adding more instruments not fix a weak first stage?
    Each extra weak instrument contributes a little overfitting to the first stage, and the finite-sample bias toward the naive estimate grows with the number of instruments. A large set of individually feeble instruments can produce results indistinguishable from what randomly fabricated instruments would give. One genuinely strong instrument beats a hundred weak ones.
  • What can you report when the instrument is weak but you must publish something?
    Report the effect of assignment itself, which needs no first-stage rescaling and is estimated with ordinary precision, alongside a weak-instrument-robust confidence set for the treatment effect. State the first-stage F explicitly. If the robust set is uninformatively wide or unbounded, say so; an honest inability to pin down the effect is a legitimate finding.

It is like measuring a room by dividing a shadow's length by the sun's angle. When the angle is tiny and poorly measured, small errors in it swing the answer wildly.

saying these in an interview costs you the question

  • Treats a significant first stage as strong enough
  • Reads agreement with the naive estimate as validation
  • Adds more weak instruments to strengthen identification
  • Reports conventional confidence intervals under weak identification
  • Never reports the first-stage statistic at all

context