How do you sign the omitted-variable bias when ability is left out of a wage-on-schooling regression?
answer
- bias is a product of two signs
- effect on outcome times correlation
- short regression versus long regression
- both positive here, so upward
- more data will not save you
basics
~20 sMultiply two signs: the effect of the omitted variable on the outcome, times its correlation with the included regressor. Ability raises wages and is positively correlated with schooling, so the schooling coefficient is biased upward.
solid answer
~40 sWrite the model you believe: `wage = b0 + b1*schooling + b2*ability + e`. If you fit the short regression `wage = a0 + a1*schooling + u` instead, then `a1 = b1 + b2*d`, where `d` is the slope from regressing ability on schooling. The bias term is `b2*d`, so its sign is `sign(b2) * sign(corr(ability, schooling))`. Ability plausibly raises wages (`b2 > 0`) and more able people tend to get more schooling (`d > 0`), so the omitted-variable bias is positive: the estimated return to schooling overstates the causal return. Flip either sign and the bias flips. The key point is that this is bias, not noise, so a bigger sample does not help — it just gives a tighter interval around the wrong number.
go deeper
Know that leaving out a common cause of the treatment and the outcome biases the remaining coefficient, and that the bias does not go away with more rows of data.
Be ready to derive it: the short-regression coefficient equals the long-regression one plus the omitted variable's outcome effect times its slope on the included regressor, and to sign that product out loud.
Demonstrate that you use the direction in practice — turning a confounded estimate into a bound, judging when a proxy helps, and pairing the argument with a sensitivity analysis rather than a disclaimer.
Decide when a directional argument is enough to act on and when the question needs a different design. Be able to explain to non-specialists why a very precise number can still be systematically wrong.
## The setup Omitted-variable bias is the algebra behind the sentence *your estimate is confounded*. It tells you not only that an estimate is off, but often which way. Assume the long model is the one you believe describes the outcome: ``` wage = b0 + b1*schooling + b2*ability + e ``` Here `b1` is the quantity you want — the return to an extra year of schooling holding ability fixed. Ability is unrecorded, so you fit the short model: ``` wage = a0 + a1*schooling + u ``` ## The result The classic relationship between the two is ``` a1 = b1 + b2 * d ``` where `d` is the slope you would get from regressing the omitted variable on the included one — ability on schooling. The whole bias lives in the product `b2 * d`, and therefore ``` sign(bias) = sign(effect of omitted on outcome) * sign(association of omitted with included regressor) ``` Both factors have to be non-zero for bias to appear. An omitted variable that strongly drives the outcome but is uncorrelated with schooling costs you precision, not correctness. An omitted variable strongly correlated with schooling but irrelevant to wages costs you nothing. ## Applying it to schooling and ability Two sign judgments, each argued in one sentence: - **`b2 > 0`**: more able workers earn more at the same level of schooling. - **`d > 0`**: more able people tend to stay in education longer. Product of two positives is positive, so `a1 > b1` in expectation: the naive regression **overstates** the return to schooling. That is a useful thing to be able to say out loud, because it converts the estimate into a bound. If the naive number is 9% per year and the bias is upward, the causal return is probably below 9% — you have learned something even without measuring ability. Now flip a sign to check you understand the machinery. Imagine the omitted variable were family financial hardship, which plausibly lowers wages (`b2 < 0`) and also shortens schooling (`d > 0`). Then the bias is negative and the naive estimate understates the return. Same formula, different inputs. ## Bias is not noise This distinction is the part interviewers most often probe. Sampling error shrinks like `1/sqrt(n)`; omitted-variable bias does not shrink at all. As `n` grows, the short-regression estimator converges to `b1 + b2*d`, not to `b1`. Ten million rows give you a very tight confidence interval centred on the wrong value, and the interval's nominal coverage of the causal parameter goes to zero. Anyone who answers *collect more data* to an omitted-variable question has missed the whole point. ## What you can do about it - **Measure the confounder, even imperfectly.** A proxy removes part of the bias; a noisy proxy leaves a residue that usually points the same way as the original bias, so the adjusted estimate is a tighter bound rather than the truth. - **Sign-reason to a bound.** As above, argue the direction and report the naive number as an upper or lower bound on the causal effect. This is honest and often decision-relevant. - **Sensitivity analysis.** Ask how strongly an unmeasured variable would have to relate to both schooling and wages to move the estimate to zero. If the required strength exceeds anything plausible in the domain, the finding is robust; if a mild confounder would erase it, say so. - **Change the design.** If a source of variation in schooling exists that is unrelated to ability, that design sidesteps the problem in a way no covariate list can. ## Common traps Do not confuse the sign of the *bias* with the sign of the *coefficient*: an upward bias on a negative true effect can leave you reporting something closer to zero, or even the wrong sign entirely. Do not assume adding more covariates always shrinks the bias — the formula applies per omitted variable and the residual biases can offset or compound. And do not present the sign argument as a correction: it gives you a direction, not a magnitude, and it rests on two judgments about the world that you should state explicitly so an interviewer can challenge them.
- Does collecting ten times more data reduce omitted-variable bias?No. The estimator converges to the biased value, so more data narrows the interval around the wrong number instead of moving toward the truth. Bias and variance are different problems: sample size buys precision only. Reporting a very tight interval from a confounded regression makes the error look more credible, not less.
- What if the omitted variable affects wages strongly but is uncorrelated with schooling?Then the coefficient is unbiased. The bias term is a product, so a zero association with the included regressor zeroes it out. You still pay a price: that variable's influence sits in the residual, inflating residual variance and widening the interval on the schooling coefficient. It is a precision cost, not a validity cost.
- Your sign argument says the estimate is an upper bound. How is that useful?It converts an unusable number into a directional claim. If the naive estimate is already too small to justify a decision, an upward bias settles the question. If it is large, you can pair it with a sensitivity analysis asking how strong an unmeasured confounder would need to be to explain the whole result, and judge whether that strength is plausible.
saying these in an interview costs you the question
- Says a larger sample removes omitted-variable bias
- Confuses the sign of the bias with the sign of the coefficient
- Claims any omitted variable biases the estimate
- Treats the sign argument as a numerical correction
- Cannot state which two quantities multiply to give the bias