skip to content

Regression Adjustment

Estimating an effect by modelling the outcome on treatment plus confounders, then averaging predictions over the covariate distribution. Interviewers probe omitted variables and the table 2 fallacy.

on this pageshow

questions

5

What does adding a confounder to a regression do to the treatment coefficient?

level: juniorimportance: must knowfreq 76%

answer

  1. compare like with like
  2. common cause of both
  3. within-strata comparison
  4. measured confounders only
  5. assumptions, not magic

basics

~20 s

Adding a measured confounder turns the treatment coefficient from a raw comparison into a within-strata one: it now compares treated and untreated units that share the same covariate value, removing the part of the gap that covariate explained.

solid answer

~40 s

A confounder is a common cause of both the treatment and the outcome, so the raw difference in outcomes mixes the treatment effect with the fact that treated units differed to begin with. Putting the confounder in the model makes the treatment coefficient an adjusted association: roughly the treated-versus-untreated contrast among units with the same value of that covariate, pooled across values. That adjusted number is only a causal effect under real assumptions: every confounder is measured and included, both treatment levels actually occur at every covariate value, and the model's functional form is right. Adjustment does nothing about causes you never measured, and it only partly removes a confounder you measured badly. So I would report it as an estimate under stated assumptions, not as the effect.

go deeper

for a junior

Be ready to define a confounder as a common cause of treatment and outcome, and to say plainly that adjusting compares treated and untreated units with the same covariate value.

for a middle

Explain the mechanics: the coefficient becomes a within-strata contrast, and it is causal only under no unmeasured confounding, positivity and a correct functional form. Name those conditions unprompted.

for a senior

Show judgment about what survives adjustment in real data: proxies measured with error, strata with only one arm present, and the habit of arguing from causal structure rather than from which variables move the estimate.

for a principal

Own the framing for the organisation: state which assumptions the estimate rests on, which unmeasured factor would overturn it, and when an observational adjusted number is simply not strong enough to base a decision on.

## The problem adjustment is trying to solve Suppose you compare an outcome `Y` between units that got a treatment `A = 1` and units that did not, `A = 0`. The raw difference in means is a *description* of the two groups. It answers the causal question only if the two groups were otherwise comparable. When some variable `X` influences both who gets treated and what the outcome is, the groups are not comparable, and `X` is called a **confounder** — a common cause of treatment and outcome. Concretely: a support team offers a premium onboarding call to accounts that already look promising. Accounts that got the call renew more often. Part of that gap is the call; part of it is that promising accounts renew more anyway. Account size is a common cause of both. ## What the regression actually does Fit `Y = b0 + b1*A + b2*X + e`. The coefficient `b1` is no longer the raw group gap. It is the treatment-versus-control contrast *holding X fixed* — in the simplest linear case, a weighted pooling of the treated-minus-untreated differences computed within levels of `X`. Everything that `X` explained about the outcome has been taken out of the comparison, so treated and untreated units are compared as if they had the same account size. Two consequences follow immediately. 1. The adjusted coefficient usually moves relative to the unadjusted one, and the direction of that move is informative but not proof of anything. 2. The coefficient's meaning has changed. It is now a conditional quantity — an effect *within* levels of `X` — not a population-averaged one. On a linear model with no interaction those coincide; on other scales they need not. ## What has to be true for it to be causal Adjustment is not magic; it buys a causal reading only under assumptions you should be able to name: - **Conditional exchangeability (no unmeasured confounding).** Within levels of the covariates you included, treatment is as good as randomly assigned. This is untestable from the data. - **Positivity.** At every covariate value that occurs, both treated and untreated units actually exist. If no small account ever got the call, no amount of modelling recovers what would have happened if one had; the model silently extrapolates instead. - **Consistency.** The treatment is well-defined enough that the observed outcome under `A = 1` is the outcome you would want to predict for anyone assigned to treatment. - **Correct functional form.** If the real relationship between `X` and `Y` is curved, or the effect of treatment differs by `X`, a straight additive model leaves some of the confounding behind. This is sometimes called residual confounding due to misspecification. ## What adjustment does not fix - **Unmeasured common causes.** If the team also picked accounts by a gut sense of engagement that is nowhere in the data, that channel is untouched. - **Mismeasured confounders.** Adjusting for a noisy proxy removes only part of the confounding; a coarse three-bucket version of a continuous confounder leaves the within-bucket imbalance in place. - **The direction of the remaining bias.** Nothing in the output tells you which way you are still wrong; you have to reason about that separately. ## Reading the output honestly The interviewer is usually listening for three things. First, that you say *common cause* rather than *anything correlated with the outcome* — the reason a variable belongs in the model is its causal role, and deciding which variables qualify is a structural question you answer before fitting, not by watching which ones move the estimate. Second, that you do not treat a small change in the coefficient as evidence of no confounding: change-in-estimate is a diagnostic, not a proof, and two biases can offset. Third, that you separate the statistical statement (this is the adjusted association) from the causal claim (this is the effect, *if* the assumptions hold). A good closing sentence in an interview is: after adjustment I have a comparison of like with like on the variables I measured; I would report the estimate with the assumption list attached, and say what unmeasured factor would worry me most.

  • What if the confounder is measured with error or only in coarse buckets?
    Then you only partly adjust for it. Regression conditions on the recorded value, so any imbalance that survives inside a bucket, or any noise between the true value and the recorded one, leaves residual confounding. The estimate moves toward the truth but does not reach it, and the leftover bias usually points the same way as the original one.
  • The coefficient barely changes when you add the covariate. Does that prove there was no confounding?
    No. A stable estimate is consistent with no confounding, but also with two biases that offset, with a covariate measured too crudely to matter, or with confounding by something you never measured. Change-in-estimate is a diagnostic, not a test, and it says nothing at all about unmeasured causes.
  • Why does adjustment need both treated and untreated units at every covariate value?
    Because a comparison inside a covariate stratum needs both arms present. Where one arm is missing, the model has no data for the contrast and fills the gap by extrapolating its functional form. The estimate then depends on a modelling assumption rather than on any observed comparison, and can be arbitrarily wrong without any visible warning.

Comparing the fuel economy of two car models by only comparing trips over similar terrain, instead of averaging one model's mountain drives against the other's motorway runs.

saying these in an interview costs you the question

  • Says any variable correlated with the outcome removes bias
  • Calls the adjusted coefficient the causal effect with no assumptions
  • Thinks a higher R-squared means confounding has been handled
  • Believes adjustment can fix unmeasured confounding
  • Treats a stable coefficient as proof of no confounding

context

open as a page

How do you sign the omitted-variable bias when ability is left out of a wage-on-schooling regression?

level: middleimportance: must knowfreq 68%

basics

~20 s

Multiply two signs: the effect of the omitted variable on the outcome, times its correlation with the included regressor. Ability raises wages and is positively correlated with schooling, so the schooling coefficient is biased upward.

open as a page

How does standardisation, the g-formula, turn a fitted outcome model into an average treatment effect?

level: middleimportance: should knowfreq 44%

basics

~20 s

Fit one outcome model on treatment and covariates, then predict every unit twice, once with treatment set to 1 and once set to 0, and average the difference of those two predictions over the whole sample.

open as a page

In an adjusted regression table, why is only the exposure coefficient causally interpretable?

level: seniorimportance: should knowfreq 38%

basics

~20 s

The adjustment set was chosen to identify one exposure's effect. Each covariate has its own confounders, so its coefficient generally mixes a partial effect with leftover confounding. Reading every row as an effect is the Table 2 fallacy.

open as a page

Why can a logistic model's conditional odds ratio differ from the marginal odds ratio?

level: seniorimportance: nice to knowfreq 26%

basics

~20 s

The odds ratio is non-collapsible: adjusting for a covariate that predicts the outcome moves it away from 1 even with no confounding. Conditional and marginal odds ratios are different quantities, not a biased and an unbiased version.

open as a page