skip to content

In an adjusted regression table, why is only the exposure coefficient causally interpretable?

level: seniorimportance: should knowfreq 38%

answer

  1. one model answers one question
  2. the covariate list serves the exposure
  3. covariates have confounders of their own
  4. conditional is not total
  5. nuisance parameters, not findings

basics

~20 s

The adjustment set was chosen to identify one exposure's effect. Each covariate has its own confounders, so its coefficient generally mixes a partial effect with leftover confounding. Reading every row as an effect is the Table 2 fallacy.

solid answer

~50 s

A model is built to answer one causal question. You pick covariates so that the exposure's coefficient is unconfounded — and that choice says nothing about whether the *covariates* are themselves unconfounded. A covariate's coefficient is a contrast conditional on the exposure and on all the other covariates, which is rarely the quantity anyone means by the effect of that covariate; it can be a partial rather than total effect, and it can carry confounding from causes that were never in the model because they were irrelevant to the exposure. This is the Table 2 fallacy, named for the results table where every adjusted coefficient is printed side by side under one heading. The fix is presentational and structural: report the exposure estimate as the effect, list the covariates without effect language, and if a second exposure matters, fit a second model with its own adjustment set.

go deeper

for a junior

Know that a regression table is built to answer one causal question, and that the other rows are there to make that one comparison valid rather than to be reported as effects.

for a middle

Explain the two mechanisms: covariates carry confounders nobody adjusted for, and a coefficient conditional on the exposure is a partial contrast rather than a total effect.

for a senior

Show how you present results so the fallacy cannot happen downstream — one effect per model, covariates without effect language, and a separate model with its own adjustment set for a second exposure.

for a principal

Own the norm across the team: results templates, review expectations, and the willingness to answer a request for five effects with five analyses or an honest refusal rather than one convenient table.

## What the fallacy is A typical results table shows one row per model term: the exposure of interest, then eight or ten covariates, each with an estimate and an interval, all under a column heading like *adjusted effect*. The Table 2 fallacy is treating all of those rows as causal effects when only one of them was ever designed to be. ## Why only the exposure row is privileged The covariate list in a causal regression is not a general-purpose collection of things that matter. It is chosen for one job: to make the exposure comparison valid. Nothing in that selection process does the same job for any other variable in the model. Three separate failures follow. **1. Each covariate has its own confounders.** Suppose the exposure is a pricing plan and you adjusted for account tenure because tenure influences both plan choice and renewal. The tenure coefficient is now on display — but you never asked what confounds tenure and renewal, because you did not need to. Whatever those causes are, if they are not in the model, the tenure row is confounded. **2. Conditional does not mean total.** The coefficient on a covariate is a contrast holding the exposure and every other covariate fixed. If that covariate influences the outcome partly through the exposure, holding the exposure fixed removes that channel, and the coefficient describes a partial pathway rather than the covariate's overall influence. The number is well defined; it just is not the quantity a reader assumes it is. **3. The estimands differ silently.** Even where a covariate coefficient is interpretable, it may answer a narrower question than the exposure coefficient does, with different assumptions behind it. Printing them in one column implies a comparability that does not exist. ## The concrete version You compare two pricing plans on renewal and add city fixed effects because plan availability and renewal both vary by city. After the fit, the plan coefficient is the within-city comparison you wanted — that is the causal row, conditional on the adjustment doing its job. The city indicators are the ones people misread. They are not the effect of a city on renewal; they absorb every stable difference between cities at once, including local competition, sales staffing, regulation, and the mix of customers who live there. They are nuisance parameters that make the plan comparison work, and a slide titled *effect of city on renewal* built from them is exactly the fallacy. ## What to do about it - **Report one effect per model.** State the exposure estimate with its interval and its assumptions. Put the covariates in an appendix, or label the column *model term* rather than *effect*. - **If several exposures genuinely matter, fit several models.** Each gets its own adjustment set, argued on its own terms. Two models with different covariate lists is the correct answer to two causal questions, and it is not a duplication of effort to be optimised away. - **Strip effect language from covariate rows.** Words matter. A reader who sees *effect* will act on it, and the phrase *adjusted for* in a caption does not undo the heading. - **Watch for it downstream.** The fallacy usually enters the organisation not in the analysis but in the summary slide, where an analyst pulls the three largest coefficients from a model and reads them as drivers. ## What covariate coefficients are still good for They are not worthless. They are legitimate for prediction, since prediction needs no causal reading at all, and they are useful for sanity checks: if a covariate whose relationship with the outcome is well understood in your domain comes out with an implausible sign or magnitude, something about the specification or the data deserves a look. Both uses are internal diagnostics, not findings, and neither justifies putting the number in front of a decision-maker with the word *effect* attached to it. ## In an interview Say the name — the Table 2 fallacy — and give the one-line reason: the covariate set was chosen for the exposure, so only the exposure row inherits the identification argument. Then give the fix in the same breath: one causal question per model, and no effect language on rows you did not design to be effects.

  • You add city fixed effects to compare two pricing plans. Which coefficient is the causal one?
    The plan coefficient, now a within-city comparison, assuming city was the confounder that needed closing. The city indicators are not effects of cities; they soak up every stable difference between cities at once — competition, staffing, customer mix — as nuisance parameters that make the plan comparison valid. Reading them as drivers of renewal is the fallacy in action.
  • Your stakeholder wants effects for all five variables in the model. What do you do?
    Fit five models, each with an adjustment set argued for its own exposure, or explain why some of those questions are not answerable from this data. The single-model shortcut looks efficient and gives four numbers that are not what was asked for. I would rather deliver one defensible estimate and a plan for the rest.
  • Are covariate coefficients useful for anything at all?
    Yes, for prediction, where no causal reading is required, and as a specification check — an implausible sign or magnitude on a well-understood covariate is a signal to look at the data or the functional form. Both are internal diagnostics. Neither belongs in a report with the word effect attached.

saying these in an interview costs you the question

  • Reads every row of the model output as a causal effect
  • Assumes one model can estimate all effects simultaneously
  • Treats a significant covariate as a proven driver
  • Uses covariate coefficient signs to judge confounding direction
  • Calls fixed-effect dummies the effect of the group they encode

context