Why center age and income before adding their interaction to a regression?
answer
- ask where the main effect is anchored
- zero income is not in the data
- the product correlates with its own parts
- reparameterisation, not a better model
- the interaction coefficient never moves
basics
~20 sCentering makes each main effect the effect at the other predictor's mean rather than at zero, which is often outside the data, and it removes the artificial correlation between main effects and their product. Fit is unchanged.
solid answer
~50 sWith an `age x income` product in the model, the age coefficient is the effect of age at `income = 0` - a point no one in the sample occupies, so the number is technically correct and practically useless. Subtracting each variable's mean before multiplying moves that anchor to the average income, which is a value the data actually contains, and the same for the income coefficient at average age. The second benefit is numerical: a raw product is strongly correlated with its own components, inflating variance inflation factors and standard errors on the main effects. Centering removes that *nonessential* collinearity. What centering does not do is change the model: the interaction coefficient, its standard error, the fitted values and R-squared are all identical, and genuine correlation between two distinct predictors is untouched. So it is a reparameterisation for interpretability, not a fix for a badly conditioned design.
go deeper
Know the headline reason: without centering, a main effect in an interaction model describes the effect at zero on the other variable, which for things like income is a value nobody has.
Explain the mechanics - what the centered product expands to, which coefficients move, and why the artificial correlation between a product and its components disappears.
Demonstrate the invariance clearly: same fit, same interaction coefficient and standard error. Also handle where to store the centering constants so later data uses the same anchor.
Own the reporting standard for the team: what anchor coefficients are quoted at, when a decision-relevant centering value beats the mean, and how to stop VIF numbers being presented as evidence of model quality.
## The interpretability problem In `y = b0 + b1*age + b2*income + b3*(age * income) + error` the effect of age is `b1 + b3*income`. The printed `b1` is that effect evaluated at `income = 0`. For an indicator partner, zero is a real group. For income it is not: nobody in the sample has zero income, so `b1` describes a hypothetical person outside the data, and its sign can even look absurd relative to the effect anywhere in the observed range. Centering fixes this by redefining the variables. Let `age_c = age - mean(age)` and `inc_c = income - mean(income)`, then fit `y = a0 + a1*age_c + a2*inc_c + a3*(age_c * inc_c) + error` Now `a1` is the effect of age at *average* income, and `a2` is the effect of income at *average* age - both statements about points the data actually covers. ## What changes and what does not This is the part interviewers probe, because it separates people who have done it from people who have read about it. Unchanged: - The interaction coefficient: `a3 = b3` exactly, along with its standard error and t-statistic. - The fitted values, residuals, R-squared and the residual standard error. The two models span the same column space, so they are the same model in different coordinates. Changed: - The intercept, which now describes an average-age, average-income case rather than a zero-zero case. - The two main effects and their standard errors, because they are now anchored at a different point. A quick sanity check for anyone doubtful: expand `(age - m)(inc - n)` and you get `age*inc` minus linear terms plus a constant, so the centered product differs from the raw product only by a linear combination of variables already in the model. That is why only the coefficients on those lower-order terms move. ## The collinearity story, told accurately A raw product term is mechanically correlated with its own components - if age runs from 30 to 60, `age * income` moves almost in lockstep with income. This shows up as large variance inflation factors and wide standard errors on the main effects, and it can make the fit numerically fragile. Centering largely removes that correlation, because a mean-centered product is roughly uncorrelated with its centered components for symmetric distributions. But this is *nonessential* collinearity: an artefact of where the origin sits, not information missing from the data. The proof is that the interaction's own standard error does not budge. So centering does not make the model better able to separate age from income; if those two are genuinely correlated in the population, they remain just as entangled after centering. Presenting a VIF drop as evidence that a collinearity problem was solved is a misreading interviewers listen for. ## When not to center, or to center elsewhere - Do not center a 0/1 indicator: zero already means something concrete, and centering replaces a clean group contrast with a fractional value nobody has. - Center at a decision-relevant value rather than the mean when one exists - a policy threshold, a target price, the start of a program. The mean is a convenient default, not a rule. - If you report predictions, nothing changes; if you report coefficients, say explicitly what the anchor is. "Effect of age at mean income" is the sentence that makes the table readable. - Standardising (dividing by the standard deviation as well as subtracting the mean) is a further step with its own tradeoffs and is not required to get the interpretability benefit. ## A common failure in practice A subtle bug: centering using the full-sample mean and then evaluating on a later cohort whose mean differs. The coefficients still refer to the *original* centering constant, so "effect at the mean" silently means the old mean. Store the constants with the model and apply the same ones everywhere, exactly as you would with any other fitted transformation. ## What interviewers listen for Two benefits (interpretable main effects, reduced artificial collinearity), one firm invariance claim (the interaction coefficient, its standard error and the fit are untouched), and one limit (it does not cure real collinearity). A candidate who claims centering improves fit or removes genuine multicollinearity has the concept backwards.
- Does centering change the interaction coefficient or the model's predictions?Neither. The centered and uncentered models span the same space, so the fitted values, residuals, R-squared and residual standard error are identical, and the interaction coefficient and its standard error are exactly the same number. Only the intercept and the two main effects change, because they are now anchored at the means rather than at zero. That invariance is the cleanest evidence that centering is a reparameterisation, not a different model.
- Does centering solve multicollinearity in general?No. It removes only the artificial correlation between a product term and its own components, which comes from where the origin sits. If age and income are genuinely correlated with each other in the population, they stay exactly as correlated after centering, and the model is no better able to separate their effects. The giveaway is that the interaction's standard error does not improve, so nothing real was gained.
- When would you center at something other than the mean, or not center at all?I would not center a 0/1 indicator, because zero is already a meaningful group and centering turns a clean contrast into a fraction nobody has. When a decision-relevant anchor exists - a policy threshold, a target price, a program start - I center there instead, so the main effect answers the question the business asks. The mean is just a sensible default when no such value stands out.
Centering is like moving the origin of a map to the middle of town instead of a point in the sea: every distance between places is unchanged, but the coordinates finally describe somewhere real.
saying these in an interview costs you the question
- Claims centering changes the interaction coefficient
- Says centering fixes multicollinearity in general
- Thinks centering improves model fit or R-squared
- Centers a 0/1 indicator whose zero is already meaningful
- Treats a lower VIF after centering as a better model