skip to content

Multicollinearity and VIF

Correlated predictors leave the model unable to split credit between them, so coefficients swing and even flip sign while overall fit looks fine. VIF puts a number on that inflation.

on this pageshow

questions

5

What is multicollinearity in a linear regression, and what does it damage?

level: juniorimportance: must knowfreq 78%

answer

  1. It is about the predictors, not the outcome
  2. Estimates stay unbiased
  3. Standard errors, not the fit, take the hit
  4. Inches and centimetres in the same model

basics

~20 s

Multicollinearity means the predictors are strongly linearly related to each other, so the fit cannot tell their separate effects apart. It inflates coefficient standard errors and makes individual coefficients unstable, while overall fit and predictions stay largely intact.

solid answer

~50 s

Multicollinearity is redundancy *among the predictors*, not between a predictor and the outcome. The limiting case is exact: put height in inches and height in centimetres into the same model and there is no unique least-squares solution, because inches and centimetres carry identical information and infinitely many coefficient pairs produce identical fitted values. Near-collinearity is the practical version: the estimates still exist and are still unbiased, but the data barely distinguishes one predictor's contribution from the other's, so standard errors blow up, confidence intervals get wide, t-statistics collapse, and a coefficient can flip sign when you add a few rows or drop a control. What does *not* break is the fit itself: `R^2`, residuals and predictions for new points that look like the training data are essentially unaffected. It is an identification problem for individual effects, not a bias or accuracy problem.

go deeper

for a junior

Be ready to define it in one sentence as overlap among predictors, and to name the two things that suffer: standard errors and the stability of individual coefficients. Knowing that overall fit and predictions survive is what separates a crisp answer from a vague one.

for a middle

Expect to explain the mechanism, not just the symptom: why shared variation leaves no data to identify a separate effect, and why the variance formula's collinearity factor grows as the other predictors explain more of this one.

for a senior

An interviewer will want you to distinguish cases where it changes a decision from cases where it is noise, and to recognise the symptoms in a real fit — flipped signs, huge standard errors, estimates that move when a control is added.

for a principal

Own the framing that this is an identification limit of the data, not a defect to be patched away. Decide and defend what the team reports when an effect genuinely cannot be separated, rather than shipping a point estimate the data does not support.

## What multicollinearity is A linear regression estimates a separate slope for each predictor, and the meaning of each slope is *the change in the outcome per unit change in that predictor, holding the other predictors fixed*. That phrase is doing all the work. To estimate it, the data has to contain observations where one predictor moves while the others stay still. Multicollinearity is the situation where that variation barely exists: one predictor is close to a linear combination of the others, so "holding the others fixed" describes a comparison the data almost never makes. Crucially this is a relationship **among predictors**. A predictor being strongly correlated with the *outcome* is not multicollinearity — that is just a useful predictor. ## Exact versus near collinearity The exact case is easiest to see. Suppose you enter height in inches and height in centimetres as two predictors. Since centimetres = 2.54 x inches, the two columns carry identical information. If one candidate solution gives coefficients `b_in` and `b_cm`, their combined contribution is `b_in * h_in + b_cm * 2.54 * h_in = (b_in + 2.54 * b_cm) * h_in`. Any pair with the same value of `b_in + 2.54 * b_cm` produces exactly the same fitted values and exactly the same residuals. There is no unique answer, so the estimates are undefined; software will typically drop one column or report a missing coefficient. Near-collinearity is the case you actually meet. The predictors are not identical, just nearly so — two engagement metrics that mostly measure the same behaviour, or spend and impressions in an advertising model. Now a unique solution exists, but it sits in a very flat, elongated valley of the error surface: many quite different coefficient pairs fit almost equally well, and small changes in the data slide you along that valley floor. ## What actually degrades The variance of an estimated slope in ordinary least squares can be written `Var(b_j) = sigma^2 / ( SST_j * (1 - R_j^2) )` where `sigma^2` is the error variance, `SST_j` is the total variation of predictor j around its own mean, and `R_j^2` is the R-squared from regressing predictor j on all the *other* predictors. The last factor is the collinearity term. If predictor j is unrelated to the others, `R_j^2 = 0` and it disappears. If the others explain 90% of predictor j, `1 - R_j^2 = 0.1` and the variance is ten times larger — the standard error is about 3.2 times larger. Nothing else in the formula changed; the coefficient simply cannot be pinned down. Consequences that follow directly: - Wide confidence intervals on the affected coefficients. - Small t-statistics, so genuinely relevant predictors look "insignificant". - Unstable signs and magnitudes: refit on a bootstrap resample, or add a quarter of new data, and the numbers move a lot. - Coefficients that contradict domain knowledge, because the fit is free to give one predictor a large positive weight and its twin a large negative one. ## What does not degrade - **Unbiasedness.** Under the usual assumptions the estimates are still unbiased; on average across repeated samples they are centred on the true values. Multicollinearity is a *variance* problem, not a *bias* problem. Saying "it biases the coefficients" is the single most common wrong answer. - **Overall fit.** `R^2`, the residual variance and the fitted values are unaffected. Two collinear predictors jointly explain what they always explained; the fit only refuses to say how to split the credit. - **Prediction.** For new observations that share the same correlation structure as the training data, predictions and their intervals are fine. The danger appears only if you predict at a combination the data never contained — for example a case with high inches and low centimetres, which the model has never seen and has no basis to handle. - **Coefficients with low collinearity.** If your predictor of interest is uncorrelated with the others, its standard error is untouched even when two control variables are tangled with each other. ## How it shows up in practice Typical tells: a coefficient with an implausible sign; estimates that change dramatically when one predictor is added or removed; and enormous standard errors on predictors you know matter. The natural next step is to quantify the redundancy for each predictor rather than eyeballing pairwise relationships, since a predictor can be nearly reproduced by a *combination* of several others while being only mildly related to any single one. ## The right mental frame Multicollinearity does not mean the model is wrong. It means a specific question — "what is the separate effect of this predictor?" — cannot be answered precisely with this data. If you only need forecasts, that may not matter at all. If you need to attribute an effect or set a policy on one lever, it matters enormously, and honest practice is to report the wide interval rather than the one point estimate that happened to come out of this sample.

  • Does multicollinearity bias the coefficient estimates?
    No. Under the usual least-squares assumptions the estimates remain unbiased — across repeated samples they centre on the true values. What changes is their variance: standard errors and confidence intervals inflate, so any single sample can land far from the truth. It is a precision problem, not a systematic-error problem.
  • If two predictors are exactly collinear, what does the fit return?
    Nothing unique. Infinitely many coefficient combinations give identical fitted values and identical residuals, so no single solution can be preferred. In practice the fitting routine either drops one of the redundant columns or reports its coefficient as undefined. The overall fit is still perfectly well defined — only the split between the two columns is not.
  • Does multicollinearity hurt predictive accuracy?
    Usually not, as long as new observations share the same relationship among the predictors that the training data had. The fitted values are stable even when the individual coefficients are not. Accuracy degrades only when you predict at predictor combinations that never occurred in training, where the unstable split between the correlated predictors suddenly matters.

Two people always push a cart together, never separately. You can measure exactly how fast the cart moves, but no amount of watching tells you how much each person contributed.

saying these in an interview costs you the question

  • Says multicollinearity biases the coefficient estimates
  • Confuses predictor-to-predictor overlap with predictor-to-outcome correlation
  • Claims it lowers R-squared or worsens the model's fit
  • Treats any correlation between two predictors as disqualifying
  • Says predictions become unreliable in every case

context

open as a page

How is the variance inflation factor computed, and what does a VIF of 12 mean?

level: middleimportance: must knowfreq 68%

basics

~20 s

The variance inflation factor for a predictor is 1 divided by (1 minus its R squared from regressing it on the other predictors). A VIF of 12 inflates that coefficient's variance twelvefold and its standard error about 3.5 times.

open as a page

A regression's overall F-test is highly significant but no coefficient's t-statistic clears 2 — why?

level: seniorimportance: should knowfreq 52%

basics

~20 s

That pattern is the classic signature of severe multicollinearity. The predictors jointly explain the outcome well, which the F-test detects, but they overlap so heavily that no single coefficient can be pinned down, so every individual t-statistic stays small.

open as a page

Two near-duplicate engagement predictors flip between +40 and -38 across bootstrap refits — how do you fix the model?

level: principalimportance: should knowfreq 41%

basics

~20 s

First decide what the model is for. If it only forecasts, the swings are harmless. If someone will act on a coefficient, the data cannot separate the two effects, so drop one, combine them into a single measure, or report the well-estimated joint effect instead.

open as a page

When is high multicollinearity safe to ignore in a regression model?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Ignore it when the inflated coefficients are ones you never interpret: pure forecasting, collinearity confined to control variables, or a predictor of interest that is itself uncorrelated with the tangled ones. It matters only when a decision rests on a separated effect.

open as a page