How is the variance inflation factor computed, and what does a VIF of 12 mean?
answer
- Regress the predictor on the other predictors
- One number per predictor, not per model
- It multiplies a variance, not a standard error
- Take the square root before you interpret it
- 5 and 10 are convention, not theory
basics
~20 sThe variance inflation factor for a predictor is 1 divided by (1 minus its R squared from regressing it on the other predictors). A VIF of 12 inflates that coefficient's variance twelvefold and its standard error about 3.5 times.
solid answer
~50 sTo get the VIF for predictor j, regress that predictor on all the *other* predictors, take the resulting `R_j^2`, and compute `VIF_j = 1 / (1 - R_j^2)`. It answers one question: how much of this predictor is already reconstructible from the rest? The number is a *variance* multiplier, so the interpretable version is its square root on the standard-error scale. A VIF of 12 means `R_j^2 = 1 - 1/12 = 0.917`, the coefficient's variance is 12 times what it would be if this predictor were unrelated to the others, and its standard error and confidence-interval width are `sqrt(12) ~ 3.5` times larger. The common thresholds are 5 (`R_j^2 = 0.8`) and 10 (`R_j^2 = 0.9`), but neither is a law — they are conventions, and what matters is whether the resulting interval is still tight enough to support the decision you are making.
go deeper
Memorise the formula 1 / (1 - R squared) and, more importantly, that the R-squared comes from predicting one predictor with the other predictors, never from the outcome. Getting that regression backwards is the usual stumble.
Be ready to derive the interpretation on the spot: variance multiplier, square root for the standard error, and the R-squared implied by any VIF you are handed. Expect to be asked what a VIF of 4 or 10 corresponds to.
Show that you convert the number into a decision. Explain when a large VIF is genuinely harmless because the affected coefficient is not the one your conclusion depends on, and why the table must be recomputed after each edit.
Argue against threshold rituals in team standards. Set the norm that reported precision, not an arbitrary cut-off, decides whether a predictor set is acceptable, and that a diagnostic number is never itself the optimisation target.
## The quantity being measured The variance inflation factor answers a narrow, precise question about one predictor: *how much of this predictor's variation is already explained by the other predictors in the model?* It does not look at the outcome at all. The recipe, for predictor j: 1. Set predictor j aside as a temporary outcome. 2. Regress it on every other predictor in the model (not the real outcome). 3. Record the R-squared of that auxiliary regression, written `R_j^2`. 4. Compute `VIF_j = 1 / (1 - R_j^2)`. Repeat for each predictor. You get one VIF per predictor, not one per model. The reciprocal, `1 - R_j^2`, is sometimes reported as **tolerance**; tolerance and VIF carry identical information. ## Why that formula, and what the number multiplies The variance of a least-squares slope is `Var(b_j) = sigma^2 / ( SST_j * (1 - R_j^2) )` where `sigma^2` is the error variance and `SST_j` is the total variation of predictor j about its own mean. If predictor j were orthogonal to the others, `R_j^2` would be zero and the variance would be `sigma^2 / SST_j`. The ratio of the actual variance to that ideal is exactly `1 / (1 - R_j^2)` — hence the name: it is the factor by which collinearity *inflates the variance* of this coefficient. That is the entire content of the statistic. ## Reading a VIF of 12 Because the VIF lives on the variance scale and people reason on the standard-error scale, always take the square root before interpreting. - `R_j^2 = 1 - 1/12 = 0.917`: about 92% of this predictor is reproducible from the others. - Variance multiplier: 12. - Standard error multiplier: `sqrt(12) = 3.46`. - Confidence-interval width multiplier: also about 3.46, since the interval is the estimate plus or minus a multiple of the standard error. - Equivalent sample-size intuition: standard errors shrink roughly like `1 / sqrt(n)`, so recovering that precision by collecting data would take roughly 12 times the sample. A useful reference table: | VIF | `R_j^2` | SE inflation | |---|---|---| | 1 | 0.00 | 1.0x | | 2 | 0.50 | 1.4x | | 4 | 0.75 | 2.0x | | 5 | 0.80 | 2.2x | | 10 | 0.90 | 3.2x | | 12 | 0.92 | 3.5x | | 100 | 0.99 | 10.0x | Note how gentle the low end is: a VIF of 2 costs only 40% on the standard error, which is why alarm below about 4 is usually overreaction. ## The 5-versus-10 argument The two conventional cut-offs correspond to `R_j^2` of 0.8 and 0.9. Neither has a theoretical basis — no distribution changes at those points, and no test is being performed. They are folklore thresholds, and treating a VIF of 9.6 as fine and 10.4 as a crisis is indefensible. The defensible reframing is to ask what the inflation costs *you*: - If the coefficient is your headline estimate and a 3.5x wider interval now spans zero and both plausible business decisions, the collinearity is a real problem at any VIF. - If the coefficient belongs to a control variable you never interpret, a VIF of 30 on it is irrelevant. Control variables are allowed to be tangled with each other. - If you have a very large sample, a high VIF can still leave a usably tight interval. Precision is the product of sample size, predictor variation and inflation — not inflation alone. So the honest answer to "is 12 too high?" is: it means a 3.5x wider interval; whether that is tolerable depends on what the interval has to decide. ## Practical cautions - **The VIF is model-relative.** It is computed against the current predictor set, so removing or adding one predictor changes every other VIF. Recompute after each change rather than trusting the original table. - **Pairwise relationships can hide it.** A predictor can be nearly reproducible from a *combination* of several others while being only moderately related to any one of them, so scanning pairs of predictors misses genuine problems. The auxiliary regression catches exactly this, which is why the VIF is the right tool. - **The intercept has no meaningful VIF**, and a VIF is defined per predictor, not per model, so "the model's VIF" is not a quantity. - **A high VIF is a description, not a verdict.** It quantifies how imprecise a coefficient is; it says nothing about whether that predictor belongs in the model, and dropping predictors purely to make a table of numbers smaller is optimising the diagnostic instead of the analysis. - **Structural cases exist.** When a model includes a predictor together with a transformation or product built from it, a large VIF can reflect the construction and the scale of the raw variable rather than genuine redundancy; centring the raw predictor before forming the derived term often shrinks the VIF without changing the fit at all. ## What a strong answer sounds like Give the formula, name the auxiliary regression, convert to the square-root scale immediately, and then refuse the threshold ritual: report what the inflation does to the interval you care about, and let that drive the decision.
- Why is the square root of the VIF the more useful number to quote?Because the VIF multiplies the *variance* of the coefficient, while people reason about standard errors and confidence-interval widths, which are on the square-root scale. A VIF of 9 sounds alarming but only triples the standard error; a VIF of 2 sounds meaningful but costs about 40%. Quoting the square root keeps the discussion tied to interval width.
- Can pairwise correlations between predictors miss a collinearity problem the VIF catches?Yes. A predictor can be almost perfectly reconstructible from a weighted combination of several others while showing only moderate relationships with each one individually. Because the VIF comes from regressing that predictor on all the others at once, it detects this multi-predictor redundancy that any pairwise scan will miss.
- Why do all the VIFs change when you drop one predictor?Each VIF is defined against the current set of other predictors. Removing one changes the auxiliary regression for every remaining predictor, usually lowering their R-squared values and hence their VIFs. The table must therefore be recomputed after each change, and a single pass of deleting everything above a threshold is not valid.
- Is a VIF of 30 on a control variable a reason to remove it?Not by itself. The VIF describes the precision of that control's own coefficient, which you were never going to interpret. Removing a legitimate control to tidy a diagnostic table trades a meaningless number for a possibly misspecified model. Act on VIFs attached to coefficients your conclusions actually rest on.
saying these in an interview costs you the question
- Computes VIF by regressing the predictor on the outcome
- Reads a VIF of 12 as a 12-fold standard error increase
- Treats 10 as a hard scientific threshold
- Reports a single VIF for the whole model
- Deletes every high-VIF predictor at once without recomputing
- Says a low VIF proves the coefficient is trustworthy