skip to content

Why does adding all four region dummies plus an intercept break an OLS regression?

level: middleimportance: must knowfreq 70%

answer

  1. count the columns and the levels
  2. every row belongs to exactly one region
  3. the intercept is a column of ones
  4. perfect, not high, collinearity
  5. singular design matrix, no unique solution

basics

~20 s

The four region dummies add to 1 in every row, reproducing the intercept's column of ones exactly. The design matrix loses full rank, so no unique least-squares coefficients exist. This is the dummy variable trap.

solid answer

~40 s

Each row belongs to exactly one region, so `North + South + East + West = 1` for every observation — which is precisely the intercept's column. That is exact, not approximate, collinearity: one column is a perfect linear combination of the others, the design matrix is rank-deficient, and `X'X` is singular so the normal equations have no unique solution. Infinitely many coefficient vectors give byte-identical fitted values, which is why individual coefficients become meaningless rather than merely imprecise. The fix is to identify the model: keep the intercept and drop one region, so the remaining three coefficients are contrasts against the dropped baseline; or drop the intercept and keep all four, so each coefficient is that region's own level rather than a contrast. Do one or the other, never neither and never both.

go deeper

for a junior

Know the rule and be able to state it without hesitation: a categorical with k levels enters as k-1 dummies when the model has an intercept, and the omitted level becomes the baseline.

for a middle

Explain the mechanism, not the rule. Show that the dummies sum to the intercept's column of ones, that this makes the design matrix rank-deficient, and that the normal equations then have no unique solution.

for a senior

Demonstrate you catch it in output — a blank or dropped coefficient, a singular-matrix error, an accidental baseline — and that you spot the disguised versions, such as nested categoricals or an aggregate dummy equal to the sum of others.

for a principal

Own the convention across the team: where baselines are set, how model specifications are reviewed so a silently dropped column never reaches a decision, and when a no-intercept parameterisation is worth its reporting costs.

## The setup A region variable takes four values — North, South, East, West — and each observation belongs to exactly one of them. Encode all four as 0/1 indicator columns and put them into a regression that also has an intercept. The model will not fit, or will fit only after something silently disappears. ## Why it breaks The intercept is not a magic constant; it is a column of ones in the design matrix `X`. Because every row belongs to exactly one region, the four dummy columns satisfy ``` North + South + East + West = 1 for every row ``` That sum *is* the intercept column. So one column of `X` is an exact linear combination of four others. `X` is **rank-deficient**: its columns are linearly dependent. Ordinary least squares solves the normal equations `X'X b = X'y`. When `X` is rank-deficient, `X'X` is singular — it has no inverse — and the system has infinitely many solutions rather than one. This is **perfect** (or exact) multicollinearity, and it is categorically different from the everyday case of two predictors being merely highly correlated. High correlation inflates standard errors; a perfect linear dependency destroys identification altogether. ## What 'infinitely many solutions' looks like Suppose one solution has intercept 100 and region coefficients (10, 4, -2, 0). Add any constant `c` to the intercept and subtract the same `c` from all four region coefficients: intercept 105 with (5, -1, -7, -5). Every fitted value is unchanged, because every row picks up `+c` from the intercept and `-c` from its own region. The residuals, the residual sum of squares and R-squared are identical too. So the *fit* is perfectly well defined — the column space of `X` is unchanged, and the projection of `y` onto it is unique. What is not defined is the **decomposition** of that fit into an intercept plus region effects. That is what 'the coefficients are not identified' means: the data contain no information that would let you prefer one of those infinitely many splits over another. A candidate who says 'the model overfits' or 'the standard errors get large' has missed the point; nothing was estimated at all. ## The two clean fixes **Fix 1 — intercept plus k-1 dummies.** Drop one region, say North. The remaining columns are linearly independent, the model is identified, and the reading changes accordingly: the intercept is the expected outcome for North (with other predictors at zero), and each of the three surviving coefficients is that region minus North. The dropped level is the reference level. This is the default and the one that generalises, because as soon as the model contains a second categorical variable you cannot drop the intercept twice. **Fix 2 — no intercept, all k dummies.** Keep all four columns and remove the intercept. The model is again identified, and each coefficient is now that region's own expected outcome rather than a contrast. This reads nicely for a single factor, but it has costs: contrasts between regions now require differencing coefficients, the reported R-squared from a no-intercept fit is computed differently and is not comparable with an intercept model's, and the trick cannot be repeated for a second categorical predictor. What you must not do is both — dropping the intercept *and* a level leaves the model unable to represent the dropped region at all. ## Detecting it The symptom is unmistakable once you know it: a coefficient reported as missing or blank, an error about a singular or non-invertible matrix, or a fit that quietly returns one fewer coefficient than you supplied columns. Many fitting routines guard against the trap by dropping a redundant column automatically, which is convenient and dangerous — the model fits, and unless you read the output carefully you may not notice that the level you thought you were estimating became the baseline. The same trap appears in less obvious dress. Two categorical variables whose levels nest — a country dummy set and a currency dummy set where each country has exactly one currency — produce a dependency of the same kind. So does a set of dummies for month plus a dummy for 'quarter one' that is exactly the sum of three of them. Any time indicator columns are constructed so that some subset adds up to another column, identification is gone. ## The one-line answer to give The dummies for an exhaustive, mutually exclusive categorisation sum to the intercept, so including all of them alongside an intercept makes the design matrix singular and the coefficients non-identified; use `k-1` dummies with an intercept, or `k` without one.

  • If the coefficients are not identified, are the fitted values also undefined?
    No. The column space of the design matrix is unchanged by the redundancy, so the projection of the outcome onto it — the fitted values, residuals and residual sum of squares — is unique. Only the split of that fit into an intercept and region effects is arbitrary: you can shift a constant from the intercept into every dummy and get identical predictions.
  • What changes in interpretation if you drop the intercept and keep all four region dummies?
    Each coefficient becomes that region's own expected outcome rather than a contrast with a baseline. Differences between regions then require subtracting coefficients, and R-squared from a no-intercept fit is computed on a different basis and should not be compared with an intercept model's. The trick also cannot be repeated once a second categorical predictor is in the model.
  • How is this different from two numeric predictors that are highly correlated?
    High correlation is imperfect collinearity: the model is still identified, the coefficients exist and are unique, but their standard errors are inflated and the estimates are unstable. The dummy variable trap is an exact linear dependency, so there is no unique solution at all. One is a precision problem; the other is an identification failure.
  • A fitting routine returns the model with one region coefficient blank. What happened?
    It detected the redundant column and dropped it to restore full rank, so the level with the blank entry silently became the reference level. The fit is valid, but the output now answers a different question than you intended, and every remaining region coefficient is a contrast against that accidental baseline. Set the baseline deliberately instead.

saying these in an interview costs you the question

  • Calls it overfitting rather than an identification failure
  • Says the standard errors merely get large
  • Drops both the intercept and one level
  • Thinks fitted values are undefined too
  • Cannot say why the dummies sum to the intercept column

context