skip to content

Why does adding any predictor to an OLS regression never lower its R-squared?

level: middleimportance: must knowfreq 74%

answer

  1. the smaller fit is still reachable
  2. set the new coefficient to zero
  3. the optimiser can only tie or improve
  4. the denominator of the ratio never moves
  5. adjusted R-squared charges for degrees of freedom

basics

~20 s

Least squares can always set the new coefficient to zero and reproduce the previous fit, so the minimised residual sum of squares can only tie or shrink. Since R-squared is 1 minus that sum over a fixed total, it never falls.

solid answer

~50 s

R-squared is `1 - RSS/TSS`, and the total sum of squares depends only on the outcome, so R-squared moves only when the residual sum of squares moves. Adding a column strictly enlarges the set of fits the estimator may choose from: the previous solution is still available with the new coefficient set to zero, so the minimised RSS is at most what it was. In practice a pure-noise column is almost never exactly uncorrelated with the leftover residuals, so RSS drops a little and R-squared creeps up - appending ten random columns can lift 0.62 to 0.64 with nothing real added. Adjusted R-squared repairs this by comparing variance estimates instead of raw sums: `1 - (RSS/(n-k-1)) / (TSS/(n-1))`, with n the sample size and k the number of predictors. Each extra column costs a degree of freedom, so a useless one pushes the adjusted value down.

go deeper

for a junior

Remember the direction: R-squared can only rise or stay flat as predictors are added, and adjusted R-squared is the version that is allowed to fall.

for a middle

Give the nesting argument out loud - the old fit is still reachable with a zero coefficient - and write the adjusted formula with its degrees of freedom correct.

for a senior

Show that you never justify keeping a variable by an R-squared bump, and that you know adjusted R-squared is too weak a filter to serve as a selection rule by itself.

for a principal

Set the team's standard for what evidence a new feature must bring, and be explicit that in-sample fit summaries are the weakest tier of that evidence.

## The mechanism Write the statistic as `R2 = 1 - RSS/TSS`. The total sum of squares `TSS = sum (y_i - y_bar)^2` is built entirely from the outcome and its mean, so no change to the predictor set can touch it. The whole behaviour of R-squared therefore comes from the residual sum of squares `RSS = sum (y_i - y_hat_i)^2`. Ordinary least squares chooses coefficients to minimise RSS over all linear combinations of the columns it is given. Add a column, and every combination that was available before is still available - just take the same coefficients and put a zero on the new one. The estimator is minimising over a strictly larger set, so the minimum it finds is at most the previous minimum. Formally: `RSS_full <= RSS_reduced` for nested models on the same rows, hence `R2_full >= R2_reduced`. Equality happens only in the knife-edge case where the new column adds nothing at all to the fit. ## Why noise still moves the number That knife-edge case is essentially never reached with real data. A column of random numbers will, by chance, line up slightly with whatever residual structure remains, and least squares will happily exploit that chance alignment. This is why appending ten columns of pure noise to a model sitting at 0.62 can move it to 0.64: each column buys a small, entirely fictitious reduction in RSS. Push the idea to its limit and it becomes obvious - with `k + 1` estimated parameters and `n = k + 1` observations, the fit can pass through every point, `RSS = 0`, and `R2 = 1` while conveying zero information. ## What adjusted R-squared changes Adjusted R-squared replaces the two raw sums with unbiased variance estimates, dividing each by its degrees of freedom: `R2_adj = 1 - (RSS/(n - k - 1)) / (TSS/(n - 1))` which is algebraically `1 - (1 - R2) * (n - 1)/(n - k - 1)`. Here `k` counts the predictors excluding the intercept. Now adding a column pulls in two directions: RSS falls (helping) while `n - k - 1` falls (hurting). A column that buys almost nothing loses that tug-of-war and the adjusted value drops. There is a sharp characterisation of the tipping point for a single added predictor: adjusted R-squared rises exactly when that predictor's t-statistic exceeds 1 in absolute value. That is a much weaker bar than the usual significance threshold near 2. So adjusted R-squared is a mild anti-padding correction, not a model-selection rule - a variable nowhere near significant at the 5 percent level can still raise it. ## The negative case Unlike R-squared, the adjusted version has no floor at zero. If the model explains almost nothing and `k` is large relative to `n`, the ratio `(RSS/(n-k-1)) / (TSS/(n-1))` can exceed 1 and the adjusted value goes negative. Read that as a blunt verdict: once the parameters spent are charged for, this model is worse than the outcome's mean. It is also undefined at `n - k - 1 = 0` and unstable as you approach it. ## Interview framing The answer an interviewer is listening for is the nesting argument stated out loud: the smaller model is still reachable inside the larger one, so the optimiser's best score cannot get worse. Everything else follows from that. The second half of a strong answer is the consequence for practice - a rise in R-squared is not evidence that a new variable belongs in the model, because the statistic was guaranteed not to fall. Any claim that a feature helped needs evidence the statistic is structurally incapable of providing. ## Related traps Comparing R-squared between models fitted on different row sets, for instance because one model dropped rows with missing values on a new predictor, breaks the argument entirely: TSS is no longer the same total, and the larger model's R-squared can then move in either direction for reasons that have nothing to do with the predictor's value.

  • When exactly does adding one predictor raise adjusted R-squared?
    When that predictor's t-statistic exceeds 1 in absolute value - a far weaker bar than the usual significance threshold near 2. Adjusted R-squared is therefore a mild penalty against padding, not a selection rule: a variable can raise it while being nowhere near significant at the 5 percent level.
  • Can adjusted R-squared be negative?
    Yes. When the model explains very little and the number of predictors is large relative to the sample size, the ratio of the two variance estimates exceeds 1 and the adjusted value drops below zero. Read it as saying the model is worse than the outcome's mean once its parameters are paid for.
  • What happens to R-squared as the number of predictors approaches the sample size?
    It marches toward 1. When the estimated parameters equal the number of observations, the fit passes through every point, the residual sum of squares is zero, and R-squared is exactly 1 with no information gained. Adjusted R-squared collapses as the residual degrees of freedom approach zero and is undefined there.

Handing least squares an extra column is like giving a contestant an extra guess. They can always ignore it, so their best score can never get worse - which is why an improved score proves nothing on its own.

saying these in an interview costs you the question

  • Claims a rising R-squared proves the new variable matters
  • Says an irrelevant predictor can lower R-squared
  • Thinks adjusted R-squared can never go negative
  • Treats adjusted R-squared as a significance test for the new term
  • Believes adding a column changes the total sum of squares

context