skip to content

Goodness of Fit

How much variation the model explains and whether extra predictors earned their place: adjusted R-squared, the overall F-test, AIC and BIC. Interviewers probe why R-squared never falls.

on this pageshow

questions

6

What does the R-squared of a fitted multiple regression tell you about the model?

level: juniorimportance: must knowfreq 78%

answer

  1. compares the model to a flat mean line
  2. one minus a ratio of two sums of squares
  3. residual variation over total variation
  4. unitless, and measured on the fitted data

basics

~20 s

R-squared is the share of the outcome's total variation that the fitted model accounts for, computed as 1 minus the residual sum of squares over the total sum of squares. It measures fit on the data used, not correctness.

solid answer

~40 s

R-squared compares the model's residual sum of squares, `RSS = sum of (y - y_hat)^2`, with the total sum of squares, `TSS = sum of (y - y_bar)^2`, as `R2 = 1 - RSS/TSS`. It answers one question: how much better does the fitted model track the outcome than always predicting the outcome's mean? For least squares with an intercept it sits between 0 and 1 and is unitless, so it travels across datasets in a way raw error sums do not. What it does not tell you is whether the functional form is right, whether an important predictor is missing, or whether any coefficient is causal. A very high value can come from a predictor that is the outcome in disguise, and a low value can still sit on a useful model.

go deeper

for a junior

Be ready to state the formula in words - one minus residual variation over total variation - and to say plainly that it is a share of variation, never an average percentage error and never a p-value.

for a middle

Explain why the comparison baseline is the outcome's mean, why the statistic is unitless, and why it cannot be compared across models that predict different outcome variables.

for a senior

Show that you interrogate a suspiciously high value before reporting it, and that you can defend a small but real R-squared when the outcome is genuinely noisy.

for a principal

Own the framing with stakeholders: agree in advance what fit level would change a decision, and push back on anyone treating R-squared as the single acceptance criterion for a model.

## The definition Every regression sits against a reference point. The laziest defensible predictor of an outcome `y` is its own mean `y_bar`: predict the same number for every row. How badly that predictor does is measured by the **total sum of squares**, `TSS = sum (y_i - y_bar)^2`. A fitted model instead produces predictions `y_hat_i`, leaving the **residual sum of squares**, `RSS = sum (y_i - y_hat_i)^2`. R-squared, also called the coefficient of determination, is the fraction of the lazy predictor's squared error that the model removed: `R2 = 1 - RSS/TSS = (TSS - RSS)/TSS` An R-squared of 0.30 says the fitted model wiped out 30 percent of the squared error that predicting the mean every time would leave behind. ## Why it is bounded and unitless For ordinary least squares fitted with an intercept, the fitted model can never do worse in-sample than the mean, because the mean-only fit is itself a special case the estimator could have chosen. So `RSS <= TSS`, and R-squared lands in `[0, 1]`. Both sums of squares carry the squared units of the outcome, and they are divided by each other, so the ratio has no units at all. That is exactly why R-squared is quoted and RSS usually is not: an RSS of 4.2 million means nothing until you know whether the outcome is measured in cents or in millions of euros, whereas 0.62 is readable on its own. ## What it does and does not establish R-squared is a summary of residual **size**, and nothing else. It is silent on: - **Functional form.** A curved relationship fitted with a straight line can still post a respectable R-squared while the residuals show obvious structure. - **Causality.** Explained variation is a statement about co-movement in the sample, not about what happens if you intervene on a predictor. - **Omitted variables.** A model missing the true driver can still explain a great deal via a correlate of it, and its coefficients can be badly biased while R-squared looks fine. - **Whether the value is impressive.** That depends entirely on the domain. In physical measurement, 0.95 is routine; in human behaviour, 0.10 can be the ceiling anyone achieves. ## Two directions of misreading The first misreading is treating a low value as automatic failure. If the outcome is close to individually unpredictable - which customer cancels next month, which visitor converts - most of the variation is idiosyncratic and no model will claim it. A model that explains a few percent while producing stable, well-sized coefficients can still be the basis of a real decision. The second misreading is treating a very high value as automatic success. In business data an R-squared near 1 is usually a defect rather than a triumph. The classic causes are a predictor computed from the outcome (a billed amount used to predict revenue), a predictor recorded after the outcome occurs, or a model carrying nearly as many free parameters as it has rows. Before reporting 0.98, find out what the strongest predictor actually is and when it becomes available. ## Things it cannot be compared across R-squared is a share of a specific total, so the total has to be the same for a comparison to mean anything. Two models of different outcomes - `y` versus `log y`, revenue versus revenue per user - have different TSS values and their R-squared numbers are not on a common scale. Likewise, comparing R-squared between two models fitted on different row sets (because one dropped rows with missing values) compares two different totals. ## How to say it in an interview A clean answer names the formula in words, names the baseline it compares against, states that it is unitless and in-sample, and then immediately gives the two-sided caution: low is not automatically bad, high is not automatically good. That last pair is what separates a candidate who has read the definition from one who has reported a model to a stakeholder.

  • Can R-squared ever be negative?
    Not for least squares fitted with an intercept on the same data, since the fitted model can always match the mean-only fit, so `RSS <= TSS`. It can go negative for a model fitted without an intercept, or when predictions come from coefficients that were set by hand rather than least-squares fitted to the data in hand.
  • Why can't you compare R-squared between a model of y and a model of log y?
    The two models explain different outcomes, so their total sums of squares are different totals. Explaining 60 percent of the variation in log spend and 40 percent of the variation in spend are answers to different questions. Compare within one outcome scale, or bring predictions back to a common scale before judging.
  • What does an R-squared of essentially 1 usually indicate in practice?
    Almost always a mistake rather than a triumph. The usual causes are a predictor that is a rescaling or component of the outcome, an identifier that indexes the outcome, or nearly as many estimated parameters as observations. Trace the strongest predictor back to its definition before reporting the number.

R-squared scores the model against a lazy competitor who always guesses the average. A value of 0.30 means the model removed 30 percent of the wobble that guessing the average leaves behind.

saying these in an interview costs you the question

  • Says R-squared shows whether the model is correctly specified
  • Treats a high R-squared as evidence of causation
  • Believes R-squared carries the units of the outcome
  • Assumes a low R-squared always means a useless model
  • Compares R-squared across models with different outcome variables

context

open as a page

Why does adding any predictor to an OLS regression never lower its R-squared?

level: middleimportance: must knowfreq 74%

basics

~20 s

Least squares can always set the new coefficient to zero and reproduce the previous fit, so the minimised residual sum of squares can only tie or shrink. Since R-squared is 1 minus that sum over a fixed total, it never falls.

open as a page

What null hypothesis does the overall F-test in a linear regression output test?

level: middleimportance: should knowfreq 58%

basics

~20 s

It tests whether every slope coefficient is zero at once, meaning the model does no better than an intercept-only model that predicts the outcome's mean. A small p-value says at least one predictor carries signal, without saying which.

open as a page

How do you test whether a block of three interaction terms improves a regression fit?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Fit the model with and without the three terms on identical rows and run a nested F-test: the drop in residual sum of squares per added parameter, divided by the full model's residual variance, on 3 and n-k-1 degrees of freedom.

open as a page

Is an R-squared of 0.04 ever good enough to ship a regression model?

level: principalimportance: should knowfreq 41%

basics

~20 s

Yes, when the model's job is to identify and size drivers in a noisy outcome rather than to predict individuals. The acceptance bar comes from the decision the model supports, not from any fixed R-squared threshold.

open as a page

Why can AIC and BIC select different models from the same 10,000-row dataset?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Both add a parameter penalty to the same fit term, but AIC charges 2 per parameter while BIC charges log n - about 9.2 at n = 10,000. BIC is far stricter there, so it favours the smaller model.

open as a page