What does the R-squared of a fitted multiple regression tell you about the model?
answer
- compares the model to a flat mean line
- one minus a ratio of two sums of squares
- residual variation over total variation
- unitless, and measured on the fitted data
basics
~20 sR-squared is the share of the outcome's total variation that the fitted model accounts for, computed as 1 minus the residual sum of squares over the total sum of squares. It measures fit on the data used, not correctness.
solid answer
~40 sR-squared compares the model's residual sum of squares, `RSS = sum of (y - y_hat)^2`, with the total sum of squares, `TSS = sum of (y - y_bar)^2`, as `R2 = 1 - RSS/TSS`. It answers one question: how much better does the fitted model track the outcome than always predicting the outcome's mean? For least squares with an intercept it sits between 0 and 1 and is unitless, so it travels across datasets in a way raw error sums do not. What it does not tell you is whether the functional form is right, whether an important predictor is missing, or whether any coefficient is causal. A very high value can come from a predictor that is the outcome in disguise, and a low value can still sit on a useful model.
go deeper
Be ready to state the formula in words - one minus residual variation over total variation - and to say plainly that it is a share of variation, never an average percentage error and never a p-value.
Explain why the comparison baseline is the outcome's mean, why the statistic is unitless, and why it cannot be compared across models that predict different outcome variables.
Show that you interrogate a suspiciously high value before reporting it, and that you can defend a small but real R-squared when the outcome is genuinely noisy.
Own the framing with stakeholders: agree in advance what fit level would change a decision, and push back on anyone treating R-squared as the single acceptance criterion for a model.
## The definition Every regression sits against a reference point. The laziest defensible predictor of an outcome `y` is its own mean `y_bar`: predict the same number for every row. How badly that predictor does is measured by the **total sum of squares**, `TSS = sum (y_i - y_bar)^2`. A fitted model instead produces predictions `y_hat_i`, leaving the **residual sum of squares**, `RSS = sum (y_i - y_hat_i)^2`. R-squared, also called the coefficient of determination, is the fraction of the lazy predictor's squared error that the model removed: `R2 = 1 - RSS/TSS = (TSS - RSS)/TSS` An R-squared of 0.30 says the fitted model wiped out 30 percent of the squared error that predicting the mean every time would leave behind. ## Why it is bounded and unitless For ordinary least squares fitted with an intercept, the fitted model can never do worse in-sample than the mean, because the mean-only fit is itself a special case the estimator could have chosen. So `RSS <= TSS`, and R-squared lands in `[0, 1]`. Both sums of squares carry the squared units of the outcome, and they are divided by each other, so the ratio has no units at all. That is exactly why R-squared is quoted and RSS usually is not: an RSS of 4.2 million means nothing until you know whether the outcome is measured in cents or in millions of euros, whereas 0.62 is readable on its own. ## What it does and does not establish R-squared is a summary of residual **size**, and nothing else. It is silent on: - **Functional form.** A curved relationship fitted with a straight line can still post a respectable R-squared while the residuals show obvious structure. - **Causality.** Explained variation is a statement about co-movement in the sample, not about what happens if you intervene on a predictor. - **Omitted variables.** A model missing the true driver can still explain a great deal via a correlate of it, and its coefficients can be badly biased while R-squared looks fine. - **Whether the value is impressive.** That depends entirely on the domain. In physical measurement, 0.95 is routine; in human behaviour, 0.10 can be the ceiling anyone achieves. ## Two directions of misreading The first misreading is treating a low value as automatic failure. If the outcome is close to individually unpredictable - which customer cancels next month, which visitor converts - most of the variation is idiosyncratic and no model will claim it. A model that explains a few percent while producing stable, well-sized coefficients can still be the basis of a real decision. The second misreading is treating a very high value as automatic success. In business data an R-squared near 1 is usually a defect rather than a triumph. The classic causes are a predictor computed from the outcome (a billed amount used to predict revenue), a predictor recorded after the outcome occurs, or a model carrying nearly as many free parameters as it has rows. Before reporting 0.98, find out what the strongest predictor actually is and when it becomes available. ## Things it cannot be compared across R-squared is a share of a specific total, so the total has to be the same for a comparison to mean anything. Two models of different outcomes - `y` versus `log y`, revenue versus revenue per user - have different TSS values and their R-squared numbers are not on a common scale. Likewise, comparing R-squared between two models fitted on different row sets (because one dropped rows with missing values) compares two different totals. ## How to say it in an interview A clean answer names the formula in words, names the baseline it compares against, states that it is unitless and in-sample, and then immediately gives the two-sided caution: low is not automatically bad, high is not automatically good. That last pair is what separates a candidate who has read the definition from one who has reported a model to a stakeholder.
- Can R-squared ever be negative?Not for least squares fitted with an intercept on the same data, since the fitted model can always match the mean-only fit, so `RSS <= TSS`. It can go negative for a model fitted without an intercept, or when predictions come from coefficients that were set by hand rather than least-squares fitted to the data in hand.
- Why can't you compare R-squared between a model of y and a model of log y?The two models explain different outcomes, so their total sums of squares are different totals. Explaining 60 percent of the variation in log spend and 40 percent of the variation in spend are answers to different questions. Compare within one outcome scale, or bring predictions back to a common scale before judging.
- What does an R-squared of essentially 1 usually indicate in practice?Almost always a mistake rather than a triumph. The usual causes are a predictor that is a rescaling or component of the outcome, an identifier that indexes the outcome, or nearly as many estimated parameters as observations. Trace the strongest predictor back to its definition before reporting the number.
R-squared scores the model against a lazy competitor who always guesses the average. A value of 0.30 means the model removed 30 percent of the wobble that guessing the average leaves behind.
saying these in an interview costs you the question
- Says R-squared shows whether the model is correctly specified
- Treats a high R-squared as evidence of causation
- Believes R-squared carries the units of the outcome
- Assumes a low R-squared always means a useless model
- Compares R-squared across models with different outcome variables