Why report a regression model's R-squared on held-out rows instead of the rows it was fitted on?
answer
- the fit already saw those rows
- residuals were minimised on that data
- one of the two has a floor
- in-sample cannot lose to the mean
- held-out has no such floor
basics
~20 sR-squared on the fitting rows is measured on data whose residuals the fit already minimised, so it flatters the model. Held-out R-squared scores rows the model never saw, the only honest estimate of future performance.
solid answer
~50 sR-squared is `1 - SS_res / SS_tot`: the model's squared error divided by the squared error of just predicting the mean of the target. On the fitting rows those coefficients were chosen precisely to make `SS_res` small, so the number is optimistic by construction; for an ordinary least-squares fit with an intercept it cannot even fall below 0, because the mean-only model is one of the candidates the fit could have picked. On held-out rows nothing pins it, so it can land anywhere, including below zero. A sensor-calibration model reading 0.92 on its own fitting data and 0.31 when scored at a second plant is the normal shape of this: the 0.31 is the number that describes the model. I would report the held-out figure, say which rows produced it, and put an absolute error in the target's units next to it.
go deeper
Be ready to say, in one sentence, that the model was fitted to minimise error on the training rows so its score there is flattering, and that the number you quote is the one from rows it never saw.
Explain the mechanics: R-squared is one minus the model's squared error over the mean baseline's, and an intercept-bearing least-squares fit cannot lose to that baseline on its own fitting rows, which is why the in-sample figure has a floor of zero and the held-out one does not.
Show you know a hold-out score is only as meaningful as the rows behind it. When a number drops on a new site or period, expect to be asked which part is generalisation failure and which part is a different population, and to name the extra hold-out that separates them.
Own the reporting convention. Decide which rows are the standard hold-out, whether the denominator uses the hold-out's own mean, and what must be quoted alongside, so two teams' numbers for the same model mean the same thing.
## What R-squared is comparing R-squared is not an absolute accuracy score. It is a ratio between two errors: ``` R2 = 1 - SS_res / SS_tot SS_res = sum over scored rows of (y - yhat)^2 SS_tot = sum over scored rows of (y - ybar)^2 ``` `yhat` is the model's prediction, `ybar` is the mean of the target over the rows being scored. `SS_tot` is therefore the squared error of the dumbest reasonable predictor: one constant, the mean, for everybody. R-squared of 1 means zero error; 0 means you tied that constant; below 0 means you lost to it. The interesting question is always *which rows* you plug into those two sums. ## Why the fitting rows flatter the model A least-squares fit picks its coefficients by minimising `SS_res` on the training rows. That is the objective — the number you later read off as R-squared is the very quantity the fitting procedure drove down. Scoring on those rows is grading an exam with the answer key the student wrote. Two consequences follow. First, an ordinary least-squares fit with an intercept can never score below 0 in-sample: predicting `ybar` for everyone is itself a member of the family the fit searched over, so the chosen coefficients cannot be worse than it on those rows. In-sample R-squared therefore lives in [0, 1] and its lower end is not evidence of anything. (This is a theorem for OLS with an intercept; for penalised or constrained fits it is not guaranteed, though in practice it holds.) Second, the more flexible the model, the more of the training rows' *noise* it absorbs into `SS_res`, and the higher the in-sample figure climbs without any of that reflecting structure that will exist in new data. Held-out rows took part in none of that minimisation. Nothing forces the model to beat their mean, so the held-out figure is free to be small, or negative, and that freedom is exactly what makes it informative. ## A worked shape A sensor-calibration model reads R-squared 0.92 on the rows it was fitted on and 0.31 when the same fitted model is scored at a second plant. Which number describes the model? The 0.31. The 0.92 answers "how closely did the fit trace the data it was handed", a question nobody deploying the model cares about. The 0.31 answers "on rows this model has never seen, did it beat predicting the average reading", and the answer is a thin yes. ## The caveat that makes the second plant awkward The 0.31 changed two things at once: the rows are unseen *and* they come from a different population, with possibly different sensor drift, different operating range, and — importantly — a different spread of the target, which changes `SS_tot` in the denominator. You cannot tell from those two numbers alone how much of the drop is the model failing to generalise and how much is the second plant being a different problem. The fix is a third number: hold out rows from the first plant too. If the first-plant hold-out reads about 0.9, the model generalises fine and the second plant is a shifted population; if it reads about 0.35, the 0.92 was mostly self-congratulation. ## Two details worth carrying **Which mean goes in the denominator.** The usual convention on a held-out set is that set's own mean. That is the harder baseline, because no constant beats a set's own mean on it. Some teams use the training mean instead — arguably more deployable, since that is the constant you would actually have available at prediction time — but it enlarges `SS_tot` and so returns a *higher* number. Different conventions give different scores for the same model, so say which one you used. **"Fraction of variance explained" stops being literally true.** That phrasing comes from the decomposition `SS_tot = SS_model + SS_res`, which holds for an OLS fit with an intercept *on its own fitting rows*, where the residuals sum to zero by construction. On held-out rows the residuals need not have mean zero, the decomposition does not hold, and the number is best read as what the formula actually says: how your squared error compares to the mean baseline's on these rows. ## What to report Quote the held-out figure, name the rows it came from, and pair it with an error in the target's units — degrees, months, dollars — so a reader who does not know the target's spread can still tell whether the model is useful.
- Can a held-out R-squared ever come out higher than the one on the fitting rows?Yes, and it is usually not good news. A small hold-out set is noisy, an easier slice of rows can land in it by chance, and a hold-out whose target spread is wider than the training set's inflates the denominator and lifts the score. None of those mean the model generalises better than it fits; they mean the estimate is unstable or the two sets are not comparable.
- The second plant reads 0.31 — how would you tell overfitting apart from a different population?Score a hold-out drawn from the first plant as well. Three numbers instead of two: fitting rows, same-plant hold-out, second plant. If the same-plant hold-out stays near 0.9, the model generalises and the second plant is genuinely a different distribution. If it drops with the second plant, the in-sample 0.92 was never real.
- Is it fair to call a held-out R-squared 'the fraction of variance explained'?Loosely, and it invites confusion. The variance-decomposition identity behind that phrase holds for a least-squares fit with an intercept on its own fitting rows, where residuals sum to zero. Out of sample the residuals can carry a bias, the decomposition breaks, and the safer reading is the literal one: your squared error relative to the mean baseline's on these rows.
Grading an exam against the answer key the student wrote themselves. The score is real arithmetic and tells you nothing about the next exam.
saying these in an interview costs you the question
- Reports the fitting-rows R-squared as the model's accuracy
- Believes R-squared is always between zero and one
- Treats a high in-sample fit as proof the model generalises
- Never says which rows the R-squared was computed on
- Assumes the same model gives the same R-squared everywhere