What does a negative R-squared on a held-out test set tell you about the model?
answer
- compare against the simplest predictor
- what does exactly zero mean here
- the mean of the scored rows
- your squared error exceeded the baseline's
- nothing bounds it below out of sample
basics
~10 sA negative held-out R-squared means the model's squared error on those rows exceeds what predicting one number, the split's mean, for every row would give. The model lost to a constant.
solid answer
~50 sR-squared is `1 - SS_res / SS_tot`, where `SS_tot` is the squared error of predicting the mean of the target for every scored row. A value of -0.08 says `SS_res` is eight percent larger than `SS_tot`: on this split the model is worse than a constant, and nothing bounds the score below zero once you leave the fitting rows. When an employee-tenure model comes back at -0.08 on a held-out cohort I check four things in order: whether the split is small enough that -0.08 is noise around zero, whether the predictions are systematically shifted, whether the population moved, and whether the target carries learnable signal at all. Beating the mean there just requires the model's mean squared error to fall below the cohort's target variance. I quote the error in months too, because -0.08 alone does not say whether the misses are two months or twenty.
code
python · 12 linesactual = [8, 12, 30, 45, 5] # tenure in months, held-out cohort
pred = [15, 22, 18, 20, 25] # what the model said
mean_actual = sum(actual) / len(actual)
ss_res = sum((a - p) ** 2 for a, p in zip(actual, pred))
ss_tot = sum((a - mean_actual) ** 2 for a in actual)
r2 = 1 - ss_res / ss_tot
print(mean_actual) # 20.0 the baseline answers 20.0 for everyone
print(ss_res) # 1318 the model's squared error
print(ss_tot) # 1158 the baseline's squared error
print(round(r2, 3)) # -0.138 the model lost to the constantgo deeper
Recall the zero point: R-squared of 0 means the model tied a single constant, the mean of the scored rows. Below zero it lost to that constant. Say so plainly rather than guessing the score must be a bug.
Explain the arithmetic — minus 0.08 means squared error eight percent above the mean baseline's — and why the out-of-sample score has no floor while the in-sample one does. Be ready to name a couple of causes, including a systematic offset in the predictions.
Show a diagnostic order: noise on a small split, a shifted mean prediction, a moved population, or no signal. Convert the sign into a concrete bar — mean squared error below the split's target variance — and be willing to recommend shipping the constant.
Decide what happens organisationally when a model loses to its baseline: whether the project stops, whether the mean becomes the shipped predictor, and what stops teams from quietly re-splitting until the number turns positive.
## The number, read literally ``` R2 = 1 - SS_res / SS_tot SS_res = sum of (y - yhat)^2 over the scored rows <- the model's error SS_tot = sum of (y - ybar)^2 over the scored rows <- the mean baseline's error ``` Rearrange: `R2 = -0.08` means `SS_res / SS_tot = 1.08`. Your model's squared error is eight percent bigger than the error of a model that ignores every feature and answers with the same number, the mean of the target on that split, every time. That is the whole content of the sign. Negative does not mean "explains negative variance" — the phrase is meaningless. It means *lost to the constant*. ## Why it is possible at all On the rows an ordinary least-squares fit with an intercept was trained on, the mean-only model is one of the candidates the fit searched over, so the fit cannot do worse than it, and the score has a floor of zero. Held-out rows were never part of that minimisation. Nothing keeps the model's predictions anywhere near those rows' mean, so `SS_res` can exceed `SS_tot` by any amount. R-squared out of sample is bounded above by 1 and unbounded below. ## Four things that produce it **Noise.** On a small held-out cohort, -0.08 and +0.02 are the same finding: the model carries no usable signal on this split. Do not narrate a story about the sign until you know the estimate's spread; re-score on a different hold-out and see how far it moves. **A systematic offset.** Predictions that track the shape of the target but sit consistently high or low destroy R-squared quickly, because squared error punishes a constant bias on every row. A tenure model fitted on a cohort averaging 30 months and scored on one averaging 14 will be wrong by roughly 16 months everywhere before it gets a chance to be right about anybody. Plotting predicted against actual, or just comparing the mean prediction to the mean actual, catches this in a minute. **A moved population.** The relationship the model learned no longer holds on the scored rows — different job market, different product, different instrument. The model is not broken; it is being asked a different question. **No signal.** Sometimes the target really is close to unpredictable from the available features and the fitted coefficients are fitting noise. Then the model must lose to the mean out of sample, because noise does not replicate. ## What beating the mean would have required Stated as a threshold: the model's mean squared error on that cohort has to be smaller than the cohort's target variance. If held-out tenure has a standard deviation of 11 months, the variance is 121, so the model needs a root mean squared error under 11 months just to reach R-squared 0. That reframing is useful in a review — it converts an abstract sign into a concrete accuracy bar in months, and often makes it obvious that the bar was never realistic with the features on hand. ## Which mean goes in the denominator By convention, the held-out set's own mean. That is the strictest constant baseline, since no other constant beats a set's own mean on it — which is why negative values show up at all. Using the training mean instead makes `SS_tot` larger (the test rows are further from a mean computed elsewhere), which *raises* the reported score. Both conventions appear in practice; the only real error is not saying which one you used. ## What to do about it Do not file a bug. Do report the number, alongside the mean absolute error in the target's units and the target's spread on that split, so a reader can see both that you lost to the constant and by how much in real terms. Then act on the diagnosis rather than the sign: correct an offset by refitting on data that matches the scoring population, treat a genuine population shift as a scoping problem rather than a modelling one, and be willing to conclude that predicting the mean — cheap, stable, explainable — is the right thing to ship for this cohort. ## A trap to avoid A negative held-out score sometimes gets "fixed" by shopping for a hold-out split on which the number turns positive. That is not a fix; it changes the denominator and the population, not the model. If you re-split, re-split for a stated reason, and report every score you computed.
- Is a held-out R-squared of -0.08 different in kind from one of +0.02?No. Both say the model carries essentially no usable signal on that split; one lost to the constant by a hair, the other beat it by a hair. On a small hold-out the difference is well inside the noise. Treating the sign as a separate failure mode leads people to chase a bug that is not there instead of asking whether the features predict the target at all.
- Which mean belongs in the denominator of a held-out R-squared — the training mean or the test mean?The test set's own mean is the usual convention and the harder baseline, because no constant beats a set's own mean on it. Using the training mean pushes the scored rows further from the baseline, enlarges the denominator, and so reports a higher number. Neither is wrong; quoting a score without saying which you used is.
- How much accuracy would the tenure model need to reach R-squared zero on that cohort?Its mean squared error has to fall below the cohort's target variance. With a held-out standard deviation of 11 months, that is a root mean squared error under 11 months. Stating the bar in months usually settles the argument faster than the ratio does — it shows immediately whether the target is reachable with the features available.
saying these in an interview costs you the question
- Calls a negative R-squared a computation bug
- Believes R-squared cannot go below zero
- Reads minus 0.08 as explaining negative eight percent of variance
- Declares the target unpredictable without checking for an offset
- Re-splits the data until the number turns positive