skip to content

Why can the same regression model post very different held-out R-squared on two test splits?

level: seniorimportance: should knowfreq 46%

answer

  1. look at the denominator first
  2. the scored split's own spread
  3. same errors, different variance
  4. a narrow split is a harsh grader
  5. quote months, not just a ratio

basics

~20 s

R-squared divides the model's squared error by the target's spread on the rows being scored. Narrow that spread and the same absolute errors give a much lower score, so the number is not comparable across splits.

solid answer

~50 s

The denominator of R-squared is the scored split's own variance around its mean, so the score is a ratio between the model's error and how hard the split is. Take an employee-tenure model with a root mean squared error of 5 months. On a broad hold-out whose tenure standard deviation is 12 months it reads `1 - 25/144`, about 0.83. Confine the hold-out to a single job family where the spread is 7 months and the identical model, making identical errors, reads `1 - 25/49`, about 0.49. Nothing about the model changed; the baseline got harder to beat because everyone in that split is already close to the mean. This is why I never compare R-squared across splits, periods or projects, and why I always quote it with an absolute error in the target's units plus the split's target spread. The months are portable; the ratio is not.

go deeper

for a junior

Remember that R-squared is a ratio, not an accuracy: the model's error divided by the spread of the target on the rows being scored. Always report an error in the real units next to it.

for a middle

Explain the arithmetic of the denominator — mean squared error over the scored split's target variance — and work an example where identical errors give very different scores because the spread differs.

for a senior

Demonstrate the diagnostic reflex: when a score moves and the model did not, check the target's spread on both splits first, and resolve two teams' disagreement by freezing one hold-out rather than auditing the arithmetic.

for a principal

Own the metric policy. A single organisation-wide R-squared threshold rewards teams with high-variance targets and is gameable through split choice; decide what the headline number is and what must be reported with it.

## The denominator belongs to the data, not the model ``` R2 = 1 - SS_res / SS_tot = 1 - (mean squared error) / (variance of the target on the scored rows) ``` Only the numerator is about the model. The denominator is a property of the rows you chose to score on. Change those rows and you change the score without touching a single coefficient. ## A worked case An employee-tenure model predicts months of tenure with a root mean squared error of 5 months — the same 5 months in both scenarios below. - **Broad hold-out**, employees drawn from across the company, tenure standard deviation 12 months. Variance is 144. `R2 = 1 - 25/144 = 0.83`. - **Narrow hold-out**, restricted to one job family where nearly everyone leaves in a similar window, tenure standard deviation 7 months. Variance is 49. `R2 = 1 - 25/49 = 0.49`. Same model, same predictions, same errors in months. The reported quality collapses because on the narrow split the mean was already a good answer — a constant is only 7 months off on average there, so beating it is a much higher bar. It runs the other way too. Deliberately or accidentally include a few extreme rows and the target spread widens, the denominator inflates, and the same model looks better. A hold-out with wide spread is a generous grader. ## Why this matters in a real review Two teams score the same fitted model and report 0.71 and 0.44. Neither made an arithmetic mistake. They held out different periods, and the two periods had different target spreads — perhaps one covered a reorganisation and the other a quiet quarter. The disagreement is not about the model at all, and an hour spent auditing the code finds nothing, because there is nothing to find. The resolution is procedural: agree one frozen hold-out that both teams score against, or compare in the units of the target instead. The same argument kills cross-project comparison. R-squared 0.6 on a noisy human-behaviour target can be a strong result; 0.6 on a sensor-calibration target where the physics is nearly deterministic is a failure. There is no threshold that transfers between problems, because the denominator is a different quantity in each. ## Pair it with an absolute error The practical habit is to never report a held-out R-squared alone. Report three things: 1. The R-squared, so a reader sees the comparison against the constant baseline. 2. A mean absolute error or root mean squared error **in the target's units** — months, dollars, degrees. This is the number that keeps its meaning when the split changes, and the one a stakeholder can weigh against the cost of being wrong. 3. The target's mean and spread on the scored rows, which is the denominator made visible. With those three, the narrow-split story reads correctly at a glance: "R-squared 0.49, MAE 4 months, hold-out spread 7 months" says clearly that the model is off by about four months on a population that varies by seven — modest lift, honestly described. The bare 0.49 says nothing you can act on. ## What R-squared is still good for This is not an argument to stop using it. Within one fixed hold-out, R-squared is a clean way to compare candidate models against each other and against the do-nothing baseline, and its zero point carries real meaning: below it you should ship the constant. The rule is narrow and absolute: compare R-squared **only** between models scored on the same rows. ## Diagnostic reflex When a held-out R-squared moves and you did not change the model, look at the denominator before you look at anything else. Compute the target's standard deviation on both splits. If it moved, the score's movement is at least partly arithmetic, and the model's error in units will tell you whether anything real happened. If the spread is stable and the score still moved, only then is there a modelling question to answer.

  • Two teams report different held-out R-squared for the same fitted model because they held out different periods — who is right?
    Both computations are right and the comparison is meaningless. Fix it procedurally rather than arithmetically: freeze one hold-out that every team scores against, and require an absolute error in the target's units plus the split's target spread with any reported R-squared. Then a difference between two numbers is a difference between two models, which is the only comparison worth having.
  • Should held-out R-squared be your organisation's single headline metric for regression models?
    No. It is not comparable across projects, so a company-wide threshold like 'ship above 0.7' rewards teams with high-variance targets and punishes teams with easy baselines, and it is quietly gameable by choosing a wider hold-out. Make the headline an error in the target's units against a stated business tolerance, and keep R-squared as the supporting check that you beat the constant.
  • Does a low held-out R-squared always mean the predictions are far off?
    No. It means they are not far enough ahead of the mean on that split. A model can be off by four months on a population whose whole spread is seven months — genuinely useful for some decisions — and still score under 0.5. Only the error in units answers 'how wrong is it'; R-squared answers 'how much better than a constant'.

Grading on a curve. The same exam paper earns a different grade depending on how spread out the rest of the class is, even though the answers never changed.

saying these in an interview costs you the question

  • Compares R-squared across datasets or projects
  • Reports R-squared with no error in the target's units
  • Blames the model when the split's target spread changed
  • Assumes a low R-squared means large prediction errors
  • Sets one company-wide R-squared threshold for shipping

context