skip to content

Why are regression residuals plotted against fitted values rather than against the observed outcome?

level: middleimportance: nice to knowfreq 24%

answer

  1. the null picture must be known
  2. orthogonal by construction, not by assumption
  3. y equals fitted plus residual
  4. the fake slope is one minus R-squared
  5. the same logic covers residual-versus-predictor plots

basics

~20 s

Because least-squares residuals are uncorrelated with the fitted values by construction, so any pattern there is a real signal. Residuals are correlated with the observed outcome, so that plot slopes upward even when the model is perfectly specified.

solid answer

~50 s

In an OLS fit with an intercept, the residuals are orthogonal to the fitted values: their correlation is exactly zero as an algebraic fact, not as a modelling assumption. That gives the residual-vs-fitted plot a clean null picture — a flat band — so any curve or funnel you see is genuine evidence. Residuals are not orthogonal to the observed outcome. Since `y = yhat + e` and `yhat` and `e` are uncorrelated, `Cov(y, e) = Var(e)`, so regressing the residuals on the observed outcome gives a slope of exactly `1 - R-squared` and a correlation of `sqrt(1 - R-squared)`. With an R-squared of 0.36 that is a slope of 0.64 and a correlation of 0.8 — a dramatic upward trend produced by nothing but the algebra. Reading it as under-prediction of large values is a classic self-inflicted wound.

go deeper

for a junior

Remember the rule and the reason in one line: plot residuals against fitted values, because plotting them against the observed outcome produces an upward trend even for a good model.

for a middle

Derive it. Show that fitted values and residuals are orthogonal in an OLS fit, so the correlation between the residuals and the observed outcome is the square root of one minus R-squared, and quote the fake slope of one minus R-squared.

for a senior

Use the principle rather than reciting it: know that every diagnostic plot needs a known null appearance, that the same orthogonality validates residual-versus-predictor plots, and that on held-out data flatness is no longer guaranteed and therefore worth testing.

for a principal

Set the standard that a diagnostic is only worth showing if its appearance under a correct model is known. Push back on dashboards whose panels manufacture patterns and then invite teams to act on them.

## The algebra behind the choice Fit a linear model by least squares with an intercept. Write each observation as `y_i = yhat_i + e_i`, where `yhat_i` is the fitted value and `e_i` the residual. Two facts follow from the normal equations, independently of whether the model is any good: - The residuals sum to zero. - The residuals are orthogonal to the fitted values and to every included predictor, so `Cov(yhat, e) = 0`. The second fact is why the residual-vs-fitted plot is the diagnostic of choice. It has a **known null appearance**: if the model's assumptions hold, the plot is a flat, even band, and it is flat by construction rather than by luck. Any systematic departure is therefore informative. ## What happens if you use the observed outcome instead Use the decomposition. Because `yhat` and `e` are uncorrelated, `Var(y) = Var(yhat) + Var(e)` and `Cov(y, e) = Cov(yhat + e, e) = Var(e)`. So if you regressed the residuals on the observed outcome, the slope would be `Cov(y, e) / Var(y) = Var(e) / Var(y) = 1 - R-squared`, and the correlation would be `Cov(y, e) / (sd(y) * sd(e)) = sd(e) / sd(y) = sqrt(1 - R-squared)`. Both are positive for any model that does not explain everything. A model with `R-squared = 0.25` gives a fake slope of 0.75 and a correlation of 0.87. The plot will look like a strong, tidy upward trend — and it means nothing. It is present for a perfectly specified model with normal, constant-variance, independent errors. The misreading it invites is specific and tempting: 'the model under-predicts high values and over-predicts low ones'. Someone then adds terms, transforms the outcome, or splits the data by outcome level, chasing a pattern that the arithmetic guarantees. ## The same reason makes residual-versus-predictor plots valid Orthogonality holds for every predictor included in the model, not only for the fitted values. That is why plotting residuals against each predictor in turn is the standard way to localise curvature: those plots also have a flat null, so a visible arc is real evidence about that variable. It is also why plotting residuals against a variable you did **not** include is informative in a different way. There is no orthogonality to protect you, so a clear relationship is evidence that the variable belongs in the model. ## Is a plot of observed against predicted useless, then? No — it is just a different tool. Observed against predicted is a **calibration** view: it shows how closely predictions track reality and over what range, and it is easy to explain to non-specialists. Two notes on reading it. First, in-sample, the least-squares line of `y` on `yhat` has slope exactly 1 and passes through the mean, again by construction, so a slope near 1 there is not evidence of anything. Second, the vertical scatter around the diagonal is the residual, which is exactly what the residual-vs-fitted plot displays with the diagonal flattened out — and flattening it is what makes changes in level and spread visible instead of being hidden along a steep line. ## Held-out data The orthogonality above is in-sample algebra: it comes from the fitting procedure. On new data the coefficients are fixed in advance, and there is no algebraic guarantee that the held-out residuals are uncorrelated with the predictions. If the model is genuinely right, they will be approximately uncorrelated because the expected error is zero everywhere. If the model is wrong — drifted, mis-specified, fitted on a different population — a trend appears. That makes a residual-versus-prediction plot on held-out data a genuinely sharper test than the same plot in-sample: in-sample, some structure was removed by the fitting itself, whereas out of sample nothing has been removed and the plot has to earn its flatness. ## How to say it in an interview Give the reason and the number: residuals are orthogonal to the fitted values by construction so the null picture is flat, while their correlation with the observed outcome is `sqrt(1 - R-squared)`, which is large for any ordinary model. Then add that the same orthogonality is why residual-versus-predictor plots are valid diagnostics, and that on held-out data the flatness is no longer automatic — which is precisely what makes it worth checking there.

  • What exactly is the slope if you regress OLS residuals on the observed outcome?
    Exactly `1 - R-squared`. Because the residuals are orthogonal to the fitted values, `Cov(y, e) = Var(e)`, and dividing by `Var(y)` gives the residual share of the outcome's variance. At an R-squared of 0.5 the spurious slope is 0.5 and the correlation is about 0.71.
  • Are residual-versus-predictor plots subject to the same problem?
    No, not for predictors that are in the model: least squares makes the residuals orthogonal to each of them too, so those plots also have a flat null and any arc is real. Plotting against a variable you excluded is different, and a visible pattern there is evidence that it belongs in the model.
  • Does the orthogonality still hold on held-out data?
    Not as algebra. The coefficients were fixed on the training data, so nothing forces held-out residuals to be uncorrelated with held-out predictions. Approximate flatness there is evidence the model is right, which makes the out-of-sample version of the plot a stricter test than the in-sample one.
  • Is a plot of observed against predicted values worth drawing at all?
    Yes, as a calibration view: it shows how well predictions track reality across the range and reads well to non-specialists. Just do not use it as a residual diagnostic, and remember that in-sample the least-squares line of observed on predicted has slope exactly one by construction.

Plotting residuals against the observed outcome is like grading a race by finishing time when the handicap was calculated from that same time: the trend you see was built in before anyone ran.

saying these in an interview costs you the question

  • Reads the upward slope against observed y as model bias
  • Believes residuals and fitted values are correlated in-sample
  • Cannot name what least-squares residuals are orthogonal to
  • Thinks the same trap applies to residual-versus-predictor plots
  • Claims the orthogonality still holds exactly on held-out data

context