skip to content

In a residual-versus-fitted plot, what does a funnel widening with the predicted value mean?

level: middleimportance: should knowfreq 48%

answer

  1. look at the spread, not the average
  2. error scales with the quantity
  3. big predictions dominate the headline
  4. constant relative, not constant absolute
  5. band the error by prediction size

basics

~20 s

Residual spread grows with the size of the prediction, so error is roughly proportional to the quantity rather than constant. Any absolute-error headline is then mostly a statement about the largest predictions. Report error banded by predicted magnitude.

solid answer

~50 s

You plot the residual — actual minus predicted — against the model's own predicted value. A funnel means a narrow band of residuals at small predictions opening out at large ones, so error scales with the magnitude of what is being predicted. On a trip-duration model, five-minute trips may be wrong by half a minute and hour-long trips by ten minutes, while the relative accuracy is much the same everywhere. It matters because a global MAE or RMSE is then largely a statement about the long trips, and squared-error training will chase them too. My response is to slice error by predicted-value band and report both absolute and relative error per band, and to consider modelling on a log scale if proportional error is the real structure. A funnel is a variance pattern; residuals drifting away from zero across the range is bias, which is a different and worse problem.

go deeper

for a junior

Know that a residual is the actual value minus the prediction and that a healthy residual plot is a shapeless band around zero. Recognise that a widening fan means bigger predictions carry bigger errors.

for a middle

Explain why a funnel makes the headline absolute error mostly a statement about the largest predictions, and be able to separate a widening spread from a drifting mean. Suggest banded reporting as the immediate response.

for a senior

Show judgement about whether the funnel is intrinsic to a positive, right-skewed target or a real defect, pick the metric that matches the downstream decision, and know the artefacts — the axis trap and the zero floor — that fake the pattern.

for a principal

Own the framing question: whether the organisation should be held to constant absolute accuracy or constant relative accuracy for this quantity, and what that choice does to the objective, the reported metric and the promises made to users.

## What the plot is For every evaluation row, compute the **residual** `r = actual - predicted` and plot it on the vertical axis against the model's **fitted value**, its own prediction, on the horizontal axis. The healthy picture is a formless horizontal band: residuals scattered symmetrically around zero, with about the same vertical spread everywhere across the range of predictions. Every interesting departure from that is a shape. This is really a continuous form of slicing. Binning predictions into deciles and reporting the metric per bin gives the same information in a table, which is what goes into a written evaluation; the plot is how you find it. ## The funnel A funnel — tight near the left, fanning out to the right — says the **magnitude of the error grows with the magnitude of the prediction**. The average residual can stay at zero throughout, so the model is not necessarily biased; its precision is simply not uniform. This is extremely common, and often it is the honest structure of the problem rather than a defect. Durations, counts, prices, demand and spend are positive, right-skewed quantities whose noise is naturally multiplicative: a long trip has more opportunities to go wrong, so being ten percent out on an hour is ten times as many minutes as being ten percent out on six minutes. Where the underlying process is count-like, the variance grows with the mean by construction. ## Why it changes how you read the headline Three consequences follow directly. 1. **The headline metric belongs to the big end.** If the largest-prediction decile carries most of the absolute error, an MAE improvement is mostly an improvement on that decile, and can coexist with the small-prediction rows getting worse. 2. **Training follows the same gradient.** A squared-error objective weights large residuals hardest, so the fit spends its capacity where the residuals are already largest. 3. **Absolute and relative accuracy disagree.** The model may be uniformly accurate in percentage terms and wildly non-uniform in minutes. Which one the user experiences depends on the decision downstream: a promise shown to a customer cares about minutes; a capacity plan may care about percentages. ## What to do about it - **Report by band, always.** Split predictions into bands or deciles and publish the count, the absolute error and the relative error for each. This is the leaf-level answer and it costs nothing. - **Match the metric to the decision.** If the cost of an error really is proportional to size, an absolute headline is the wrong summary and a relative one is closer. - **Consider modelling the target on a log scale**, or with a multiplicative form, when the error genuinely is proportional. This stabilises the spread rather than merely reporting it, but it changes what the model optimises, so verify on the original scale. - **Do nothing but document**, when the funnel is intrinsic and the bands are all acceptable. Not every pattern is a bug. ## Funnel versus curvature versus clumps The same plot shows three different diseases and they need distinguishing: - **Funnel**: mean residual near zero everywhere, spread increasing. A variance pattern; the model is less precise on big values. - **Curvature**: the mean residual is systematically positive in one part of the range and negative in another — for instance the model under-predicts the smallest and largest values and over-predicts the middle. That is **bias**: a missing nonlinearity, a missing interaction, or a model that is too constrained. It is the more serious finding, because it is a correctable and systematic error rather than an unavoidable spread. - **Clumps or stripes**: distinct blobs sitting off the main band usually mark a cohort — a product line, a region, a data-source — behaving differently. That is the cue to switch from the plot to a named-cohort table. ## Two traps **Plot against the prediction, not the actual value.** Residual `y - yhat` contains `y`, so plotting residuals against `y` produces an upward trend even when the model is unbiased. That trend is an artefact of the axis choice, not a finding. **Watch the floor.** If the target cannot be negative, the residual satisfies `r >= -predicted`, so at small predictions residuals simply cannot be very negative. That alone creates an asymmetric narrowing at the left that can look like a funnel. Check whether the spread on the positive side widens too before concluding anything. ## What interviewers listen for A correct definition of residual and fitted value; the ability to separate a variance pattern from a bias pattern in the same picture; the instinct to convert the plot into a banded error table that a stakeholder can read; and the judgement not to treat every funnel as a defect requiring a fix.

  • How do you tell a funnel from curvature in the same plot?
    Track the average residual across the range, not just the spread. If the mean stays near zero while the band widens, it is a variance pattern and the model is simply less precise on large values. If the mean drifts systematically positive in one region and negative in another, the model is biased there, usually from a missing nonlinearity, and that needs a modelling change rather than a reporting change.
  • The target cannot be negative and many predictions sit near zero. Could the funnel be an artefact?
    Partly, yes. Since the actual value is at least zero, the residual is at least minus the prediction, so near-zero predictions physically cannot produce large negative residuals. That clips the lower half of the band on the left and mimics a funnel. Check whether the positive side widens as well before treating it as real heteroscedasticity.
  • What would you actually put in a stakeholder report from this plot?
    A table rather than the scatter: predicted-value bands down the rows, and for each band the row count, the absolute error and the relative error. Then one sentence saying where accuracy stops being good enough for the decision the model feeds, so the reader knows the range over which they can trust it.

A scale accurate to one percent is off by a gram on a feather and by ten kilograms on a car. Same instrument, wildly different absolute errors.

saying these in an interview costs you the question

  • Reads any residual pattern as overfitting
  • Calls a widening funnel evidence of bias
  • Plots residuals against the actual value instead of the prediction
  • Quotes one global error for a target spanning orders of magnitude
  • Assumes a funnel must be fixed rather than reported

context