skip to content

In regression, how does a prediction interval differ from a confidence interval for the mean response?

level: middleimportance: must knowfreq 72%

answer

  1. two different questions, same centre
  2. average case versus one case
  3. coefficient error plus new-case error
  4. the extra 1 under the square root
  5. one shrinks with n, one has a floor

basics

~20 s

A confidence interval for the mean response covers the average outcome at a given predictor value. A prediction interval covers one new individual case, so it is always wider: it adds the irreducible spread of a single observation.

solid answer

~50 s

Both are centred on the same fitted value `y_hat(x0)` but answer different questions. The confidence interval for the mean response asks where the *average* outcome sits among all cases with predictor value `x0`; its width comes only from coefficient uncertainty, so it shrinks toward zero as the sample grows. The prediction interval asks where a *single new* outcome lands, so it carries that plus the error of that one case. In simple regression the half-widths are `t * s * sqrt(1/n + (x0 - xbar)^2 / Sxx)` versus `t * s * sqrt(1 + 1/n + (x0 - xbar)^2 / Sxx)` — the extra `1` is the new case's own variance. Concretely: the average delivery time across all 10 km orders might be 31 to 35 minutes, while the next single 10 km order sits in 18 to 48. With unlimited data the first collapses to a point; the second never narrows past about two residual standard deviations.

go deeper

for a junior

Be ready to say which of the two is wider and why, in one sentence: the prediction interval covers one new case, so it also carries that case's own random error.

for a middle

Expect to write both half-widths and point at the extra 1 under the square root as the new observation's variance. Explain why only the mean-response interval shrinks toward zero as the sample grows.

for a senior

Show you pick the interval from the decision: aggregate planning gets the mean-response interval, per-case commitments get the prediction interval. Mention that skewed residuals damage the prediction interval and are not fixed by more data.

for a principal

Own the framing question of what your team publishes by default. Decide whether dashboards and model outputs ship point estimates, mean-response bands, or per-case intervals, and be able to defend the choice against the pressure to show tight numbers.

## Two questions, one fitted line After fitting a regression you can ask a value `x0` two very different questions: 1. **What is the average outcome for cases with this predictor value?** That is a statement about a *parameter* — the true mean response at `x0`, written `E[Y | X = x0]`. 2. **What will the next single case with this predictor value actually do?** That is a statement about a *random variable* — one future observation `Y_new`. Both are estimated by the same number, the fitted value `y_hat(x0) = b0 + b1 * x0`. The intervals around that number are not the same, because the second question has an extra source of randomness. ## Where the width comes from A new observation is `Y_new = (true mean at x0) + error`. Your prediction misses it for two independent reasons: - **Estimation error.** Your fitted line is not the true line, because the coefficients were estimated from a finite sample. This shrinks as `n` grows. - **Individual error.** Even if you knew the true line exactly, a single case scatters around it with variance `sigma^2`. This never shrinks, no matter how much data you collect. The mean-response interval prices only the first. The prediction interval prices both, and because the two sources are independent their variances add. ## The formulas, in plain notation For simple linear regression, with `s` the residual standard error, `xbar` the mean of the predictor and `Sxx = sum (xi - xbar)^2`: - Confidence interval for the mean response at `x0`: `y_hat(x0) +/- t * s * sqrt( 1/n + (x0 - xbar)^2 / Sxx )` - Prediction interval for a new case at `x0`: `y_hat(x0) +/- t * s * sqrt( 1 + 1/n + (x0 - xbar)^2 / Sxx )` The only difference is the leading `1` under the square root, and it is the whole story: it represents the variance of the new case's own error term relative to `s^2`. The multiplier `t` is the same for both, taken on the residual degrees of freedom. ## Why the difference is dramatic in practice With a decent sample, `1/n` and the leverage term are both small — say they sum to `0.05`. Then the mean-response half-width is about `t * s * 0.22` while the prediction half-width is about `t * s * 1.02`. The prediction interval is roughly **five times wider**, and the gap grows with `n`. A worked feel for it: a delivery-time model on distance, with a residual standard error of about 7.5 minutes and a fitted value of 33 minutes at 10 km. The average delivery time across all 10 km orders might be quoted as 31 to 35 minutes. The next single 10 km order is quoted as roughly 18 to 48 minutes. Both are 95% intervals from the same model; they are simply about different things. ## The limiting behaviour As `n` grows without bound, the mean-response interval collapses toward a single point: with enough data you learn the true average exactly. The prediction interval converges to approximately `y_hat(x0) +/- 1.96 * sigma` for a 95% level — a floor set by the irreducible noise in individual outcomes. If someone claims that with enough data they can predict a single order's delivery time to the minute, this is the sentence that answers them. ## Choosing between them Ask what the number is for. - **Aggregate decisions** — how many couriers to staff for a day of 10 km orders, what average revenue to book from a spend level — are about the mean, so the mean-response interval is the honest one. - **Individual commitments** — a promise to one customer, an alert threshold for one transaction, a capacity guarantee for one request — are about a single case, so the prediction interval is the honest one. Quoting the narrow interval here is the most common way regression output gets oversold. ## Assumptions worth stating Both intervals assume the model form is right, the errors are independent with constant variance, and `x0` is inside the range the model actually saw. Normality of the errors matters more for the prediction interval: for the mean response, averaging pulls the sampling distribution toward normal as `n` grows, but a single new observation is governed by the error distribution itself, so skew or heavy tails are not washed out by sample size. If the residuals are visibly skewed, a symmetric prediction interval will be wrong on both ends even with thousands of rows.

  • As the sample size grows without bound, what happens to each of the two intervals?
    The mean-response interval collapses toward a single point, because the coefficients are eventually known essentially exactly. The prediction interval converges to about the fitted value plus or minus `1.96 * sigma` at 95%, since the individual error term never goes away. That floor is the honest limit of what any amount of data buys you for a single case.
  • Which interval would you quote when promising a delivery time to one customer?
    The prediction interval, because the promise is about one order, not about the long-run average of orders like it. Quoting the mean-response interval there would advertise a range that most individual orders fall outside of, and the failures would show up as broken promises rather than as statistical error.
  • Which of the two is more sensitive to non-normal residuals, and why?
    The prediction interval. The mean-response interval concerns an average, and averaging drives the sampling distribution toward normal as the sample grows, so mild skew washes out. A prediction interval is about one draw from the error distribution itself, so skew or heavy tails distort it no matter how large the sample is.

Predicting the average height of all 30-year-olds in a city is far easier than predicting the height of the next one who walks through the door. Both guesses start at the same number; only the second has to cover individual variation.

saying these in an interview costs you the question

  • Treats the two intervals as the same thing under different names
  • Says a 95% mean-response interval covers 95% of individual cases
  • Expects the prediction interval to shrink to zero given enough data
  • Quotes the narrow interval when committing to a single case
  • Forgets that both intervals share the same centre

context