skip to content

Why is a posterior predictive interval for one new observation wider than the credible interval for the parameter?

level: middleimportance: must knowfreq 58%

answer

  1. different questions, different scales
  2. law of total variance
  3. one term shrinks, one does not
  4. noise floor survives infinite data

basics

~20 s

A predictive interval must cover a single random observation, so it carries the model's sampling noise on top of the uncertainty about the parameter. A parameter interval carries only the second piece, and shrinks toward zero width as data accumulate.

solid answer

~50 s

The two intervals answer different questions and live on different scales. A credible interval says where the parameter plausibly is; a predictive interval says where the next observed value plausibly lands. The law of total variance splits the predictive variance into two terms: `Var(y_new | data) = E[Var(y_new | theta)] + Var(E[y_new | theta])`. The first term is the irreducible noise the model says data have even when the parameter is known; the second is the parameter uncertainty that the credible interval alone captures. Only the second term shrinks with sample size, so with a large dataset the predictive interval converges to the noise floor while the credible interval keeps narrowing toward a point. Quoting a parameter interval when a stakeholder asked what the next observation will be is one of the most common and most costly interview mistakes.

go deeper

for a junior

Remember the one-line reason: predicting a single observation includes the randomness of that observation, which estimating a parameter does not. Do not quote a parameter interval as a forecast.

for a middle

Write the variance decomposition and name both terms. Be ready to say which term shrinks with sample size and which is fixed by the likelihood.

for a senior

Diagnose an under-covered forecast in production: recognise the symptom of a parameter interval being reported as a prediction, and know how the horizon shifts which term dominates.

for a principal

Set the reporting convention for the org — which decisions get a parameter interval and which get a predictive one — and be able to defend the extra width to stakeholders who read wide intervals as weak analysis.

## Two different questions A **credible interval** is a statement about a parameter: given the data, there is a 95% posterior probability that the conversion rate `p` lies between two numbers. A **posterior predictive interval** is a statement about data you have not seen: given the data, there is a 95% predictive probability that the next observation — or the count among the next 100 visitors — falls between two numbers. They are on different scales (a rate versus a count, a mean versus an individual measurement) and they cover different objects. Mixing them up produces forecasts that are confidently, reproducibly wrong. ## The variance decomposition The cleanest way to see the width difference is the **law of total variance**, conditioning on the parameter `theta` and averaging over its posterior: ``` Var(y_new | data) = E_theta[ Var(y_new | theta) ] + Var_theta( E[y_new | theta] ) = expected sampling noise + parameter uncertainty ``` **First term — sampling noise.** Even if someone told you the true parameter exactly, the next observation would still be random: a coin with a known bias still lands unpredictably. This term is a property of the likelihood, not of how much data you have. It does not shrink. Ever. **Second term — parameter uncertainty.** This is the spread of the posterior, pushed through onto the data scale. This is the *only* thing a credible interval measures, and it is the piece that shrinks roughly like `1/n` in variance as the sample grows. Because the predictive variance is the sum of a non-negative term and the parameter term, the predictive interval is always at least as wide as the parameter interval rescaled to the data scale, and in practice strictly wider. ## A worked feel for the two terms Take the conversion example: 12 conversions from 100 visitors, posterior for `p` centred near 0.127 with standard deviation about 0.033. Forecast the count among the next 100 visitors. - Parameter-uncertainty contribution on the count scale: `100 * 0.033`, about 3.3 conversions. - Sampling-noise contribution: about `sqrt(100 * 0.127 * 0.873)`, also about 3.3 conversions. - Combined (the two variances add, not the standard deviations): `sqrt(3.3^2 + 3.3^2)`, about 4.7 conversions. The near-equality is not a coincidence — you observed 100 visitors and are forecasting 100 more, so the two sources are comparable. Change the horizon and the balance changes sharply. Forecast **one** future visitor and sampling noise dominates: the parameter is known well enough that the only real question is a coin flip. Forecast **10,000** future visitors and parameter uncertainty dominates, because the sampling-noise variance grows linearly in the horizon while the parameter-uncertainty variance grows quadratically. That asymmetry is worth being able to state out loud; it is the reason long-horizon forecasts stay wide no matter how much history you have. ## The limiting behaviour As the observed sample size goes to infinity, the posterior collapses onto the true parameter, the second term vanishes, and the predictive distribution converges to the likelihood at that value. The credible interval shrinks toward zero width; the predictive interval converges to a **positive** floor set by the model's own noise. A candidate who claims 'with enough data the prediction interval also goes to zero' has confused a statement about a parameter with a statement about a random observation. ## How this shows up in practice A product manager asks 'how many conversions will we get next week?' and receives an interval that turns out to be a credible interval for the rate multiplied by traffic. That interval will be badly under-covered — actual weeks will land outside it far more than 5% of the time — and the miss will be blamed on the model rather than on the reporting. The fix is not a wider arbitrary fudge factor; it is generating replicate observations through the full two-step draw (parameter from the posterior, then data from the likelihood) and quoting quantiles of *those*. ## What to say in the interview One sentence for the framing (parameter versus observation), one for the decomposition (noise plus parameter uncertainty), one for the limit (the noise floor never disappears), and one for the consequence (never hand a parameter interval to someone who asked about future data). That answer covers the ground without hand-waving.

  • Does the predictive interval shrink to zero width as the sample size grows?
    No. The parameter-uncertainty term vanishes, but the sampling-noise term is a property of the likelihood and stays. The predictive interval converges to the width implied by the data-generating process with the parameter known exactly, which is strictly positive for any non-degenerate model.
  • How does the balance between the two variance terms change with the forecast horizon?
    Sampling-noise variance grows roughly linearly in the number of future observations, while parameter uncertainty enters through the squared horizon. So for a single future observation noise dominates, and for a long horizon parameter uncertainty dominates — which is why long-range forecasts stay wide even with a large history.
  • When would you legitimately report the parameter interval rather than the predictive one?
    When the decision is about the underlying rate itself — is the true conversion rate above 10%, is the effect positive — rather than about a specific future outcome. Capacity planning, alert thresholds and any per-case commitment need the predictive interval instead.

Knowing a factory's average defect rate to four decimal places still tells you little about whether the specific unit in your hand is defective.

saying these in an interview costs you the question

  • Says both intervals answer the same question
  • Claims the predictive interval shrinks to zero with enough data
  • Adds the two standard deviations instead of the variances
  • Uses a parameter interval to forecast next week's observed count

context