An 80% prediction interval covers only 55% of held-out actuals. How do you diagnose it?
answer
- nominal is a promise, empirical is the audit
- too narrow, not biased
- do not pool the horizons
- training residuals are optimistically small
- coverage alone is gameable without width
basics
~20 sEmpirical coverage far below the nominal level means the intervals are too narrow. Recompute coverage per forecast horizon, check whether the interval width grows with horizon, and confirm the shortfall is larger than the sampling noise in the count of covered points.
solid answer
~50 sCoverage is measured by counting how many held-out actuals fall inside the interval; an 80% interval should contain roughly 80% of them. At 55% the intervals are badly too narrow, so the model is understating its own uncertainty. First check whether the gap is real: with 20 held-out points the standard error of an 80% coverage rate is about `sqrt(0.8 * 0.2 / 20)`, roughly 9 percentage points, so 55% is about 2.8 standard errors low — suspicious but worth more points. Then break coverage out by horizon rather than pooling. Intervals should widen with horizon, and a pooled 55% often hides near-nominal coverage at one step and severe under-coverage at twelve. Typical causes are interval width derived from in-sample residuals, which are optimistically small, a Gaussian assumption on skewed or heavy-tailed errors, and variance that ignores parameter and model uncertainty. The usual repair is to widen using empirical quantiles of backtest errors computed separately for each horizon, then re-measure.
go deeper
Be ready to define coverage as the share of held-out actuals falling inside the interval, and to say that far below nominal means the interval is too narrow.
Explain why in-sample residuals give optimistically small widths and why width must grow with horizon, and compute coverage per horizon rather than pooled.
Show the diagnostic sequence: rule out sampling noise, split by horizon and segment, name the likely cause, recalibrate on backtest error quantiles, then re-measure on untouched data.
Own the reporting contract. Decide that calibration and sharpness ship together, so no team can present coverage without width, and set what downstream consumers may assume about interval quality.
## What coverage means A prediction interval carries a nominal level: an 80% interval claims that a future actual falls inside it 80% of the time. Empirical coverage is the check on that claim — over a held-out set of forecast-actual pairs, the share of actuals that landed between the lower and upper bound. Nominal is a promise, empirical is the audit. Coverage of 55% against a nominal 80% is a calibration failure in the dangerous direction. The intervals are too narrow, so every downstream consumer — a safety-stock rule, a capacity plan, an alerting threshold — is being told the future is more predictable than it is. ## Step one: is the gap larger than noise? Coverage is a proportion estimated from a finite number of held-out points, so it has sampling error. If the true coverage were the nominal 80% and you scored k independent points, the standard error of the observed share is `sqrt(0.8 * 0.2 / k)`. With 20 points that is about 0.089, so an observed 55% sits roughly 2.8 standard errors below nominal — unlikely but not impossible. With 200 points the standard error drops to about 0.028 and a 55% reading is unambiguous. There is a second reason to distrust a small count: forecast errors from overlapping windows are not independent. Neighbouring origins share most of their training data and the same underlying shocks, so the effective sample size is smaller than the raw count of points. Treat a coverage estimate from a handful of origins as directional, not decisive. ## Step two: split coverage by horizon This is the highest-yield diagnostic. Uncertainty about a series grows with how far ahead you look, so an honest interval widens with the horizon. If width is roughly constant across horizons — a common symptom when the interval is built from a single residual standard deviation — then coverage will look acceptable at one step and collapse at longer horizons. Pooling hides exactly that. Report coverage as a small table: one row per horizon, the nominal level, the observed share covered, and the number of points behind it. The shape of that table tells you whether you have a uniform width problem or a horizon-scaling problem, and they have different fixes. It is also worth splitting by segment — by series, by season, by regime. Coverage that is fine on stable series and terrible on volatile ones points at a variance model that does not adapt, rather than at a global width that is too small. ## Step three: the usual causes **In-sample residuals.** Interval width is often derived from the spread of residuals on the training data. Those residuals are optimistically small: the model was fitted to minimise them. Out-of-sample errors are systematically larger, so intervals built this way under-cover from the start. **Only one source of uncertainty.** Many interval formulas capture the noise of the process while assuming the model form and its fitted parameters are correct. Parameter uncertainty, model-selection uncertainty and the possibility that the data-generating process shifted are all excluded, and each of them widens the honest interval. **A distributional assumption that does not hold.** Symmetric, Gaussian-shaped intervals under-cover when errors are heavy-tailed or skewed, because the tail mass sits further out than a normal quantile allows. Demand and revenue series are routinely right-skewed. **Heteroscedasticity.** If error size scales with the level of the series or varies by day of week, one global width is too wide in quiet periods and far too narrow in busy ones. Pooled coverage can look mediocre while the busy periods, which are the ones that matter, are catastrophic. **A regime change in the holdout.** If the held-out window contains a shift the training period never saw, no interval built on history will cover it. That is a different finding from a mis-specified width, and the response is a monitoring and retraining question rather than a wider interval. ## Step four: fix and re-measure The most reliable repair is empirical rather than parametric: collect the forecast errors your backtest already produced, group them by horizon, and set the interval bounds at the appropriate empirical quantiles of those errors — the 10th and 90th for an 80% interval. This inherits the true error distribution, including skew and heavy tails, and it widens naturally with horizon because longer-horizon errors are larger. It needs enough errors per horizon to estimate a tail quantile, which is the practical constraint. Then re-measure coverage on data that was not used to set the width. Calibrating and evaluating on the same errors reproduces the original sin of using in-sample residuals. ## Do not stop at coverage Coverage alone is trivially gameable: intervals from minus infinity to plus infinity cover 100% of everything and are useless. Calibration must be reported with sharpness — the average interval width — and the honest summary is a score that trades the two off. An interval score does this by charging the width plus a penalty proportional to how far outside the interval a missed actual landed, scaled up as the nominal level gets tighter. Reporting coverage and mean width side by side is the minimum; two forecasters at identical coverage are ranked by the narrower one.
- Two forecasters both hit exactly 80% empirical coverage. How do you rank them?By sharpness: at equal calibration, the narrower intervals are more useful, because they constrain the decision more. Report mean interval width alongside coverage, or use an interval score that charges width and adds a penalty for actuals falling outside, scaled by the nominal level. Coverage alone cannot rank them — an arbitrarily wide interval achieves any coverage you like.
- Why report coverage per horizon rather than pooled across all forecasts?Because honest intervals widen with horizon, and pooling averages a well-calibrated one-step forecast together with a badly over-confident twelve-step one. The pooled number can look like a mild, uniform shortfall while the real defect is that width does not scale with horizon at all. The per-horizon table also tells you where to stop trusting the forecast.
- How would you widen the intervals without assuming a distribution?Use the empirical quantiles of the backtest errors, grouped by horizon: for an 80% interval, offset the point forecast by the 10th and 90th percentile of the errors seen at that horizon. This inherits skew and heavy tails automatically and widens with horizon on its own. It needs enough errors per horizon to estimate the tails, and the calibration errors must not be the same ones you then evaluate on.
It is a weather service that says 80% of its forecast ranges will contain the real temperature, and only 55% do. The forecast is not necessarily wrong on average — it is overconfident about how wrong it can be.
saying these in an interview costs you the question
- Treats under-coverage as a bias in the point forecast
- Pools all horizons into one coverage number
- Sets interval width from in-sample residual spread
- Declares a coverage gap real without checking sample size
- Reports coverage with no mention of interval width