skip to content

What can go wrong when you build a P10-P90 band from two separately fitted quantile models?

level: seniorimportance: nice to knowfreq 28%

answer

  1. two models, no shared constraint
  2. what stops the edges swapping order?
  3. nominal width versus realised width
  4. coverage overall can hide coverage by segment
  5. tail fits rest on the fewest rows

basics

~20 s

Two independent fits can cross, putting the P10 above the P90 for some inputs, and their 80 percent width is only nominal. Tail quantiles rest on few effective rows, so verify empirical coverage per segment.

solid answer

~50 s

Nothing in two separate fits forces the low quantile to stay below the high one. Each model minimises its own pinball loss with its own parameters, so in sparse or extrapolated regions of feature space the P10 curve can rise above the P90 curve -- quantile crossing -- and the band inverts. Second, the 80 percent is a target, not a guarantee: it holds only insofar as each conditional quantile is estimated well, and the tails are the part of the data you have least of, so both edges are the noisiest things the model produces. Third, coverage can be right overall and wrong everywhere in particular, too wide for busy segments and too narrow for quiet ones. The remedies are to enforce monotonicity -- fit the lower bound plus a non-negative width, or sort the two outputs -- and to measure realised coverage and average width on held-out data, sliced by segment.

go deeper

for a junior

Know that each edge of the band comes from its own model fitted at its own quantile level, and that the two are trained separately rather than produced by one model.

for a middle

Be ready to explain why nothing forces the lower fit to stay below the upper one, and to describe measuring realised coverage on held-out data instead of trusting the nominal width.

for a senior

Show the operating discipline: log the crossing rate, report coverage together with average width, slice both by segment, and treat instability across retrains as a signal that the taus are too extreme for the data volume.

for a principal

Own how uncertainty is communicated to the people acting on it -- what the band promises, how often it is allowed to be wrong, and whether the organisation should pay for the tighter band or accept a less demanding level.

## What the band is Fit one model with pinball loss at `tau = 0.1` and another at `tau = 0.9`, then report the pair as a band. For call-centre volume feeding a staffing plan this is attractive: the lower edge is the volume you will almost certainly exceed, the upper edge the volume you will almost certainly stay under, and the two together let a planner size a shift with the spread visible rather than hidden. Nominally 80 percent of outcomes should land inside. The construction is simple, which is exactly why its failure modes are easy to miss. ## Failure 1: quantile crossing The two models are trained independently. Each minimises its own loss, each has its own coefficients, and no term in either objective refers to the other. There is therefore no mechanism that guarantees ``` prediction(tau=0.1, x) <= prediction(tau=0.9, x) ``` for every input `x`. On the dense middle of the training distribution it holds almost automatically because the data enforces it. In sparse regions -- an unusual combination of day, hour and campaign flag; anything close to extrapolation -- the two fitted surfaces can and do swap order. The band inverts and the consumer sees a lower bound above the upper bound, or an implausibly narrow one just before it flips. Three standard remedies: 1. **Sort at prediction time.** Cheap, and it removes the visible absurdity, but it papers over the fact that both estimates were unreliable at that point. 2. **Reparameterise.** Fit the lower quantile directly and fit a strictly non-negative width on top of it, so the upper edge is the lower edge plus something that cannot go negative. Monotonicity holds by construction. 3. **Fit the quantiles jointly** under a single objective that sums the pinball losses and penalises or forbids crossing. Whichever you use, count how often crossing occurs before you patch it -- the crossing rate is a useful signal about where the model is out of its depth. ## Failure 2: nominal is not actual Eighty percent is a property of the *true* conditional quantiles. Your two fits are estimates, and a misestimated quantile does not deliver its advertised coverage. Two mechanisms push in opposite directions: model bias in the tails tends to shrink the band toward the centre and under-cover, while a very flexible model that overfits the training tails can produce bands that look tight in training and are wildly wrong out of sample. The only honest check is empirical. On held-out data, compute the share of actuals falling inside the band and compare with 0.8. Pair it with the average band width, because coverage alone is trivially gamed -- a band from minus infinity to plus infinity has perfect coverage and zero value. Good quantile models are judged on both: coverage close to nominal at the smallest width that achieves it. ## Failure 3: right on average, wrong everywhere Overall coverage of 80 percent is compatible with 95 percent coverage on weekday mornings and 60 percent on Monday peaks. The average hides the failure, and the failure is concentrated exactly where a staffing decision costs the most. Always slice coverage by the segments the decision cares about -- time of day, day of week, site, campaign -- and look at the worst segment, not the mean. ## Failure 4: the tails have the least data A `tau = 0.9` fit is effectively pinned by the top slice of the conditional distribution; a `tau = 0.99` fit even more so. Estimates at extreme taus have high variance, respond strongly to a few observations, and drift more between retrainings than a central fit does. If the band looks unstable across weekly retrains, extreme taus are the first suspect. Widening tau toward the centre -- P20 to P80 -- buys stability at the cost of a less demanding promise, and that trade should be made deliberately. ## Failure 5: two models to own Operationally this is a pair of models. Two training runs, two sets of hyperparameters, two monitoring dashboards, two things that can silently fail to retrain, two artefacts that can fall out of version sync so the band is built from a fresh lower edge and a stale upper one. Teams that ship quantile bands and treat them as "one model" discover this at the worst moment. Either package them as a single deployable unit or monitor them as separate models with an explicit consistency check between them. ## How to present it Finally, say what the band means when you hand it over. It describes the spread of *outcomes* the model expects for inputs like these, learned from the data. Its width is not a constant padding -- it varies with the features, which is precisely why it is more useful than a fixed margin, and precisely why a consumer who sees it narrow one week and wide the next needs to be told that is the model working, not the model breaking.

  • How would you stop the two quantile fits from crossing?
    Either reparameterise so the upper edge is the lower edge plus a strictly non-negative width, fit both quantiles under one joint objective that penalises crossing, or sort the two outputs at prediction time. Sorting is the cheapest and removes the visible absurdity, but it hides that both estimates were unreliable at that input, so log the crossing rate before patching it.
  • Coverage on held-out data comes out at 80 percent. Why is that not enough to accept the band?
    Coverage alone is trivially satisfied by a very wide band, so it must be read together with average width -- the goal is nominal coverage at the narrowest width achieving it. And an overall 80 percent can hide 95 percent in one segment and 60 percent in another, so slice by the segments the decision depends on and judge the worst one.
  • The band's width swings noticeably between weekly retrains. What is your first suspicion?
    The extreme taus. Each edge is pinned by a small slice of the conditional distribution, so its estimate has high variance and moves with a handful of observations. Confirm by checking whether a central fit on the same data is stable, and consider pulling the levels inward -- P20 to P80 -- if the promise allows it.

saying these in an interview costs you the question

  • Assumes the two fits cannot cross because one tau is higher
  • Treats the nominal 80 percent as a guaranteed coverage rate
  • Reports overall coverage without slicing by segment
  • Judges the band on coverage alone, ignoring its width
  • Forgets that the band is two independently trained, separately deployable models

context