skip to content

Robust and Quantile Losses

Swapping squared error for absolute, Huber or pinball loss changes what the fit chases: the median, a compromise, or a chosen quantile. Interviewers ask what to do when outliers dominate the error.

on this pageshow

questions

5

How does minimising pinball loss at tau = 0.9 produce a 90th-percentile prediction?

level: middleimportance: must knowfreq 52%

answer

  1. two arms, two different slopes
  2. one direction of error is charged more
  3. balance the probability mass either side
  4. tau equals the share falling below the prediction
  5. at one half it collapses to absolute error

basics

~20 s

Pinball loss at tau = 0.9 charges 0.9 per unit of under-prediction and 0.1 per unit of over-prediction. That nine-to-one asymmetry is minimised by the value the target falls below 90 percent of the time.

solid answer

~50 s

Pinball loss, also called the check or quantile loss, weights the two directions of error differently. For actual `y` and prediction `q`, it is `tau * (y - q)` when the prediction is too low and `(1 - tau) * (q - y)` when it is too high. At `tau = 0.9`, being ten units short costs 9 while being ten units over costs 1, so the optimiser keeps raising the prediction until the extra cost of over-predicting on the many rows below balances the saved cost of under-predicting on the few rows above. That balance is reached exactly where the fraction of outcomes falling below the prediction equals tau -- the 90th percentile. Setting `tau = 0.5` makes the weights equal, recovering half the absolute error and the median. Fitting one model per tau gives you specific quantiles of the conditional distribution, not the mean.

code

python · 15 lines
python
def pinball(actual, quote, tau):
    """Under-quoting is charged tau per unit; over-quoting, 1 - tau per unit."""
    if actual >= quote:
        return tau * (actual - quote)
    return (1 - tau) * (quote - actual)

tau = 0.9
minutes = [18, 20, 21, 22, 24, 25, 27, 28, 29, 33, 38, 52]

# arriving 10 minutes late hurts nine times as much as arriving 10 early
print(round(pinball(40, 30, tau), 2), round(pinball(20, 30, tau), 2))  # 9.0 1.0

best = min(minutes, key=lambda q: sum(pinball(a, q, tau) for a in minutes))
late = sum(1 for a in minutes if a > best)
print(best, late, len(minutes))                     # 38 1 12

go deeper

for a junior

Know that pinball loss punishes under-prediction and over-prediction by different amounts, and that the level tau you choose is the percentile the fitted prediction targets.

for a middle

Be ready to write both branches with their tau and one-minus-tau weights, and to argue why the optimum lands where the share of outcomes below the prediction equals tau. Note that tau of one half recovers absolute error.

for a senior

Demonstrate the operating side: verify empirical coverage on held-out data by segment, expect extreme taus to be noisy, and flag that every downstream consumer of the prediction is now reading a quantile rather than an expectation.

for a principal

Own the framing that the choice of tau is a business promise, not a modelling detail. Decide who sets it, how often it is revisited, and what service commitment the chosen level is actually underwriting.

## The loss Pinball loss (also the check loss, or quantile loss) is defined for a chosen level `tau` between 0 and 1. With actual value `y` and prediction `q`: ``` L(y, q) = tau * (y - q) if y >= q (under-predicted) L(y, q) = (1 - tau) * (q - y) if y < q (over-predicted) ``` Both branches are non-negative and both are zero when `q = y`. The loss is piecewise linear with a kink at zero -- the shape gives it its name, the two straight arms hinged like a pinball flipper. The only thing that distinguishes it from absolute error is that the two arms have different slopes: `tau` on the under-prediction side, `1 - tau` on the over-prediction side. At `tau = 0.5` the slopes are equal at one half each, so pinball loss becomes half the absolute error and the fitted value is the median. Every other tau tilts the flipper. ## Why the optimum is the tau-quantile Take a single number `q` predicted for a whole distribution. Expected loss is ``` E[L] = tau * E[(y - q) when y > q] + (1 - tau) * E[(q - y) when y < q] ``` Nudge `q` upward by a tiny amount. Every outcome below `q` gets worse, at rate `(1 - tau)` each, and there is a `P(y < q)` share of them. Every outcome above `q` gets better, at rate `tau` each, and there is a `1 - P(y < q)` share of them. The derivative is therefore ``` (1 - tau) * P(y < q) - tau * (1 - P(y < q)) ``` Setting it to zero and solving gives `P(y < q) = tau`. The minimiser is precisely the value the target falls below with probability tau -- the tau-quantile. With features in the model, the same argument applies pointwise and the fit estimates the **conditional** tau-quantile. Notice what the argument used: only the *probability* on each side of the prediction, never how far away the outcomes on the far side were. That is why quantile fits inherit absolute error's insensitivity to how extreme the tail values are. ## Reading the asymmetry as a promise A food-delivery ETA is a good illustration. The customer sees one number. Quoting the conditional mean means roughly half of deliveries arrive later than promised, which reads as unreliable however good the model's average error is. Quote instead the conditional 90th percentile, fitted with pinball loss at `tau = 0.9`: the quote is deliberately pessimistic, and by construction only about one delivery in ten runs past it. The model has not become more accurate; it has been asked a different question -- "what time will this beat nine times out of ten?" rather than "what is the expected time?". That also tells you how to sanity-check such a model in production: measure the share of actual outcomes that fall below the prediction on held-out data. For `tau = 0.9` it should sit near 90 percent. If it is 97 percent your quotes are needlessly padded; if it is 78 percent you are breaking a promise you advertised. ## Practical points - **One tau, one model.** A single quantile fit gives one slice of the conditional distribution. Wanting several quantiles means fitting several models, each with its own tau. - **Extreme taus are data-hungry.** At `tau = 0.99` the fit is pinned by the top one percent of rows, so the estimate is noisy unless the dataset is large. - **It is not differentiable at zero.** Like absolute error, the loss has a kink where the residual crosses zero, so it is fitted with methods that tolerate a subgradient, or with a smoothed approximation. - **The loss is also the evaluation metric.** Average pinball loss at the same tau on held-out data is the natural way to compare two candidate quantile models -- and the natural companion check is the empirical coverage described above, because a model can score reasonably on average pinball loss while being systematically mis-calibrated in one segment. - **The prediction moves, and everything downstream notices.** Swapping from a mean fit to a `tau = 0.9` fit lifts essentially every prediction. Anything that consumes those numbers -- capacity plans, promised times, alerting thresholds -- is now reading a different quantity, and should be told so. ## The common misreading Candidates often describe a `tau = 0.9` fit as "a model with a safety margin added". It is not an additive offset. The gap between the median fit and the 90th-percentile fit is whatever the local spread of the conditional distribution demands: wide during a chaotic Friday dinner rush, narrow on a quiet Tuesday morning. A quantile fit learns that varying width from the features; a constant margin bolted onto a mean fit cannot.

  • How would you check in production that a tau = 0.9 model is doing its job?
    Measure empirical coverage on held-out data: the share of actual outcomes falling below the prediction should sit near 90 percent. Check it by segment as well as overall, since a model can hit 90 percent on average while over-padding quiet periods and under-padding busy ones. Average pinball loss at the same tau compares candidate models.
  • What is pinball loss at tau = 0.5 equivalent to?
    Half the absolute error. The two arms carry equal slope, so the loss becomes symmetric and its minimiser is the median rather than a tail quantile. This makes absolute error a special case of pinball loss, not a different family.
  • Why is a tau = 0.9 fit not just a fixed safety margin added to a mean prediction?
    The gap between the central fit and the 90th percentile is set by the local spread of the conditional distribution, which the quantile fit learns from the features. Where outcomes are volatile the gap widens; where they are tight it narrows. A constant offset applies the same padding everywhere and is wrong in both directions.
  • Does one quantile fit tell you the whole conditional distribution?
    No. It estimates a single slice at the tau you chose. Describing the distribution means fitting several taus, each a separate model with its own parameters, and then reasoning about them jointly.

It is a tilted see-saw. Under-shooting sits nine times heavier than over-shooting, so the balance point slides up the distribution until only a tenth of the outcomes remain on the heavy side.

saying these in an interview costs you the question

  • Describes a quantile fit as a mean fit plus a constant safety margin
  • Says pinball loss at tau equals 0.9 predicts the top 10 percent of rows
  • Cannot state which direction of error tau penalises
  • Assumes one quantile fit yields the whole distribution
  • Thinks a higher tau makes the model more accurate rather than more pessimistic

context

open as a page

Why does one extreme target value distort a squared-error regression fit more than an absolute-error fit?

level: middleimportance: must knowfreq 62%

basics

~20 s

Squared error penalises a residual by its square, so its gradient grows with the residual: a point ten times further off pulls ten times harder. Absolute error's gradient has fixed size, so extreme points get no extra vote.

open as a page

What does the delta parameter control in Huber loss for a regression fit?

level: middleimportance: should knowfreq 46%

basics

~20 s

Delta is the residual size at which Huber switches from squared to absolute behaviour. Below delta a row's pull grows with its error; above delta the pull is capped at delta, so extreme rows stop dominating the fit.

open as a page

Your demand model's stockouts cost four times what excess inventory does — how do you reflect that in the loss?

level: principalimportance: should knowfreq 34%

basics

~20 s

A four-to-one cost ratio is pinball loss with tau = 4/(4+1) = 0.8, so forecast the 80th percentile of demand rather than the mean. Confirm the ratio is real, linear and stable before baking it into training.

open as a page

What can go wrong when you build a P10-P90 band from two separately fitted quantile models?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Two independent fits can cross, putting the P10 above the P90 for some inputs, and their 80 percent width is only nominal. Tail quantiles rest on few effective rows, so verify empirical coverage per segment.

open as a page