skip to content

What does quantile regression at tau = 0.9 estimate that an OLS fit does not?

level: middleimportance: should knowfreq 26%

answer

  1. ask which functional is being fitted
  2. mean versus a conditional quantile
  3. the loss weights residual signs unequally
  4. nine to one at tau equals 0.9
  5. coefficient describes the 90th percentile

basics

~20 s

Quantile regression at tau = 0.9 estimates the conditional 90th percentile of the outcome; least squares estimates the conditional mean. Its coefficients say how a predictor moves the upper tail, which can differ from its effect on the average.

solid answer

~50 s

Ordinary least squares minimises squared residuals and so fits the conditional mean, `E[Y | X]`. Quantile regression at `tau` minimises an asymmetrically weighted sum of absolute residuals - the check loss, which charges `tau` per unit of under-prediction and `1 - tau` per unit of over-prediction - and so fits the conditional quantile `Q_tau(Y | X)`. At `tau = 0.9` an under-prediction costs nine times an over-prediction, which pulls the fitted line up until roughly 90% of the conditional mass sits below it. The coefficient on a predictor then reads as the change in the 90th percentile of the outcome per unit change in that predictor. That is a different quantity from the least-squares slope, not a robust version of it: a feature can leave the conditional mean untouched while steepening the upper tail. The topic forces you to say which functional of the conditional distribution you actually care about.

go deeper

for a junior

Know that least squares fits the average outcome and that quantile regression fits a chosen percentile of it instead, so the two answer different questions about the same data.

for a middle

Be ready to write the check loss, explain the nine-to-one asymmetry at tau = 0.9, and state the coefficient interpretation as a change in the 90th percentile of the outcome per unit of the predictor.

for a senior

Show judgment about when a tail model is the right estimand, fit a range of tau values rather than one, and flag that coefficient uncertainty at extreme tau inherits the same sparse-tail problem as any tail quantile.

for a principal

Own the choice of estimand across the organisation. Argue for when a team should commit to a tail-conditional target instead of an average, and be clear about the extra interpretation burden that choice puts on every downstream consumer.

## Two different targets A regression does not estimate "the relationship" between predictors and an outcome; it estimates a specific summary of the conditional distribution of the outcome given the predictors. Ordinary least squares targets the conditional mean `E[Y | X]`, because the value that minimises expected squared error is the mean. Quantile regression targets a conditional quantile `Q_tau(Y | X)` - the value below which a fraction `tau` of the conditional distribution lies. Both are legitimate. Which one you want is a modelling decision, and the interviewer is checking that you know it is a decision at all. ## The loss function that gets you there Quantile regression minimises the check loss, sometimes called the pinball loss. For a residual `u = y - y_hat`: `rho_tau(u) = tau * u` when `u >= 0` (the fit is too low) `rho_tau(u) = (1 - tau) * |u|` when `u < 0` (the fit is too high) Equivalently `rho_tau(u) = u * (tau - 1[u < 0])`. Three sanity checks on that formula. At `tau = 0.5` both arms have weight 0.5, so the loss is proportional to absolute error and the fit is the conditional median. At `tau = 0.9` an under-prediction costs 0.9 per unit and an over-prediction only 0.1, so the fit is pushed upward until being too low is as expensive at the margin as being too high - which happens when about 90% of the conditional mass sits below the fitted value. At `tau = 0.1` the asymmetry reverses and the fit is pushed down. Because the loss is piecewise linear rather than smooth, there is no closed-form normal-equations solution; the fit is obtained by a linear-programming-style optimisation. This is a mechanical difference from OLS, not a conceptual one. ## Reading a coefficient An OLS slope says: a one-unit increase in this predictor is associated with a `beta` change in the *average* outcome. A `tau = 0.9` quantile-regression slope says: a one-unit increase is associated with a `beta_0.9` change in the *90th percentile* of the outcome. These are different numbers about different things, and they can disagree dramatically. A payload-size feature might leave median response time flat - the fast path is unaffected - while sharply raising the 90th percentile, because only requests that already take a slow path are hurt. In wage data, a covariate can compress the bottom of the distribution while expanding the top, so the low-`tau` and high-`tau` slopes have different magnitudes and sometimes different signs. Fitting a whole family of `tau` values and reading the coefficients as a function of `tau` is how practitioners see that. ## What it buys, and what it costs Benefits. It answers tail questions directly instead of by proxy. It inherits the median's robustness to outliers in the outcome, since the loss is linear rather than quadratic in the residual, so one absurd value cannot drag the fit. It does not require constant residual variance to be interpretable - in fact heteroskedasticity is exactly the situation where different `tau` slopes carry different information, and the model is designed to express it. And the fitted quantile curves are equivariant under monotone transformations of the outcome, which the mean is not: the quantile of a log-transformed outcome exponentiates back to the quantile of the outcome, whereas the mean of the log does not. Costs. Each `tau` is a separate fit, so you spend more computation and more interpretation effort. Estimates in extreme `tau` regions are noisy for the same reason any tail quantile is noisy - little data lives there. And inference on the coefficients is harder than for OLS: the asymptotic covariance of the quantile-regression estimator involves the conditional density of the outcome at the fitted quantile, the same awkward nuisance quantity that makes a plain sample quantile resist a closed-form interval. Practitioners therefore usually get standard errors for quantile-regression coefficients by resampling. ## What a weak answer sounds like "Quantile regression is just robust regression" is the common miss. Robustness is a side effect; the point is that it estimates a different functional. "It removes outliers" is worse - it removes nothing, it merely weights residuals linearly and asymmetrically. ## The one-line version OLS answers "how does this predictor move the average outcome"; quantile regression at `tau = 0.9` answers "how does it move the 90th percentile", and in a skewed metric those are genuinely different questions.

  • What loss does quantile regression minimise, and what does it reduce to at tau = 0.5?
    It minimises the check loss `rho_tau(u) = u * (tau - 1[u < 0])`, which charges `tau` per unit of under-prediction and `1 - tau` per unit of over-prediction. At `tau = 0.5` both weights are 0.5, so the objective is proportional to the sum of absolute residuals and the fit becomes the conditional median rather than the conditional mean.
  • Why are standard errors for quantile-regression coefficients usually obtained by resampling?
    Because the asymptotic covariance involves the conditional density of the outcome evaluated at the fitted quantile. That density is an unknown local quantity, hard to estimate well and sitting in a denominator, so plug-in standard errors are fragile. Resampling avoids estimating it explicitly, at the cost of computation.
  • Can a predictor have a positive OLS slope and a near-zero slope at tau = 0.9?
    Yes. The two coefficients describe different functionals of the conditional distribution. A predictor can lift the bulk of the distribution - moving the mean - while leaving the upper tail unchanged, so the mean slope is positive and the `tau = 0.9` slope is flat. Disagreement between them is informative, not an error.

saying these in an interview costs you the question

  • Calls it simply a robust version of least squares
  • Says it removes or downweights outliers in the predictors
  • Interprets a tau = 0.9 coefficient as an effect on the average
  • Expects one fit to give all quantiles at once
  • Assumes plug-in standard errors work as easily as for least squares

context