skip to content

Why does a sample p95 resist the closed-form confidence interval a sample mean gets?

level: middleimportance: must knowfreq 55%

answer

  1. the mean's recipe reuses what the sample measures
  2. the quantile's standard error hides a nuisance term
  3. that term is local, not global
  4. density at the quantile, in the denominator
  5. few points live near a p95

basics

~20 s

A sample quantile's standard error contains the unknown probability density at that quantile: roughly sqrt(p(1-p)/n) divided by that density. The density has to be estimated from the few points sitting near a p95, so practitioners resample instead.

solid answer

~50 s

For a mean, the standard error is `s / sqrt(n)` because the spread parameter you need is exactly what the sample standard deviation estimates. For a sample quantile it is different: the asymptotic standard error of the sample p-quantile is `sqrt(p(1-p)/n) / f(q_p)`, where `f(q_p)` is the probability density of the data at the population quantile. Nothing in the sample hands you that density directly, and near a p95 only a thin slice of observations sits in the neighbourhood that carries the information, so any density estimate there is noisy. Because the density sits in the denominator, that noise is amplified. The sampling distribution of an extreme sample quantile is also skewed, so a symmetric estimate-plus-or-minus-critical-value interval misstates it. That is why quantile intervals are normally produced by resampling or by a distribution-free order-statistic construction rather than a plug-in formula.

go deeper

for a junior

Be ready to say that a percentile computed from a sample is an estimate with error, and that the mean's s over sqrt(n) formula does not transfer to it. Knowing that you would resample to get the interval is enough here.

for a middle

You are expected to write the quantile standard error as sqrt(p(1-p)/n) divided by the density at the quantile, and explain that the density is the unknown, locally estimated piece that makes the formula impractical.

for a senior

Show that you check the practical preconditions: how many observations sit near the quantile, whether the metric is tied or rounded, and whether the sampling distribution is skewed enough that a symmetric interval would mislead.

for a principal

Own the reporting standard. Decide which quantiles the organisation is allowed to quote at a given traffic volume, and require an interval alongside any tail number so teams stop arguing over noise dressed up as a regression.

## The question behind the question An interviewer asking this wants to see whether you understand that "put an interval on it" is not one universal recipe. A confidence interval is an interval, computed from data, built so that over repeated samples it captures the unknown population value a stated fraction of the time. Getting one requires knowing how much the estimator moves from sample to sample. For some estimators that quantity falls out of the sample; for a quantile it does not. ## Why the mean is easy The sample mean of n independent draws has variance `sigma^2 / n`, where `sigma` is the population standard deviation. The sample standard deviation `s` estimates `sigma` directly and uses every observation to do it. Plugging `s` in gives `SE = s / sqrt(n)`, a number you can compute from the same data with no extra machinery. The unknown nuisance quantity and the thing the sample measures well are the same quantity. ## Why a quantile is not The population p-quantile `q_p` is the value with a fraction `p` of the distribution below it. The sample p-quantile `q_hat_p` is the corresponding order statistic of the sorted sample. Under the usual assumptions - independent draws from a distribution with a density `f` that is continuous and strictly positive at `q_p` - the estimator is asymptotically normal: `sqrt(n) * (q_hat_p - q_p)` converges to Normal(0, p(1-p) / f(q_p)^2) so the standard error is approximately `SE(q_hat_p) = sqrt(p(1-p)/n) / f(q_p)` Read that formula slowly. The numerator is benign: `p(1-p)` is 0.25 at the median and 0.0475 at p95, and the `1/sqrt(n)` shrinkage is the familiar one. The trouble is entirely in `f(q_p)`, the height of the density curve at the quantile. It is an unknown population feature, it is a *local* feature (only data near `q_p` speaks to it), and it appears in the denominator, so a 20% error in the density estimate is a 25% error in the standard error. ## Estimating that density is the hard part Estimating a density at a point requires choosing how wide a neighbourhood to average over. Too narrow and the estimate is pure noise; too wide and you are measuring the wrong place on the curve. That bias-variance choice is bad enough in the middle of a distribution, and it is much worse at p95, where by construction only 5% of the sample lies above the target and the effective number of observations informing the local shape is small. With n = 500 that is roughly 25 points above the quantile. The information about `f` at the p95 comes from a handful of them. ## Two more reasons the plug-in recipe misfires First, convergence to normality is slower for tail quantiles than for a mean, so at realistic sample sizes the sampling distribution of the sample p95 is still visibly skewed - longer to the right, since the sparse upper tail can push the estimate far up but the denser body cannot pull it far down. A symmetric interval centred on the estimate then has the wrong shape even if its width is roughly right. Second, the asymptotic result needs `f(q_p) > 0`. If the distribution has a flat gap where the quantile falls, or if the metric is heavily tied - millisecond-rounded latencies, integer counts, a capped score - the continuity assumption behind the formula breaks and the sample quantile can sit on a single repeated value across many resamples. Check for ties before trusting any quantile interval. ## What is actually done in practice Three routes, all of which sidestep estimating `f` explicitly. Resampling-based intervals let the data supply the sampling variability empirically. Distribution-free order-statistic intervals use the fact that the count of observations below the population quantile is binomial, which needs no density at all. And where you must extrapolate past the observed range, a parametric tail model fitted above a threshold replaces the empirical quantile entirely. Whichever you choose, the discipline is the same: report the interval, and do not quote a p99 to three significant figures from a few hundred observations. ## The one-line version The mean's uncertainty depends on a global spread the sample measures well; a quantile's uncertainty depends on the local density where the quantile falls, which the sample measures badly - especially in a tail.

  • Does the standard error of a sample p95 still shrink at the 1/sqrt(n) rate?
    Yes, the rate is the same - the `sqrt(n)` sits in the same place. What differs is the constant in front: `p(1-p)/f(q_p)^2` is typically far larger in a sparse tail than the variance term governing a mean. So a p95 needs a much bigger sample than a mean to reach comparable relative precision, but it does eventually get there.
  • Why is a well-built interval on a p99 usually asymmetric around the point estimate?
    Because the density is lower above the quantile than below it in a right-skewed metric. The estimate can be pulled a long way up by a sparse upper tail and only a short way down by the dense body, so the sampling distribution is right-skewed. An interval that respects that is wider on the upper side; forcing symmetry understates the upside risk.
  • What breaks if the metric is rounded to whole milliseconds?
    Ties break the continuity assumption the asymptotic formula rests on. With enough repeated values the sample quantile can land on the same value across many samples, so its distribution is lumpy rather than continuous, and interval endpoints jump in discrete steps. Check the number of distinct values near the quantile before reporting any interval on it.

Measuring a mean is like weighing a whole crate; measuring a p95 is like judging the height of a fence from the handful of people tall enough to see over it.

saying these in an interview costs you the question

  • Uses s divided by sqrt(n) as the standard error of a percentile
  • Assumes the sample p95 is normally distributed at any sample size
  • Thinks quantile uncertainty depends only on n, never on distribution shape
  • Reports a p99 as an exact number with no interval at all
  • Ignores heavy ties or rounding when putting an interval on a quantile

context