Why do two reporting pipelines compute different 90th percentiles from the same 10 numbers?
answer
- The data have no value at rank 9.1
- Take an observation or interpolate between two
- Several conventions, all defensible
- Worst on small samples and far tails
- Pin the definition in the metric spec
basics
~20 sBecause there is no single definition of a sample quantile. With 10 sorted values the requested rank falls between two observations, and the competing conventions either take a specific observation or interpolate between neighbours, which yields different numbers from identical data.
solid answer
~50 sA percentile is defined exactly for a distribution, but a sample of 10 values gives only 10 order statistics, so the 90th percentile has to be pinned down by a convention. Take the sorted values 10, 20, ..., 100. The nearest-rank or inverse-empirical-CDF definition takes the smallest value whose cumulative share reaches 0.9, giving 90. A convention that averages at the jump returns (90 + 100) / 2 = 95. A common linear-interpolation convention uses rank h = (n - 1) * p + 1 = 9.1 and returns 90 + 0.1 * (100 - 90) = 91. Another uses h = (n + 1) * p = 9.9 and returns 99. Same 10 numbers, four defensible answers. The statistics literature catalogues nine such definitions. The fix is organisational: pin one definition in the metric's specification and make sure every producer of the number uses it.
go deeper
Be ready to say that a percentile of a small sample can fall between two observations, so the reported value depends on whether the method picks a neighbour or interpolates between them.
Work a concrete example: state the rank each convention computes for p = 0.9 with n = 10 and the value it returns. Explain why the disagreement shrinks as the sample grows.
Show the triage instinct. Treat a metric shift with no data change as a definitional difference first, and know the other suspects: aggregation across shards and windows, window length, and approximate quantile summaries.
Own the metric contract across teams. Decide whether percentile targets or threshold-share targets are the organisation's standard, since the latter removes the ambiguity entirely at the cost of a fixed threshold.
## Why a sample quantile needs a convention For a probability distribution the p-th quantile is well defined. For a *sample*, you have n observations and nothing in between them. Asking for the 90th percentile of 10 values asks for a point at 'rank 9-point-something', and the data simply do not contain such a point. Every practical definition therefore has to make a choice, and reasonable people made different ones. ## Four answers to the same question Sort the values: 10, 20, 30, 40, 50, 60, 70, 80, 90, 100. So n = 10 and p = 0.9. **Inverse empirical CDF (nearest rank).** The empirical CDF steps up by 1/n at each observation. Take the smallest value whose cumulative share is at least 0.9. Nine of the ten values are at or below 90, so the cumulative share at 90 is exactly 0.9 and the answer is **90**. This definition always returns an actual observation, never an invented value. **Averaging at the discontinuity.** When n * p is exactly an integer, the empirical CDF has a flat stretch between two observations and one convention splits the difference: (x9 + x10) / 2 = (90 + 100) / 2 = **95**. **Linear interpolation with rank h = (n - 1) * p + 1.** Here h = 9 * 0.9 + 1 = 9.1, so take the 9th value plus 0.1 of the way to the 10th: 90 + 0.1 * (100 - 90) = **91**. This convention maps p = 0 to the minimum and p = 1 to the maximum, which is convenient and is a widespread default. **Linear interpolation with rank h = (n + 1) * p.** Here h = 11 * 0.9 = 9.9, giving 90 + 0.9 * (100 - 90) = **99**. This convention treats the n observations as splitting the distribution into n + 1 equal-probability gaps and is favoured when the goal is an approximately unbiased estimate of the underlying distribution's quantile. So the same ten numbers legitimately produce 90, 91, 95 or 99. None of these is a bug. Hyndman and Fan's well-known survey of sample-quantile definitions enumerates nine distinct types in use. ## When the difference matters and when it does not The gap between definitions is at most the distance between two neighbouring order statistics. Two things follow. First, **the disagreement shrinks as n grows**, because with many observations consecutive order statistics are close together. Across a million request latencies the four definitions of p90 agree to within noise, and arguing about them is a waste of time. Second, **the disagreement is worst exactly where it hurts**: small samples and extreme quantiles. In the example above the definitions differ by 9 out of 90, roughly 10 percent, because the spacing in the tail is wide relative to the values. This is why a p99 computed from a few hundred observations can differ visibly between two systems that both believe they are correct. ## The deeper problem: resolution With n = 10 the empirical CDF only takes the values 0.1, 0.2, ..., 1.0. There is no information at all about the 95th or 99th percentile beyond 'somewhere at or above the largest observation'. Any p99 reported from 10 samples is an artifact of the interpolation rule rather than a measurement. As a rule of thumb, a quantile at level p needs enough observations that several data points sit above the cut, so a stable p99 wants at least several hundred observations per window and a p99.9 wants thousands. ## What to do about it in production The engineering answer is not to pick the 'true' definition, because there isn't one. It is to remove the ambiguity: - **Write the definition into the metric specification.** If two teams publish 'p95 checkout latency', the spec should say which convention computes it, just as it says which requests are in scope. - **Expect drift when systems change.** A metric that moved when a pipeline was rewritten, with no change in the underlying data, is a classic false regression, and the quantile convention is one of the first things to check. - **Prefer count-based statements when precision matters.** 'What fraction of requests completed under 300 ms' has no interpolation ambiguity at all: it is a count divided by a count. Many SLOs are cleaner stated this way. - **Be careful with approximate quantile structures.** Systems that compute quantiles over large streams typically store a compressed summary rather than every observation, and return a value with a stated error bound on the rank. That error is usually much larger than the difference between the textbook definitions, so it dominates the discussion at scale. ## The interview answer A strong answer states that sample quantiles require a convention, demonstrates two or three conventions numerically on a small sorted list, notes that the discrepancy shrinks with n and blows up in small-sample tails, and ends on the organisational fix: pin the definition, and prefer threshold-based formulations when two systems must agree exactly.
- Which of these conventions always returns a value that actually appears in the data?The nearest-rank or inverse-empirical-CDF definition, which returns the smallest observation whose cumulative share reaches p. That property matters when the quantity must be a real observed value, for example when reporting a representative record or when the values are categorical-ordinal and an interpolated in-between value would be meaningless.
- A latency metric jumped 8 percent the week a reporting pipeline was rewritten. How do you triage it?Treat it as a suspected definitional change before treating it as a regression. Recompute the old and new metric from the same raw window and compare; check the quantile convention, the aggregation across servers and time windows, the window length, and whether an approximate quantile structure replaced exact computation. Only if all of those match should you look for a real performance change.
- Does the choice of convention matter for the median?Much less, because the observations are densest near the middle of the distribution, so neighbouring order statistics are close together and every convention lands in a narrow range. The one visible difference is with an even sample size, where some conventions average the two middle values and others return one of them. In the tails the same conventions can differ substantially.
saying these in an interview costs you the question
- Assumes one system must have a bug in it
- Believes there is a single correct sample-quantile formula
- Reports p99 from a sample of a few dozen values
- Thinks interpolation error matters most at the median
- Ignores approximate quantile summaries when comparing systems