Why do latency SLOs report p95 and p99 instead of the mean response time?
answer
- The average is not what users feel
- Thin tails barely move a mean
- One in a hundred requests
- Twenty calls, twenty chances to be slow
- Percentiles cannot be averaged together
basics
~20 sMean latency averages away the slow requests that users actually notice. The 95th and 99th percentiles report the experience of the worst 5 percent and 1 percent of requests, which is where dissatisfaction, timeouts and retries live.
solid answer
~40 sA p99 latency of 2,400 ms means one request in a hundred is slower than 2,400 ms; the mean is the arithmetic average over all requests. The two answer different questions. A service can average 120 ms and still make one user in a hundred wait more than two seconds, and the mean will barely move when that tail doubles. Tail latency is what users, timeouts and retry budgets respond to, so SLOs are written against percentiles. The effect compounds with fan-out: if a page makes 20 independent backend calls and each call exceeds its p95 five percent of the time, roughly 1 - 0.95^20, about 64 percent of page loads contain at least one slow call. Reporting p95 and p99 alongside the median gives you both the typical case and the pain.
go deeper
Be ready to state precisely what p95 and p99 mean in words: 5 percent and 1 percent of requests were slower than that value. Knowing that the mean hides a thin slow tail is the bar here.
Explain the mechanics: why a thin tail barely moves an average, why fan-out multiplies the chance of hitting the tail, and why the percentile is a single cut point rather than a summary of everything above it.
Show you have operated this. Talk about aggregating percentiles correctly across servers and windows, segmenting a stable global tail to find the cohort that is actually suffering, and choosing a percentile the traffic volume can support.
Own the SLO design. Decide whether the objective is written per request or per user session, what percentile the business can afford to defend, and what the error budget and alerting cost of chasing p99.9 across every service really is.
## What a latency percentile means Collect every request's response time over a window and sort them. The **p95** is the value below which 95 percent of those response times fall; equivalently, 5 percent of requests were slower than it. The **p99** is the same at 99 percent: one request in a hundred was slower. These are empirical quantiles of the observed latencies, computed by position in the sorted list, not by any arithmetic on the values. Saying 'our p99 is 2,400 ms' therefore means exactly one thing: about 1 percent of requests took longer than 2.4 seconds. It does not mean the slowest request took 2.4 seconds, and it does not mean 99 percent of requests took roughly 2.4 seconds. ## Why the mean is the wrong headline The mean is a balance point over all observations. Latency distributions are almost always right-stretched: a large mass of fast requests plus a thin, long tail of slow ones caused by cold caches, garbage collection pauses, lock contention, retries, or an unlucky shard. A thin tail contributes little to the average precisely because it is thin. Concretely: 99 requests at 100 ms and one at 5,000 ms gives a mean of about 149 ms. Double that outlier to 10,000 ms and the mean moves to 199 ms, a change most dashboards would not flag. But the p99 moved from 5,000 to 10,000 ms, and the user behind that request went from annoyed to gone. The mean is structurally insensitive to exactly the events that matter, and the percentile is structurally sensitive to them. There is a second problem. The mean is a number no user experiences. Every reported percentile corresponds to a real request that a real user actually waited for. ## Fan-out amplifies the tail Modern requests are rarely single calls. Suppose rendering a page requires 20 independent backend calls, and each call exceeds its own p95 with probability 0.05. The chance that all 20 come in under their p95 is 0.95^20, about 0.36. So roughly 64 percent of page loads include at least one call in its slow 5 percent. At p99 the same arithmetic gives 1 - 0.99^20, about 18 percent of page loads containing at least one p99-slow call. This is why teams that fan out widely care about p99 and even p99.9: a per-call tail that looks negligible becomes a common user-visible event once you multiply the opportunities to hit it. It is also why the *user-facing* percentile and the *per-call* percentile are different metrics, and the SLO should be explicit about which one it governs. ## Percentiles do not average or add A quantile is not a linear function of the data, so aggregating percentiles arithmetically is invalid. The average of ten servers' p99 values is not the fleet p99, and the average of twelve five-minute p99 values is not the hourly p99. Both can be badly wrong in either direction: if one server serves most of the traffic, the fleet tail is dominated by that server; if one server is small and pathological, averaging inflates the fleet number. The correct aggregation recomputes the quantile from the pooled observations. In practice systems store latency as a histogram or a sketch per server per interval, merge those structures, and read the percentile off the merged distribution. Merging counts is valid because the pooled distribution is genuinely the sum of the parts; averaging quantiles is not. ## Choosing which percentile to report Higher percentiles are estimated from fewer observations, so they are noisier and need more data to be meaningful. A p99 computed from 100 requests is essentially the largest or second-largest observation and will jump around wildly between windows; a p99.9 needs thousands of requests per window before it is stable. Picking p99.9 for a low-traffic endpoint produces an SLO that alerts on noise. A defensible reporting set is the median (typical experience), p95 or p99 (the pain), and the maximum or a count of requests over a hard threshold (the outright failures). Some teams instead report the *share of requests under a threshold*, for example '99 percent of requests complete in under 300 ms', which is the same information stated as a service level objective rather than as a latency number, and is often easier to reason about because the threshold is fixed and only the percentage moves. ## Common interview traps The two most frequent misreads are describing p99 as 'the slowest 1 percent averaged together' (it is a single cut point, not an average of the tail) and treating the p99 as an upper bound (a quarter of the requests above it may be far slower). If you need to characterise how bad the tail gets beyond the cut point, report a further percentile or the maximum, because the p99 by construction says nothing about the shape above itself.
- Ten servers each report a p99 latency for the same minute. How do you get the fleet's p99?Not by averaging them. A quantile is not a linear function of the data, so the average of per-server p99s can be far from the pooled value, especially with uneven traffic. Recompute it from the pooled request latencies, or have each server emit a latency histogram or sketch, merge those counts, and read the percentile off the merged distribution.
- Your p99 latency is stable but users complain. What else would you look at?Segment before you escalate the percentile. A stable global p99 can hide a specific cohort, region, endpoint, or client version whose own p99 is terrible but whose traffic share is too small to move the aggregate. Also check p99.9 and the maximum, and check whether complaints track sessions rather than requests, since a multi-call session hits the tail far more often than a single request does.
- Why is a p99 measured over 100 requests unreliable?With 100 observations the 99th percentile is essentially the largest or second-largest value, so it is determined by a single request and swings wildly between windows. High percentiles need enough observations that several data points sit above the cut. For low-traffic endpoints, widen the window, report a lower percentile, or report the count of requests over a fixed threshold instead.
Reporting mean latency is like judging an airline by the average delay; passengers remember the flight that sat on the tarmac for three hours.
saying these in an interview costs you the question
- Describes p99 as the average of the slowest one percent
- Treats the p99 as a hard upper bound on latency
- Averages per-server or per-window percentiles into one number
- Says 99 percent of requests took about the p99 value
- Picks p99.9 for an endpoint with a handful of requests per minute