A performance run's response times form two distinct clusters. What does one 95th-percentile figure get wrong?
answer
- Two populations, not one slow end
- It lands in whichever cluster holds the rank
- It tracks the mix, not the speed
- Near the crossover it swings wildly
- Publish a share above a boundary
basics
~20 sIt implies one population with one centre. With two clusters the figure lands in whichever cluster holds the rank, describes only that one, and moves with the proportion between them rather than with either cluster's own speed.
solid answer
~50 sTwo clusters mean two populations sharing one report — a fast path and a slow one, separated by times almost nothing produced. A single high percentile silently claims one population with one centre, and what it actually reports depends only on the **share** in the slow cluster. If that share is 3%, the 95th-percentile figure is read out of the fast cluster and the slow population is invisible. If it is 12%, the figure sits inside the slow cluster and says nothing about the 88% who were fast. If it is near 5%, the figure straddles the gap and swings wildly on a small change in proportion. In every case the number tracks the **mix**, not either cluster's own speed: it can improve while the slow path gets worse. Report the share in each cluster, a centre for each, and the splitting variable if you can name it.
code
pseudocode · 14 lines# one figure hides which cluster it was read from
slow = [t for t in times if t > CROSSOVER]
fast = [t for t in times if t <= CROSSOVER]
slow_share = count(slow) / count(times)
report = {
fast_cluster: { share: 1 - slow_share, middle: median(fast) },
slow_cluster: { share: slow_share, middle: median(slow) },
share_above_boundary: slow_share, # stable where a rank is not
curve: counts_per_range(times) # the flat stretch is the finding
}
# value_at_rank(sort(times), 0.95) lands in fast, in slow,
# or in the empty gap - decided only by slow_sharego deeper
Recall that response times do not always form a single group. When they fall into a fast set and a slow set, one summary figure describes at most one of them.
Explain that the reported value is decided by the share of requests in the slow cluster, and be able to say which cluster the figure comes from for a given share.
Show that you look at the shape before quoting a figure, recognise a swinging figure as a crossover rather than noise, and split the samples by whatever variable separates the clusters before reporting again.
Own how results are stated across teams: shares above agreed boundaries rather than single crossover-sensitive figures, and a rule that a distribution is published alongside any headline number.
## Two populations wearing one number A bimodal response-time distribution is two populations that happen to share a report: a fast group and a slow group with a sparsely-occupied stretch between them. It is one of the most common shapes in real systems, and it always has a cause worth naming — a served-from-memory read against one that goes to storage, a request on an already-open connection against one that establishes a new one, a small payload against a large one, one instance behaving differently from its peers, a first request after a periodic runtime pause. A percentile is defined perfectly well on such a distribution. The problem is what a single percentile **implies**: that the samples form one body with one centre and a slow end, so a single figure summarises them. On a two-cluster distribution that implication is false, and everything downstream of it goes wrong. ## Where the figure actually lands The reported value depends on one thing — the share of requests in the slow cluster relative to the portion above the cut. Three cases: 1. **Slow share below 5%** (say 3%): the whole slow cluster sits above the 95% rank, so the figure is read out of the fast cluster. Every slow request is invisible in the headline number, and making them slower will not change it. 2. **Slow share well above 5%** (say 12%): the rank falls inside the slow cluster, so the figure describes the slow population. It now says nothing about the 88% of requests that were fast, and it also understates how bad the slow group gets. 3. **Slow share near 5%**: the rank straddles the gap. The reported figure jumps between the two clusters — potentially by an order of magnitude — on a change in mix of a fraction of a percent, with nothing in the system having changed at all. Case three is the one that wastes weeks, because a figure that swings between runs looks like an unstable system rather than an unstable statistic. ## What the number implies that is not true - **That some requests took about that long.** In case three the figure lands in the gap, and it names a duration that essentially nothing in the run experienced. - **That the figure moves when performance moves.** It tracks the proportion between clusters. A run can improve its headline figure purely by sending more work down the fast path while the slow path is unchanged or worse. - **That neighbouring percentiles describe a smooth curve.** Reading between two percentiles across the gap interpolates a region with almost no samples in it, producing values the system never produced. - **That one figure can be compared across runs.** Two runs with the same two clusters and different proportions produce very different figures. ## Reporting the shape instead of the point | What to publish | What it survives | Why it helps | |---|---|---| | The share of requests above a fixed time boundary | Mix changes near a crossover | A share moves smoothly where a rank jumps | | A centre and a spread per cluster | Either cluster changing on its own | Separates speed from proportion | | The cumulative curve or the counts per latency range | Everything | The flat stretch across the gap *is* the finding | | A figure per group once the splitting variable is known | Mix changes entirely | Turns one dishonest report into two honest ones | The strongest move is the last one. If you can identify what separates the clusters — served from memory or not, connection reused or not, payload above or below some size — then split the samples by it and report each group as its own unimodal population with its own request share. The bimodal report disappears, replaced by two reports that each mean what they say, plus a proportion that is itself a useful number to watch. Where the splitting variable is not yet known, finding it is the work: take the samples above the gap, look at what they have in common that the samples below do not, and check the obvious candidates first. ## What to do - Look at the shape before quoting any figure from it; a single high percentile on an unexamined distribution is a guess about its shape. - Publish the share in each cluster alongside any percentile you do quote. - Where a rule has to be stated as a number, prefer *"no more than 2% of requests above 500 ms"* to a percentile point, because a share stays stable exactly where a rank does not. - Treat a figure that swings between runs with no change to the system as evidence of a crossover, not of noise, until you have looked at the curve.
- The headline 95th-percentile figure improved sharply and neither cluster's own centre moved. What do you report?That the proportion of requests taking the fast path rose, and that no request got faster. The figure is a function of the mix as much as of the clusters, so an improvement of that shape is a change in how work is routed or how often a fast condition is met — worth understanding, but not evidence that the slow path improved, and reversible the moment the proportion shifts back.
- How would you find the variable that separates the two clusters?Split the samples at the sparse stretch between them and compare the two groups on everything recorded against each request — which operation, which instance served it, payload size, whether a cached or established resource was reused, position in the run. Whatever differs systematically is the candidate. Then re-report by that variable and confirm each group is now a single population.
- Why prefer a share above a fixed boundary to a percentile point on this shape?Because a share is continuous in the mix while a rank is not. Near a crossover a small proportion change moves a percentile figure between clusters, while the share above the boundary moves by the same small amount. The share is also directly meaningful — it states how many requests were slow — where a jumping figure states nothing reliable at all.
Quoting one high percentile for a two-cluster distribution is like reporting a single travel time for a route where some journeys catch the ferry and some wait for the next one. The average traveller experience is not a number between the two.
saying these in an interview costs you the question
- Treats a two-cluster shape as one centre with a slow end
- Reads a percentile improvement as the slow path getting faster
- Interpolates between percentiles across the empty gap
- Publishes one figure without ever looking at the curve
- Blames run-to-run swing on noise near a crossover