Why can a performance run's per-interval 95th-percentile times not be averaged into one figure for the whole run?
answer
- A rank, not a quantity
- Averages assume things that add
- Quiet intervals get equal weight
- The pooled figure sits between the extremes
- Merge counts or raw samples
basics
~20 sPercentiles are ranks over a set of samples, not quantities that add or average. The mean of per-interval figures weights a quiet minute like a busy one and throws away each interval's shape. Merge raw samples or bucket counts instead.
solid answer
~50 sA percentile is a **rank over a set of samples**, not a quantity that combines arithmetically. Averaging one 95th-percentile figure per interval gives every interval equal weight however many requests it served, and each stored figure has already discarded the shape the pooled rank depends on. Weighting by request count does not repair it: a weighted mean of quantiles is still a mean. The stored figures do *bound* the answer — the run's true 95th percentile lies between the smallest and the largest of them — so publish that interval rather than an invented point inside it. A single honest whole-run figure exists only if the result store kept either every request's time, or counts of requests per latency bucket that add interval by interval and can be re-read. That is a retention decision taken before the run, not afterwards.
code
pseudocode · 18 lines# WRONG: averaging the stored per-interval quantiles
run_p95 = mean(iv.stored_p95 for iv in intervals)
# RIGHT, if raw per-request times were retained
all_times = concat(iv.samples for iv in intervals)
run_p95 = value_at_rank(sort(all_times), 0.95)
# RIGHT, if per-interval bucket counts were retained
merged = [0] * len(BUCKET_EDGES)
for iv in intervals:
for b in range(len(BUCKET_EDGES)):
merged[b] = merged[b] + iv.counts[b]
run_p95 = percentile_from_counts(merged, BUCKET_EDGES, 0.95)
# What the stored quantiles alone still license
lower_bound = min(iv.stored_p95 for iv in intervals)
upper_bound = max(iv.stored_p95 for iv in intervals)
# report the interval, not its midpointgo deeper
Recall that a percentile is a position in a sorted list of measured times, not an amount. Combining two sets of measurements means going back to the measurements and sorting them together.
Be ready to explain why an equal-weight average of per-interval figures is wrong on two counts — sparse intervals count as much as busy ones, and each interval's shape was already discarded — and why weighting by request count does not rescue it.
Show that you settle retention before a run starts: raw times or per-interval bucket counts when a merged figure will be needed, and a published bound rather than an invented number when only per-interval summaries survive.
Own the standard: what every team's result store retains, which figures a run report is permitted to carry, and the cost of raw retention weighed against the questions you will be asked after a bad run.
## What a percentile is, mechanically A percentile is a **rank statistic**. The 95th percentile of a set of response times is the value you reach after sorting every measured time and walking 95% of the way along the sorted list. It is defined only with respect to *one particular set of samples*. Add samples, remove samples or reweight them and the answer changes, and the only way to learn the new answer is to go back to the samples and rank them again. That is why the arithmetic that works on other summary figures does not work here. A total is a sum over samples, so totals from two windows add. An arithmetic average is a sum divided by a count, so two of them combine correctly as long as you carry both counts. A percentile is neither a sum nor a ratio of sums — it is a **position in a sorted list**, and positions in two separate lists do not combine into a position in the concatenated list. ## Why averaging the per-interval figures is not merely approximate Two independent things go wrong, and repairing one leaves the other standing. 1. **Every interval gets equal weight regardless of size.** A minute that served 12,000 requests and a minute that served 400 each contribute one number to the average. The pooled rank weights every *request* equally instead, so the two answers diverge whenever the request rate varied — which it does during ramp, during a pause, and during exactly the minutes when the system was struggling. 2. **The shape inside each interval has already been discarded.** A minute in which 30% of requests were slow and a minute in which 6% were slow can report almost the same per-minute 95th-percentile figure while contributing completely differently to the pooled rank, because the pooled rank depends on how many samples sit above a given value, not on where one interval's own cut happened to land. A stored scalar no longer carries that. Weighting the average by each interval's request count fixes the first problem and not the second. A count-weighted mean of quantiles is still a mean, and a mean of quantiles is not the quantile of the pooled samples. There is a subtler honesty problem underneath both. When a system degrades it often *completes fewer requests*, so the worst period contributes the fewest samples to the pooled rank. Pooling is the correct arithmetic, but it is not automatically the most alarming arithmetic — which is why the request count per interval belongs beside the figure in any report. ## The one thing the stored per-interval figures still tell you They **bound** the true value, and the bound costs nothing: - No more than 5% of any interval's samples exceed that interval's own 95th percentile. So no more than 5% of the pooled samples exceed the **largest** of the per-interval figures: the run's true 95th percentile is at most that maximum. - At the **smallest** of the per-interval figures, every interval has at most 95% of its own samples at or below it, so the pooled fraction at or below it is at most 95%: the run's true 95th percentile is at least that minimum. - Both bounds hold whatever the per-interval sample counts were, so they survive a wildly uneven run. The honest statement available from stored per-interval quantiles alone is therefore an interval — *"the run's 95th percentile lies somewhere between 420 ms and 1.9 s"* — and an average is one arbitrary point inside that interval with no property that recommends it over any other. ## What has to have been retained What the result store keeps decides what a merged run report can honestly compute, and the decision cannot be revisited once the run is over. | What was retained | What a whole-run figure can be | Cost | |---|---|---| | Every request's time | Any percentile, exactly; any regrouping by interval, operation or outcome | Grows with request count; by far the largest | | Counts per latency bucket, per interval | Any percentile, to within one bucket's width; counts add bucket by bucket | Small and roughly fixed per interval | | One quantile figure per interval | Only the minimum-to-maximum bound above; no merge is possible | Smallest, and it forecloses the question | A store that keeps only pre-aggregated per-interval summaries can draw an excellent trend over time and cannot answer *"what was this run's 95th percentile"* at all. A store that keeps raw per-request samples answers every regrouping question anyone thinks of afterwards, and pays for it in volume and in retention policy. A mergeable structure of bucket counts sits between the two, which is where long runs usually settle: the merge is a bucket-by-bucket addition, and the price is a bounded resolution error rather than a lost question. ## What to do - Decide retention **before** the run, working backwards from the figures the report will have to carry. - Merge raw samples or bucket counts; never merge stored quantiles. - If only per-interval quantiles survive, publish the bound and label it as a bound. - Keep the per-interval series as a series — a shape over time is what it is genuinely good for — and never collapse it into one number by averaging.
- What would the result store have to have kept for one honest whole-run 95th percentile to exist?Either every request's response time, so the whole set can be ranked once, or counts of requests per latency bucket recorded per interval, which add bucket by bucket into merged counts the percentile can be re-read from. A stored scalar quantile per interval supports neither — it can only be turned into a bound.
- If a team kept only per-interval figures, what is the strongest honest statement they can still make?That the run's 95th percentile lies between the smallest and the largest of the per-interval 95th percentiles. Both bounds hold no matter how many requests each interval served. Anything narrower — an average, a weighted average, the middle value of the series — is a number with no rank meaning and should not be compared against an agreed response-time limit.
- Does weighting the per-interval percentiles by request count make the average correct?No. Weighting repairs only the equal-weight problem; the result is still a mean of quantiles, and a mean of quantiles is not the quantile of the pooled samples. The pooled rank depends on how each interval's samples were spread above and below its own cut, and a stored scalar figure no longer carries that information.
Averaging per-minute 95th percentiles is like averaging a runner's finishing position across several races and calling it their position in one combined field. A position describes a field; it is not an amount you can add up.
saying these in an interview costs you the question
- Averages the per-interval figures and calls it the run
- Believes percentiles combine the way totals do
- Thinks a count-weighted average of quantiles is exact
- Keeps only per-interval summaries, then asks for one figure
- Discards the worst interval as an outlier before merging