In a performance run mixing several operation types, how can one slow operation vanish from the aggregate 95th percentile?
answer
- The pooled rank counts requests, not importance
- A rare operation sits above the cut
- Under five percent cannot move it
- Publish a row per operation type
- Merging only runs one way
basics
~20 sAn aggregate percentile ranks every request together, so weight follows request share, not importance. An operation carrying under 5% of requests can be entirely above the 95% cut and never move it. Report a distribution per operation type.
solid answer
~50 sAn aggregate percentile is a rank over the **pooled** requests of every operation type, so each type's influence is its share of the request count and nothing else. If report generation is 3% of the requests, then even when every one of those calls takes three seconds they all sit above the 95% cut, and the aggregate 95th-percentile figure is read entirely out of the cheap reads. The general rule: an operation type below one minus the percentile level cannot move that figure at all, however slow it is. Above that share it still competes with higher-volume types, and the aggregate then tracks the **mix** as much as the system — the figure moves when the proportions change with no operation's own behaviour changing. So the report has to keep a distribution and a request count per operation type, with the aggregate published only beside them.
code
pseudocode · 17 linesby_operation = group(run.samples, key = s.operation)
total = count(run.samples)
report_rows = []
for op, samples in by_operation:
share = count(samples) / total
report_rows.append({
operation: op,
requests: count(samples),
share: share,
middle: value_at_rank(sort(samples), 0.50),
high: value_at_rank(sort(samples), 0.95),
# below this share the operation cannot move the aggregate at all
invisible_in_aggregate_p95: share < 0.05
})
aggregate_p95 = value_at_rank(sort(run.samples), 0.95) # supplementary onlygo deeper
Recall that an aggregate figure ranks every request together, so an operation that is only a small fraction of the traffic barely shows up in it however slow it is.
Be ready to work the share rule out loud: an operation type below one minus the percentile level cannot move that figure at all, and above it the aggregate tracks the request mix as much as the system.
Show that you fix the operation grouping before the run, publish a row per operation with its request share, and check the mix before attributing an aggregate movement to a code change.
Own what a run report is allowed to claim: which figures teams may publish as headlines, why the mix is part of the result rather than merely how the run was arranged, and the cost of tagging every sample with its grouping.
## How an aggregate percentile is computed An aggregate percentile for a mixed workload is produced by pooling every request the run made, whatever operation it was, sorting the times and walking to the rank. Nothing in that procedure knows which operation a sample came from, and nothing in it knows which operation anybody cares about. **The weight an operation type carries is exactly its share of the request count.** That is a reasonable definition of "what a typical request experienced". It is a very poor definition of "is every part of this system healthy", and the two get confused constantly because the same number is used for both. ## The share rule, worked Suppose a run applies a hundred requests a second in a fixed mix: 1. **97 lightweight reads**, all completing under 80 ms. 2. **3 report generations**, every one of them taking about three seconds. Rank the pooled requests and walk to 95%. The three slow calls occupy the top 3% of the sorted list, so the value at the 95% rank sits comfortably inside the read distribution — around 80 ms. The run's headline figure says 80 ms while a whole operation type is a hundred times slower than that, on every single call, with no exception anywhere in the run. This generalises to a rule worth remembering: **an operation type whose share of requests is smaller than one minus the percentile level cannot influence that percentile at all.** Below 5% of requests, an operation is invisible in the aggregate 95th percentile; below 1%, invisible in the aggregate 99th. Making it slower changes nothing in the figure. Making it faster changes nothing either — which is the same problem wearing the other face, because it means the aggregate cannot show that the fix worked. ## Why the aggregate moves when the mix moves Above the invisibility threshold, an operation type competes with the others in proportion to volume, and the consequence is that the aggregate figure is a function of two things: how fast each operation is, and how many of each the run sent. - Shift the mix towards a slower operation and the aggregate worsens with no operation's own distribution changing. - Shift it towards a faster one and the aggregate improves the same way — a flattering result that no code change earned. - Two runs of the same build with different mixes produce different aggregates, so an aggregate figure without its mix beside it is not interpretable at all. The mix is therefore part of the result, not merely how the run was arranged. A run report that states the aggregate without the per-operation shares has published a number whose meaning it did not record. ## What the report has to keep separate | Reported | Answers | Does not answer | |---|---|---| | One aggregate percentile | What a typical request experienced across the whole mix | Whether any particular operation is healthy | | A percentile per operation type, from that type's own samples | Whether each operation is healthy, and which one changed | What a typical request experienced | | The request share per operation type | Whether an aggregate movement came from the mix | Anything about latency on its own | The minimum honest run report is a row per operation type carrying its request count, its share, a middle figure and a high figure computed **from that operation's own samples only**, with the aggregate offered underneath as a supplementary summary rather than as the headline. Two further points make the practice stick: - The same hiding recurs *inside* an operation type whenever it has cheap and expensive variants — a lookup by identifier against a lookup by free text, a small payload against a large one. Whatever variable splits the population is the variable the report should split on. - **Merging runs one way only.** Per-operation results can always be merged into an aggregate; an aggregate can never be split back into operations, because the pooled rank did not record which sample came from where. That asymmetry is why the grouping has to be decided before the run and recorded with the samples. ## What to do - Fix the operation grouping before the run and tag every sample with it. - Publish a per-operation row first, request count and share included. - Treat the aggregate as a supplementary figure, and never compare two aggregates without comparing their mixes as well. - When an aggregate moves, check the mix before you look for a code change.
- The aggregate 95th percentile worsened but every operation type's own figure is unchanged. What happened?The mix changed. The aggregate is a rank over the pooled requests, so shifting proportion towards a slower operation raises it even when no operation behaved differently. Compare the request shares between the two runs first, and only then look for a change in the system. It is also why the shares belong in the report alongside the figures.
- Why can per-operation figures not be recovered from an aggregate afterwards?Because the pooled rank discarded which operation each sample came from. Merging is one-way: per-operation results always compose into an aggregate, and an aggregate never decomposes back. That makes the operation grouping a decision taken before the run and recorded against every sample, not a reporting choice made afterwards.
saying these in an interview costs you the question
- Reports one aggregate percentile for a mixed workload
- Assumes a healthy aggregate means every operation is healthy
- Thinks the aggregate weights operations by business importance
- Tries to recover per-operation figures from a pooled result
- Explains an aggregate movement without checking the mix