A performance run passed on average response time although many requests were far slower. How do you write the pass rule so that cannot happen?
answer
- A summary of everyone hides a few
- Bind at a stated distribution point
- Pair the bound with an outer cap
- Say how many intervals may breach
- Central measure describes, never decides
basics
~20 sAttach the bound to a stated point in the distribution rather than a central measure, pair it with an outer cap on the worst reported interval, and say how many intervals may breach before the run fails.
solid answer
~50 sA central measure summarises everyone, so a slow minority is diluted by a fast majority and the rule can be satisfied while a real group of users is served badly. Write the rule as a **pair of bounds**: one at a named point well into the slow end - say ninety-five per cent of calls under 400 ms - and an **outer cap** on the worst observation or the highest reported interval, so a small but severely-served group can still fail the run. Then say **how intervals combine**: the bound holds in every reported interval, or in all but a stated number of them. Keep the central measure in the report as description, carrying no verdict. The rule still has to name the operation, the window and the workload, or a tail-aware pair of bounds is just a better-looking version of the same argument.
code
pseudocode · 13 linesverdict(run):
intervals = split(run, into = 1 minute)
per_interval = [ percentile(i.response_times, 95) for i in intervals ]
breached = count(v in per_interval where v > 400 ms)
worst = max(per_interval)
if breached > 1: return FAIL // agreed: at most one may breach
if worst > 1200 ms: return FAIL // outer cap, no exceptions
return PASS
// reported beside the outcome, deciding nothing:
report.middle = median(run.response_times)go deeper
Know that a run's pass rule can attach to different parts of the result, and that a limit on the average across all calls is a different claim from a limit that most calls individually have to meet.
Be able to write the rule: a bound at a named point in the distribution, an outer cap on what lies beyond it, and a statement of how many reported intervals may breach before the run fails.
Show that you have met the failure - a run recorded as a pass while users complained - and explain how you rewrote the rule so that the report and the complaint stopped disagreeing.
Own the shape teams reuse: a standard rule template in which a central measure alone may never carry an outcome, and the judgement about where the bound sits for each kind of operation.
## Why a central bound can pass a run that hurt people A central measure - the arithmetic average, or the middle value - summarises everyone. That is exactly what makes it a poor place to attach a pass rule. Response time distributions in real systems are lopsided: most calls finish quickly, a minority take far longer, and the fast majority pulls the summary down over the slow minority. A rule that binds only that summary can be satisfied while an identifiable group of users is served badly, and the run is recorded as a pass. The group is not small in practice. If one call in twenty takes four seconds and the rest take eighty milliseconds, the average is still comfortably under half a second. Every twentieth attempt is a bad experience, and the rule as written cannot see it. ## Bind at a stated point, then cap what lies beyond it The repair has two halves, and teams routinely do only the first. 1. **Move the bound into the slow end.** State the point in the distribution the bound attaches to - for example, ninety-five per cent of calls under 400 ms. The rule is now about the fraction of users you are unwilling to serve badly rather than about a summary of everyone. 2. **Cap what lies beyond that point.** A bound at the ninety-fifth percentile says nothing about the five per cent above it; they may be at 600 ms or at forty seconds. Add an outer clause - a limit on the worst observation, or on the highest reported interval - so a small but severely-served group can still fail the run. | Rule shape | What it catches | What it lets through | | --- | --- | --- | | Central measure only | A slowdown that affects nearly everyone | A minority served far worse than the rest | | Single high percentile | A minority served badly | Anything at all beyond that point | | High percentile plus outer cap | Both of the above | Only the rare breaches the rule deliberately tolerates | ## Say how intervals combine into one outcome A run is reported as a series of intervals, and a bound that holds in most of them and fails in one has no defined result unless the rule says so. Write the combining clause explicitly. Any of these is defensible: - The bound holds in **every** reported interval. - The bound holds in **all but a stated number** of intervals, with an outer cap no interval may exceed. - The bound is evaluated once over the whole reported window - in which case say so, and accept that a short severe episode can vanish inside it. Leaving the choice unwritten is the one option that is not defensible, because the choice is then made after the breach is visible, by someone who now knows which reading produces the answer they want. ## Keep the central measure, give it no authority None of this makes the average useless. It moves for reasons a high percentile does not: a change that makes everything slightly slower shows in the middle first, while a change that creates a new slow class shows only at the slow end. Report both, side by side, and let the pair describe the run - a move in one and not the other tells you whether a change hurt everybody a little or a few a lot. Only the bounded figures carry an outcome. There is a second reason to keep it. A rule expressed purely at the slow end can be gamed in the other direction: work can be made uniformly slower in a way that never breaches the high percentile bound but degrades the ordinary experience for everyone. Reporting the middle beside the bounded figures makes that visible even though it decides nothing. ## What still has to accompany the pair of bounds A tail-aware rule is still arguable if the rest of it is loose. The same document must name the operation being timed, the workload the bound holds under, the window and the interval size, and the vantage point the timings come from. A pair of bounds attached to "the system" under "normal load" has moved the ambiguity rather than removed it. ## The habit this prevents The failure this rule shape defends against is not a bad measurement. It is a good measurement summarised badly and then defended: the run passed, the number is in the report, and complaints from users are treated as anecdote. Once the rule bounds a stated point in the distribution and caps what lies beyond it, the report and the complaint stop disagreeing, because the report is finally measuring the thing the complaint is about.
- If the central measure carries no verdict, why report it at all?It describes the ordinary experience and it moves for reasons the slow end does not. A change that makes everything slightly slower shows there first; a change that creates a new slow class shows only at the far end. Reported together, the two say whether a change hurt everyone a little or a few a lot. Neither replaces the other, and only the bounded figures decide the run.
- A rule bounds ninety-five per cent of calls in every one-minute interval, and one interval breaches by a hair. Is the run a fail?By the rule as written, yes - and that is the point of writing it first. If a single marginal interval should not fail the run, the rule should have said all but one interval must hold, with an outer cap no interval may exceed. Amending the rule once the result is visible is exactly how a rule stops meaning anything.
An average water depth of one metre says nothing about the stretch of the river that is three metres deep.
saying these in an interview costs you the question
- Summarising a whole run with one central number
- Assuming a passing average means everyone was served well
- Bounding a high percentile with no cap beyond it
- Leaving unstated how many intervals may breach
- Loosening the bound once the result is visible