skip to content

Before a performance run starts, what must its pass rule fix for the run's verdict to mean anything?

level: middleimportance: must knowfreq 70%

answer

  1. Agreed before any numbers exist
  2. Five parts, none left open
  3. Measurement, distribution point, window, workload, vantage
  4. An open part is filled in afterwards
  5. Name who may accept a breach

basics

~20 s

Fix five things before the run starts: which measurement, at which point in its distribution, over which window, under which workload, and read from which vantage point. Anything left open is chosen afterwards to suit the numbers.

solid answer

~50 s

A pass rule is only worth having if every part of it is written down before the run and nobody can pick a part afterwards. Fix **the measurement** - which quantity, for which operation; **the point in the distribution** the bound attaches to, since a central measure and the slow end of the same run disagree by design; **the window** the numbers come from and the intervals they are reported in; **the workload** the rule holds under, meaning arrival rate, operation mix and data size; and **the vantage point**, since a call timed by its caller and the same call timed inside the service are different durations. Add the number with its unit, and name who may accept a breach and record it. An unstated part is not neutral - once the numbers exist, whoever wants a pass reads it the way that gives one.

code

pseudocode · 12 lines
pseudocode
pass_rule for run "checkout-hold-2026-03-11":
  measurement   = response time of the "submit order" operation
  point         = 95th percentile, plus the worst 1-minute interval
  bound         = 400 ms (95th percentile), 1200 ms (worst interval)
  window        = the 30-minute hold, reported in 1-minute intervals
  workload      = 200 arrivals/second, 7 browse : 1 submit,
                  catalogue of 2,000,000 items
  vantage       = timings taken by the calling client
  breach_policy = service owner may accept one breached interval,
                  recorded with a reason

verdict = PASS only if every clause above holds

go deeper

for a junior

Be ready to say that a performance run needs a rule agreed before it starts, and to name what that rule fixes: the measurement, the point in the distribution, the window, the workload, and where the timings are taken.

for a middle

Explain each part's mechanics: why a time bound with no workload attached can be met by driving the system more gently, and why the same call has several honest durations depending on where it is timed.

for a senior

Bring a run whose rule left a part open and the argument that followed once the numbers arrived. Say what you changed in the rule, and how a breach becomes a recorded acceptance rather than a debate.

for a principal

Own the standard: a template every team fills in before a run is commissioned, a named person who may accept a breach, and a way of keeping the practice from decaying into a form nobody reads.

## What a pass rule is, and why it must exist first A **pass rule** is the sentence that turns a performance run's numbers into a verdict, written down and agreed before the run is started. It is not the product requirement it descends from, and it is not a promise to anyone outside the team. Its single job is to be **decidable**: two people reading the same result document must reach the same verdict without negotiating. It has to exist beforehand because a run does not produce *a* number. It produces a cloud of them - thousands of timings, split by operation, spread across intervals, taken at several points in the system. Almost any story can be told from that cloud once it exists, and fixing the rule first removes the freedom to pick which one. ## The five parts | Part | The question it settles | What leaving it open licenses | | --- | --- | --- | | Measurement | Which quantity, for which operation? | Quoting whichever operation happened to be fast | | Distribution point | A central measure, a high percentile, the worst observation? | Retreating to the central measure when the slow end looks bad | | Window | Which stretch of the run, reported in which intervals? | Trimming away the stretch that held the bad minutes | | Workload | Which arrival rate, operation mix and data size? | Passing by driving the system more gently than intended | | Vantage point | Where is the timing taken? | Reading whichever of two honest measurements is friendlier | 1. **The measurement.** "Response time" is not yet a measurement. Response time *of which operation*, counted from which moment to which moment? A rule that binds one figure over a mixed workload is satisfied by a flood of cheap reads while the operation anyone cares about is slow. 2. **The distribution point.** A bound has to attach to a stated place in the spread of results - the middle, a high percentile, or the worst observation. Those disagree by design, and a rule that does not say which one it means is not yet a rule. 3. **The window.** Numbers must come from a named stretch of the run and be reported in stated intervals, because the interval size changes what a breach even looks like: one bad thirty-second stretch disappears inside a ten-minute interval and dominates a one-minute one. 4. **The workload.** A time bound is meaningless without the load it holds under. Arrival rate, the mix of operations, the size and shape of the data, and the number of concurrent callers all belong in the rule, because every one of them moves the number. 5. **The vantage point.** The same call has several honest durations: timed by the caller, timed at the edge of the service, timed around the handler inside it. They differ by connection setup, by waiting before the work starts, and by any retry. The rule must name which span it bounds. ## Two clauses teams forget - **The number with its unit, written once.** "Under 400 ms" and "under 0.4 seconds" are the same rule; "under 400" is not a rule at all, and a figure each reader recomputes is a figure each reader can recompute favourably. - **Who may accept a breach, and how.** Runs breach for reasons everybody agrees are irrelevant. The rule should name a person who may accept a breach and require that the acceptance be recorded with a reason. An accepted breach with a name on it is a decision; an argument after the fact is not. ## A worked example Suppose a team agrees, before starting: the measurement is the response time of the *submit order* operation; the bound is 400 ms at the 95th percentile with no one-minute interval above 1200 ms; the window is the thirty minutes after load has stabilised, reported in one-minute intervals; the workload is 200 arrivals per second at seven browse calls per submit against a two-million-item catalogue; timings are taken by the caller that issued the request; and the service owner may accept a single breached interval, recorded with a reason. Nothing there is clever. Its virtue is that when the run finishes, the verdict falls out of the document rather than out of a meeting. ## What an unstated part licenses Every gap gets filled in afterwards, by whoever has the strongest interest in a particular outcome: - No workload named, so the run is repeated at a gentler rate until it passes. - No distribution point named, so the central measure is quoted and the slow end goes unmentioned. - No window named, so the report covers the calm stretch and stops there. - No vantage named, so the internal timings are shown and the caller-side ones are not. - No breach owner named, so the breach is explained away rather than accepted, and nobody records that it happened. None of those require dishonesty. They are what reasonable people do with an ambiguous rule and a deadline, so the fix is structural rather than moral: fix the rule while the numbers do not yet exist and the temptation never arises. Worth saying plainly, too - this is a rule for one run and nothing more. A long-lived reliability commitment about live traffic is a different subject with different owners.

  • A pass rule bounds response time but says nothing about where the timing is taken. What goes wrong?
    Every call has several honest durations - measured by the caller, at the service edge, or around the handler inside it. They differ by connection setup, by waiting before work starts, and by retries. With the vantage unstated, one reader passes the run on internal timings while another fails it on caller timings, and neither is misreading the rule.
  • The rule fixes one bound, but the run produces many intervals. What else has to be said?
    How intervals combine into a verdict: whether the bound must hold in every reported interval, whether a stated number may breach, and what the worst single interval may reach. Without that clause, a run with one catastrophic minute and fifty good ones has no defined outcome, and the argument happens with the numbers already on screen.
  • Who should agree the rule before the run, and what does that agreement buy?
    Whoever will act on the outcome - the owner of the service and the person who commissioned the run. Their agreement buys a decidable result: a breach becomes an accepted exception with a name and a reason attached, rather than a debate about whether the rule ever meant that. It also makes the run repeatable by someone else.

It is the difference between agreeing the finish line before the race and painting one once everybody has stopped running.

saying these in an interview costs you the question

  • Deciding what counts as a pass after seeing the numbers
  • Stating a time bound with no workload attached
  • Treating caller-side and service-side timings as interchangeable
  • Quoting a bound without its unit or its window
  • Assuming an unstated part will be read charitably