Before you call a difference between two performance runs a regression, what must you measure first?
answer
- know the noise before naming a regression
- repeat the same configuration unchanged
- report dispersion, not just an average
- repeat across provisionings, not only back to back
- spread differs per figure and expires
basics
~10 sMeasure the run-to-run spread first: repeat one unchanged configuration several times and record how much the compared figure varies on its own. A difference smaller than that spread is not evidence of anything.
solid answer
~50 sYou need the **noise floor** — how far the figure moves when nothing has changed. Get it by repeating the identical configuration several times: same build, same dataset, same environment specification, same warm-up, same applied workload. Record the figure from each repeat and report the dispersion, not just an average: the range across repeats, or the spread relative to their middle value. Do some of the repeats across separate provisionings of the environment rather than only back to back, because creating the environment afresh is itself a source of variation. Then any gap between two runs is read against that spread. Without it, a four percent move is uninterpretable — it might be a genuine loss or it might be a Tuesday. Noise also differs per figure and per environment, so measure it for the figure you actually compare and refresh it when the environment changes.
code
pseudocode · 15 linesfigures = []
repeat 5 times:
provision environment from the same specification
deploy the unchanged build
load the same dataset snapshot
warm with the same operation mix for the same duration
apply the same workload for the same hold
figures.append(reported_95th_percentile_ms)
centre = median(figures)
spread = (max(figures) - min(figures)) / centre
record(individual = figures, centre = centre, spread = spread)
# spread is the noise floor: differences below it are not evidencego deeper
Be ready to say that the same test run twice does not give the same number, and that you have to know how much it moves on its own before a difference can mean anything.
An interviewer expects the mechanics: repeat the identical configuration several times, report the dispersion of the figures rather than only their average, and measure it separately for each figure you compare.
Demonstrate that you repeat across separate provisionings, inspect the individual figures for outliers or two clusters, and treat the spread as a dated measurement that expires when the environment changes.
Own the policy: how much repetition the team pays for, which environments are quiet enough to be worth measuring in, and how the spread is published so that every result is read against it.
## Why the spread has to exist first A performance figure is not a constant. Run the same build against the same data on the same machine twice and the numbers differ — because of scheduling, cache state, disk and network variation, shared infrastructure, background activity, and the randomness inside the applied workload itself. That variation has a size. Until you know it, every comparison you make is a claim without a scale. The consequences of skipping it run in both directions. A team without a spread figure chases differences that were never there, burning days on a *regression* that repeats would have dissolved. The same team, a month later, waves through a genuine loss because *it is only a few percent* — with no idea whether a few percent is inside the noise or five times it. Both failures come from the same missing number. The figure is also the input to almost everything else in this area: it sets the smallest change that can be detected at all, it sets any regression alarm level worth having, and it is the evidence that a shift in level was reproduced rather than observed once. ## How to measure it 1. **Fix the configuration completely.** One build, one dataset snapshot, one environment specification, one warm-up procedure, one workload definition, one measured window. Nothing is under study; you are measuring the apparatus. 2. **Repeat it several times.** Three is a hint, five is a working minimum, more is better for a noisy environment. Fewer repeats give a spread figure that is itself noisy. 3. **Vary what you cannot hold, deliberately.** Run some repeats back to back and some across separate provisionings of the environment, on different days or times. Back-to-back repeats measure only the variation *within* one setup and will flatter you; the variation that bites in practice includes being placed on a different physical host. 4. **Record the individual figures, not only a summary.** You want to see whether the repeats cluster with one outlier, or spread evenly, or fall into two groups — those three shapes have different causes and different remedies. 5. **Express the dispersion.** The range across repeats relative to their middle value is enough for most teams and is easy to explain. A standard deviation is fine if the repeats look unimodal. Publish it beside the result, in the same units and framing as the comparison it governs. 6. **Do it per figure.** Completed work over a long hold is usually steadier than a high percentile, and the slow end of the latency distribution is usually the noisiest thing you measure — which is unfortunate, because it is often the thing you care about most. ## What the number is worth, and when it expires | Situation | What the spread figure tells you | | --- | --- | | Repeats cluster tightly | Small differences are meaningful; you can detect fine changes | | Repeats spread widely | Only large differences are detectable; fine changes need more repeats or a quieter environment | | One repeat far from the rest | Something intermittent is present; investigate it before trusting any comparison | | Repeats fall into two clear groups | The apparatus has two states — placement, cache warmth or a configuration that is not fully pinned | A spread figure describes one apparatus at one time. It expires when the environment changes shape, when the dataset is refreshed, when the workload definition is edited, or when the infrastructure underneath is re-provisioned onto different hardware. Treat it as a measurement with a date, re-take it periodically, and be suspicious when it improves dramatically without explanation — that usually means the repeats stopped being independent. ## Common ways this goes wrong - **Taking the best run.** The fastest of five is not a reference; it is the left tail of your noise, and comparing a later best-of-five to it is comparing two extremes. - **Repeating without re-provisioning.** All-back-to-back repeats systematically understate spread and produce alarm levels that trip in real use. - **Reusing a spread figure across metrics.** Applying the completed-work spread to a high percentile understates the noise where it is worst. - **Treating the gap between two builds as the spread.** That gap is the thing under test, not the apparatus's variability. - **Never refreshing it.** A figure measured on last year's environment says nothing about this one. The habit that makes this stick is publishing the spread beside every reported result, so that any reader — including the version of you reading it next quarter — can immediately see whether a difference of a given size means anything at all.
- Why is repeating a configuration back to back not enough?Back-to-back repeats hold constant everything about how the environment was created and placed, so they measure only the variation inside one setup. The variation that actually affects comparisons includes landing on different physical hardware, different neighbours and a freshly built dataset. Spread measured only back to back is systematically optimistic, and alarm levels derived from it trip constantly in real use.
- The repeats fall into two clear groups rather than scattering. What does that suggest?Two groups mean the apparatus has two states rather than random jitter. Common causes are placement on two different hardware generations, a cache or pool that is warm on some repeats and not others, or a setting that is not actually pinned. Find and eliminate the state before reporting a spread figure, because a bimodal spread describes neither group honestly.
- How many repeats are enough?Enough that the spread figure itself stops moving much as you add more. Three gives a hint, five is a workable minimum for a reasonably quiet environment, and a noisy shared environment may need considerably more. The practical test is to compute the spread from the first three and again from all of them; if the two disagree substantially, you have not repeated enough.
saying these in an interview costs you the question
- Calls any difference a regression without repeating anything
- Uses the fastest run as the reference figure
- Repeats only back to back on one provisioned environment
- Reports an average of repeats but never their spread
- Reuses one spread figure for every metric
- Never re-measures spread after the environment changes