skip to content

How do you set a regression threshold for performance runs from measured noise rather than a round number?

level: seniorimportance: should knowfreq 46%

answer

  1. derive it, do not pick a round number
  2. above the measured spread, not below
  3. one level per metric, not one overall
  4. below the noise floor retrains the readers
  5. too wide hides losses that accumulate

basics

~20 s

Derive it from the measured run-to-run spread: set the alarm level a multiple above the spread of repeats, per metric, then check that size against what change actually matters. A round percentage chosen by habit is either noise or nothing.

solid answer

~50 s

Start from the spread you measured by repeating an unchanged configuration. The alarm level has to sit **above** that spread — commonly two to three times the observed dispersion, or outside the range the repeats covered — because anything at or below it will be tripped by the apparatus alone. Then sanity-check the number the other way: is a difference of that size one the team would actually act on? Set the level **per metric**, since a high percentile is typically far noisier than completed work over the same hold. If the smallest change worth catching is smaller than the noise, no threshold fixes that; you must reduce the noise (more repeats, a quieter environment, a longer hold, less randomness in the applied workload) or accept the detection limit honestly. A level below the noise floor produces alarms with no information, and the results stop being read at all.

code

pseudocode · 13 lines
pseudocode
noise = relative_spread_from_repeats(metric)   # e.g. 0.06 from 5 repeats
smallest_change_worth_acting_on = 0.03         # agreed with the service owner

if smallest_change_worth_acting_on <= noise:
    # not detectable by comparing single runs
    reduce_noise() or state_detection_limit(noise)

level = max(2 * noise, smallest_change_worth_acting_on)

delta = (new_figure - reference_figure) / reference_figure
flag_regression = delta > level

record(metric, noise, repeats, level, derived_on = today)

go deeper

for a junior

Be ready to say that a difference only counts as a regression once it is bigger than the amount the measurement moves on its own, and that this amount has to be measured rather than guessed.

for a middle

An interviewer expects the derivation: take the spread from repeats, place the level above it, set it separately for each metric, and record the spread it came from.

for a senior

Show that you can argue both failure modes — a level below the noise that retrains people to ignore results, and one so wide that real losses accumulate — and that you know what to do when the change that matters is smaller than the noise.

for a principal

Own the tradeoff between measurement cost and detection limit: how much repetition and how quiet an environment the organisation buys, and where you state honestly that a class of change is simply not detectable this way.

## The noise floor sets the smallest detectable change Comparing two performance runs is signal detection. The apparatus has a variability — the spread you get by repeating an unchanged configuration — and any real change smaller than that variability cannot be distinguished from it by looking at two numbers. That is not a policy choice; it is a property of the measurement. So a regression alarm level is not a value someone picks because it sounds strict. It is derived. A round percentage chosen by habit lands in one of two bad places: **below the noise**, where the apparatus trips it on its own, or **far above it**, where changes that genuinely matter slip past unremarked. Neither number is wrong because it is round; it is wrong because nobody checked it against the spread. It is also worth separating two instruments that get muddled. Comparing a run to an **earlier run** asks whether something changed, and that is what a regression alarm level governs. Comparing a run to an **agreed target** asks whether the system is good enough, which is a different question with a different number. A team can be well inside its target and still have lost twenty percent since last quarter, and only the first comparison will say so. ## Deriving the number 1. **Measure the spread** for the specific figure you will compare, on the environment you will compare in, from repeats of an unchanged configuration. 2. **Place the level above it.** Two to three times the observed dispersion, or just outside the full range the repeats covered, are both defensible starting points. The more repeats behind the spread figure, the tighter you can afford to be. 3. **Check it against what matters.** Ask what size of change the team would act on. If the derived level is far larger than that, you have a detection problem to solve, not a number to write down. 4. **Set it per metric.** One level covering every figure is a false economy: it will be too tight for the noisy ones and too loose for the steady ones. 5. **Record how it was derived** — the spread figure, the repeats behind it, the date — so that when it starts misbehaving the next person can tell whether the level was wrong or the environment changed. 6. **Re-derive it** when the environment, dataset or workload definition changes, and treat a level that has never been revisited as suspect. ## Both failure modes, and what each costs | Level relative to noise | What happens | What it costs the team | | --- | --- | --- | | At or below the spread | Trips on unchanged code, repeatedly | Alarms carry no information; people re-run until it passes, then stop reading results entirely, and a real loss arrives amid noise nobody trusts | | Just above the spread | Trips on real changes and occasionally on outliers | The intended behaviour: investigations are mostly worthwhile | | Far above the spread | Silent for everything short of a disaster | Real losses of an actionable size pass unnoticed, and repeated sub-level slips accumulate into a large decline nobody ever decided to accept | The first row is the one teams walk into. It is worth being explicit about the damage: an alarm that fires without information does not merely waste the investigation, it **retrains the readers**. Once a result is assumed to be noise, the one that was not noise is assumed to be noise too, and the whole comparison stops being a control. The third row is quieter and is the reason a level should not simply be widened whenever it becomes annoying. Widening is the cheapest response to a noisy environment and the one that removes the measurement's value while leaving its cost in place. ## When the change that matters is smaller than the noise This is the genuinely hard case and the one a strong answer names. Options, roughly in order of cost: - **More repeats.** Comparing the middle of several repeats on each side rather than one figure against one figure shrinks the effective spread, at a linear cost in machine time. - **A quieter environment.** Dedicated hardware, no neighbours, no background activity, pinned placement — usually the single largest reduction available. - **A longer measured hold.** More completed work per run reduces the influence of short disturbances, though it does nothing about differences between provisionings. - **Less randomness in the applied workload.** A fixed sequence rather than a randomly drawn one removes one source of variation, at the cost of some realism. - **A more stable figure.** A steadier metric may detect the same underlying change earlier than the noisiest one, if it moves for the same reason. - **Accept the detection limit.** Write down that changes below a given size are invisible to this comparison, and rely on something else for them. An honest stated limit is far better than an alarm level pretending to a precision the apparatus does not have. What is not an option is setting the level below the noise and hoping the readers will use judgement. They will, for about a month.

  • A team widens the alarm level every time it becomes annoying. What is the failure there?
    Widening treats the symptom and destroys the instrument. Each widening raises the size of loss that can pass unnoticed, and because the changes are individually small nobody sees the accumulated decline. The right response to frequent trips is to check whether the level sits below the measured spread and, if it does, reduce the noise or state the detection limit rather than quietly raising the number.
  • Should the alarm level be the same for completed work and for a high percentile?
    No. Their noise differs substantially: work completed over a long hold averages out short disturbances, while a high percentile is driven by exactly those disturbances and is typically the noisiest figure measured. One level applied to both is simultaneously too tight for the percentile and too loose for the throughput figure, so derive each from its own spread.
  • How would you notice that a derived level has gone stale?
    Watch the trip rate against unchanged code. A level that suddenly trips often, with no change to explain it, usually means the environment got noisier rather than that the system got slower. A level that has not tripped in a very long time may mean the opposite. Both are prompts to re-measure the spread and re-derive rather than to argue about the last result.

A smoke alarm set to trigger on toast is removed from the ceiling within a week, and the fire it was bought for goes unannounced.

saying these in an interview costs you the question

  • Picks a round percentage without measuring spread
  • Uses one alarm level for every metric
  • Widens the level whenever it becomes inconvenient
  • Treats every trip as a confirmed regression
  • Ignores that a level below the noise floor teaches people to ignore results
  • Assumes a tighter level is always the stricter engineering choice