skip to content

questions

4

What has to be held constant for two performance runs to be comparable at all?

level: middleimportance: must knowfreq 62%

answer

  1. one variable moves, everything else does not
  2. comparable to another run, not in general
  3. build, data, environment, warm state, workload
  4. same measurement point on both sides
  5. an unrecorded axis makes the figure unreusable

basics

~20 s

Comparable performance runs hold the same build, the same dataset size and shape, the same environment and its neighbours, the same warm state and the same applied workload. When one of those differs, a change in the numbers cannot be attributed.

solid answer

~50 s

Two performance runs can be compared only when everything except the one thing under study was the same. In practice that is five or six axes: the **build and its runtime settings**; the **dataset** — its size and its distribution, not just a row count; the **environment**, including the machine class, whatever else shares it and the state of the dependencies behind it; the **warm state** the system was in when measurement began; the **applied workload** — operation mix, arrival pattern and hold duration; and the **measurement point**, since a figure taken at the client and one taken inside the service are different quantities. If two axes move at once, the difference is real but unattributable. That is also why a run report has to record those axes: without them, the older figure cannot be used as a reference a quarter later.

code

yaml · 21 lines
yaml
run_manifest:
  run_id: perf-2026-03-14-a
  build:
    artifact_id: checkout-service-4812
    settings: { worker_pool: 64, log_level: warn }
  dataset:
    snapshot: orders-2026-03-01
    rows: 12000000
    shape: 80 percent of orders sit on 5 percent of accounts
  environment:
    class: 8-core-32gb
    shared_with: nothing else scheduled
    dependencies: dedicated instances, not shared
  warm_state: same mix applied for 10 minutes before measurement
  workload:
    mix: { browse: 70, search: 20, checkout: 10 }
    arrival: { workload_model: open, rate_per_second: 400 }
    hold_minutes: 30
  measurement_point: client side, includes network time
  repeats: 5
  spread_pct: 6

go deeper

for a junior

Be ready to say that a performance figure describes a set of conditions rather than the system in general, and to name the obvious ones: the same build, the same data, the same machine, the same applied load.

for a middle

An interviewer expects you to list the axes and explain what moves when one is not held — a neighbour on the host, a dataset that grew between runs, a cold start compared against a warm reference.

for a senior

Show that you treat the conditions as part of the result: a generated manifest beside the numbers, and the habit of refusing a comparison whose manifests differ on an axis nobody meant to change.

for a principal

Own the standard. Decide what every run report must carry, who keeps reference runs and their environments, and how comparability survives re-provisioning, a dataset refresh and a team handover.

## Comparability is a property of a pair, not of one run A performance run measures a system under a particular set of conditions. Its numbers describe that combination, not the system in the abstract. The moment two runs are placed side by side and the gap is read as *the change made it slower*, an experiment is being claimed: one variable moved, everything else was held. If in fact three things moved, the gap is real but **unattributable** — it is the sum of every uncontrolled difference, and no statistic recovers the parts afterwards. So comparability is not a quality a single run can have. A run is comparable *to another run*, or it is not. This also means the axes are only worth holding pairwise: you do not need a perfect environment, you need the *same* environment twice. A distinction worth drawing early: **comparability is not realism**. Two runs on a modest environment can detect a relative change between them perfectly well, even though neither says much about production capacity in absolute terms. Realism governs how far a result generalises; comparability governs whether a difference means anything at all. Teams often throw away a valid regression signal on the grounds that the environment is not production-like, which confuses the two questions. ## The axes that have to be held | Axis | What it covers | What moves if you do not hold it | | --- | --- | --- | | Build and settings | The exact artefact plus pool sizes, feature switches, logging verbosity | A single verbose logging setting can move the slow end of the latency distribution | | Dataset | Row counts **and** distribution: key skew, how many rows a typical request touches | The plan the data store chooses changes as data grows, with no code change at all | | Environment | Machine class, storage, network path, and whatever else runs on the same host | A neighbour's overnight job arrives in your results as a regression | | Warm state | Whether caches, connection pools and lazily initialised paths were in the same condition | A cold start compared against a warm reference looks like a large loss | | Applied workload | Operation mix, arrival pattern, concurrency, think time, hold duration | A slightly different mix changes the average cost of a request | | Dependencies | Whether downstream systems were real, shared or substituted, and in what state | A shared downstream under someone else's load becomes your slowdown | | Measurement point | Where the timing was taken and what it includes | A client-side figure and a server-side figure are different quantities, not two views of one | ## What the report has to state The axes only help if the next person can check them, which is a records problem rather than an engineering one. A run report that will still be usable a quarter later carries: 1. The **build identifier** and the settings that differ from the shipped defaults. 2. How the **dataset** was produced, how large it was, and what shape it had. 3. The **environment specification** and what shared it during the measured window. 4. How the system was **warmed**, and for how long, before measurement began. 5. The **workload definition** — mix, arrival pattern, concurrency, hold duration. 6. The **measurement point** and the window the figures summarise. 7. The date, and the **run-to-run spread** if repeats were executed. A figure without that record is not a reference. It is a memory, and a quarter later nobody can tell whether a new run matches its conditions or whether the old machine even exists any more. ## Where comparability is lost quietly None of these announce themselves; each has retired a perfectly good comparison: - A dataset that grew between runs because the earlier run left its rows behind. - An environment re-provisioned onto a different machine generation with the same label. - A default that changed inside a dependency rather than in your own build. - Repeats taken at different times of day on infrastructure shared with other teams. - A workload definition edited *slightly* to work around a scripting problem mid-campaign. - A report that records the numbers but not the conditions that produced them. The cheap discipline is to emit a **run manifest** with every result, generated rather than typed, and to refuse to compare two runs whose manifests differ on an axis nobody intended to change. That refusal is not pedantry: it is the only thing standing between a team and a quarter spent chasing a regression that was a neighbour on a host. Finally, keep separate the two things a run's numbers get compared to. Comparing a run to an **earlier run** asks whether something changed; comparing it to an **agreed target** asks whether it is good enough. The axes above make the first comparison possible. They do not, on their own, make either verdict.

  • Which of those axes must a run report write down for its numbers to still be usable a quarter later?
    All of them. The build identifier and its settings, how the dataset was produced and its size and shape, the environment specification and what shared it, how the system was warmed, the full workload definition, and the measurement point and window. A number without that record cannot serve as a reference, because nobody can reconstruct whether a later run matches the conditions that produced it.
  • Does a run have to resemble production before it can be compared to another run?
    No. Comparability is a same-to-same property between two runs; how far the result generalises to production is a separate question. A pair of runs on a modest environment can detect a relative change reliably, even though what their absolute figures imply about production capacity is another matter entirely. Confusing the two makes teams discard valid regression signals because the environment is not production-like.
  • Two runs match on every axis except that one used substituted dependencies. Can they be compared?
    Only for what happens above the substitution. If both runs substitute the same dependency in the same way, a difference still isolates changes in your own service. If one used a real dependency and the other a substitute, the gap includes the difference between the substitute and the real thing, which is usually larger than whatever you were trying to measure.

It is the difference between weighing yourself twice on the same scales at the same time of day, and weighing yourself on two different scales a season apart. Only the first tells you anything about you.

saying these in an interview costs you the question

  • Compares runs from different builds without saying so
  • Treats a dataset row count as the only data axis
  • Ignores what else was running on the same host
  • Reports figures with no record of the conditions
  • Measures one run from cold and the other warm
  • Thinks comparability requires the environment to mirror production
open as a page

Before you call a difference between two performance runs a regression, what must you measure first?

level: middleimportance: should knowfreq 52%

basics

~10 s

Measure the run-to-run spread first: repeat one unchanged configuration several times and record how much the compared figure varies on its own. A difference smaller than that spread is not evidence of anything.

open as a page

How do you set a regression threshold for performance runs from measured noise rather than a round number?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Derive it from the measured run-to-run spread: set the alarm level a multiple above the spread of repeats, per metric, then check that size against what change actually matters. A round percentage chosen by habit is either noise or nothing.

open as a page

When is adopting a new performance reference run legitimate rather than moving the goalposts?

level: principalimportance: should knowfreq 38%

basics

~20 s

It is legitimate when the shift in level has an identified, intentional cause, was reproduced across repeats, and still meets the obligation the team owes. It is a moved goalpost when the reason is that meeting the old figure became inconvenient.

open as a page