When is adopting a new performance reference run legitimate rather than moving the goalposts?
answer
- attributed and intentional, or it is a loss
- reproduced across repeats, not seen once
- still meets the obligation with margin
- keep the old reference, record the decision
- watch cumulative drift from a long-term anchor
basics
~20 sIt is legitimate when the shift in level has an identified, intentional cause, was reproduced across repeats, and still meets the obligation the team owes. It is a moved goalpost when the reason is that meeting the old figure became inconvenient.
solid answer
~50 sTwo positions are defensible and the evidence picks between them. **Adopt the new reference** when the cause of the shift is identified and deliberate, reproduced across repeats rather than seen once, and the new level still satisfies the obligation with margin — holding a stale reference makes every later run fail for a reason nobody will act on, and the comparison stops being read. **Refuse it** when the cause is unattributed, or when the only argument is that the old figure became awkward to meet; the difference is then a loss carrying a debt, and the reference should stand so the debt stays visible. Either way, record the cause, before and after figures with their spread, the approver, and whether it is permanent. Each adoption breaks the chain back to the original level, and small justified steps sum to a decline nobody chose.
code
yaml · 19 linesreference_change:
date: 2026-04-02
metric: checkout 95th percentile, measured client side
previous_reference:
run_id: perf-2026-01-08-c
value_ms: 240
spread_pct: 5
new_reference:
run_id: perf-2026-04-02-b
value_ms: 305
spread_pct: 5
cause: per-item tax lookup added to the checkout path
evidence: reproduced in 5 of 5 repeats; absent in repeats of the previous build
outside_noise: yes, shift 27 percent against a 5 percent spread
meets_obligation: yes, agreed limit 400 ms, margin 95 ms
permanence: temporary
recovery: batch the lookup, owner platform team, review 2026-07-01
approved_by: service owner
drift_since_anchor_ms: 65 of an agreed 100go deeper
Be ready to say that the earlier result later runs are compared against can be replaced, that replacing it is a decision rather than housekeeping, and that the old figures are kept rather than overwritten.
An interviewer expects you to distinguish an attributed, reproduced and intentional shift from an unexplained one, and to know that a shift inside the measured run-to-run spread is not a shift at all.
Show the evidence you would demand before signing: attribution, repeats on both configurations, the size against the spread, remaining margin against the obligation, and a recovery owner for anything temporary.
Own both sides of the tradeoff. Argue when holding a stale reference destroys the comparison, argue when adoption converts debt into accounting, and put a mechanism against cumulative drift in place before you need it.
## Two defensible positions A reference run is the earlier result later runs are compared against. Sooner or later a change moves the level and someone proposes that the new figures become the reference. The question is genuinely open, and both answers are defensible depending on the evidence. **The case for adopting.** Some level shifts are the intended consequence of a decision the organisation made with its eyes open: a feature that costs work per request, a security control on a hot path, a data structure traded for correctness. Refusing to move the reference in that situation does not preserve rigour; it produces a permanent red state that nobody can act on, and a permanently failing comparison is quickly reclassified as noise. Worse, it hides the *next* regression underneath the one everybody has agreed to ignore. **The case for refusing.** Other shifts are losses wearing an explanation. The cause is vaguely described, the change was never reproduced, or the only real argument is that meeting the old figure had become inconvenient. Adopting there converts an engineering debt into an accounting adjustment. The reference exists precisely to keep the debt visible; moving it is the one action that makes the debt disappear without anyone paying it. ## The evidence the decision needs A proposal to adopt a new reference should have to answer all of these, in writing: 1. **Is the cause identified?** Not *the release was slower*, but which change and which mechanism. An unattributed shift is an open investigation, not a candidate for adoption. 2. **Was it intentional?** Something the team chose, or something that happened to it. Both can be adopted, but only the first is adopted without a recovery plan. 3. **Was it reproduced?** The shift must appear across repeats of the new configuration and not in repeats of the old one. A single run showing a difference is a hypothesis. 4. **Is it larger than the measured spread?** If the shift sits inside the run-to-run noise, there is nothing to adopt; the reference has not moved. 5. **Does the new level still meet the obligation, with margin?** Adopting a reference does not renegotiate what the system owes its users. If the new level consumes the remaining margin, the decision is not a re-baselining at all — it is a decision to ship at the edge. 6. **Is it permanent or temporary?** A temporary concession needs an owner, a plan and a date, or it is a permanent one with better manners. 7. **Who approves?** The person accountable for the service, not the person whose change caused the shift. ## What gets recorded The record is the difference between a re-baselining and a quiet edit. It should carry, at minimum: - The **date** and the run that becomes the new reference, with its full comparability record. - The **previous reference** and its figures — kept, never overwritten. - **Before and after values with their spreads**, so a later reader can see the shift was outside the noise. - The **cause** and the reason it was accepted. - The **approver**. - **Permanence**: permanent, or temporary with an owner and a recovery date. - The **cumulative drift** from a long-term anchor, so the sum of past adoptions is visible at the moment a new one is proposed. That last line is the one teams omit and the one that does the most work. ## What is lost every time | What breaks | Why it matters | | --- | --- | | The chain of comparison | You can compare to the current reference, but the direct link back to the original level is gone unless drift is tracked separately | | The sense of normal | Each adoption resets what the team believes the system costs, and the new level immediately feels like it was always so | | The incentive | Explaining a shift becomes cheaper than fixing it, and teams optimise for the cheaper path | | The ratchet guard | Individually justified steps of a few percent each sum, over a year, to a decline nobody would have approved in one go | The ratchet is the serious one. Every step can be defensible on its own evidence and the total still be indefensible. That is why the anchor matters: keep one long-term reference that is *not* re-baselined, track cumulative drift against it, and require the total — not the increment — to be re-approved past an agreed amount. Some teams also cap adoptions per period, which is cruder but has the same effect of forcing the conversation upward. ## How the decision usually plays out Asked to sign this off, a strong answer refuses the framing of a single yes or no. It asks for the attribution first, insists the shift be reproduced against the measured spread, checks the remaining margin against what the service owes, and then chooses between adoption with a recorded rationale and rejection with an owner for the recovery. It also says the uncomfortable part out loud: if the team cannot articulate why the level moved, the honest answer is that the reference stands and the investigation continues, however inconvenient that is for the release everyone wants to ship this week.
- A shift in level cannot be attributed to any particular change. Do you adopt the new figures?No. An unattributed shift is an open investigation, and adopting it closes the investigation by definition rather than by evidence. Keep the reference, record the discrepancy, and look for the cause — most often a change in the dataset, the environment or the applied workload rather than in the code. If the cause is genuinely never found, that itself is a finding about the measurement, and it should be written down instead of absorbed.
- How do you stop a series of individually justified adoptions from becoming a large decline?Keep one long-term anchor that is never re-baselined, and record cumulative drift against it every time a new reference is adopted. Require the total, not the increment, to be re-approved once it passes an agreed amount. The mechanism works because it moves the conversation from a small local decision, which is always easy to defend, to the aggregate, which is the thing nobody would have approved in one step.
- Is there a case for adopting a new reference that is faster than the old one?Yes, and it deserves the same discipline. A genuine improvement should become the reference so that later losses are measured from the new level rather than being absorbed by the headroom the improvement created. The same evidence applies: attributed, intentional, reproduced across repeats, outside the measured spread, and recorded — otherwise a measurement artefact gets locked in as the new expectation.
saying these in an interview costs you the question
- Adopts a new reference because the old figure was inconvenient
- Overwrites the previous reference instead of keeping it
- Accepts a shift seen in one run without repeats
- Never tracks cumulative drift across successive adoptions
- Lets the author of the change approve their own reference move
- Treats a temporary concession as needing no owner or date