skip to content

Your treatment lifts sign-ups but degrades p95 page latency by 20ms — how do you decide whether to ship?

level: seniorimportance: must knowfreq 56%

answer

  1. verify the signal before trading it away
  2. interval width, then instrumentation symmetry
  3. the threshold was set beforehand for a reason
  4. pooled p95 hides device segments
  5. cumulative budget, not one decision

basics

~10 s

Confirm the 20ms is real and not an instrumentation artefact, compare it against the pre-agreed latency budget, segment it by device, then ship, fix, or escalate. Never move the threshold to fit the result.

solid answer

~50 s

First establish that the degradation is real: check the confidence interval on the latency delta, confirm both arms are instrumented identically, and rule out artefacts such as cold caches or a new code path being measured differently. Second, compare 20ms against the threshold agreed before launch. If it is inside the budget, the guardrail did its job and you ship. If it breaches, the decision is not mine alone to overturn — it goes to the named owner as a documented tradeoff. Third, look past the pooled number: p95 hides segments, and low-end devices or slow networks often absorb several times the average degradation. Fourth, ask whether the cost is necessary — extra latency is frequently an implementation artefact that a deferred load or an async call removes, turning a tradeoff into a fix. The one thing I will not do is move the threshold after seeing the result.

go deeper

for a junior

Be ready to say that a primary-metric win does not automatically override a guardrail, and that the latency threshold should have been agreed before the experiment started.

for a middle

Explain how you verify the degradation is real — interval width, instrumentation symmetry, cold-start effects — and why p95 moving while the median is flat points at a specific slow path.

for a senior

Demonstrate the full sequence: verify, compare against the pre-set budget, segment by device and flow, look for a fix that removes the cost, then choose among ship, ship-with-fix, escalate, or block.

for a principal

Own the accounting problem. Argue for a standing performance budget that launches draw down, a named owner for breaches, and the discipline that stops many individually acceptable regressions compounding into a slow product.

### Step 1 — is the 20 ms real? Before trading anything away, verify the measurement. - **Look at the interval, not the point estimate.** A reported 20 ms with an interval of [4 ms, 36 ms] is a very different fact from [18 ms, 22 ms]. Latency distributions are skewed, and percentile estimates are noisier than means. - **Check instrumentation symmetry.** If the treatment introduced a new code path, the timer may start or stop at a different point, producing a difference that exists only in the measurement. - **Check for transient causes.** Freshly deployed code can carry cold caches, unwarmed connection pools, or first-run compilation costs that fade. Compare the daily trend across the test window rather than the pooled number. - **Check the invariants.** Skewed traffic split or different bot filtering between arms can move a percentile without any real change in user experience. A guardrail alarm that turns out to be a measurement artefact is common enough that skipping this step is itself a red flag. ### Step 2 — compare against the pre-agreed threshold The decision rule should already exist: a latency budget written down when the experiment was designed, for instance "p95 must not degrade by more than 10 ms". Then the answer is mechanical. - Inside the budget: the guardrail passed. Ship, and record the consumed budget. - Outside the budget: the guardrail is breached. It is not the experiment owner's call to override, which is the entire point of writing the number down beforehand. If no threshold was set, say so honestly — the correct lesson is that the experiment lacked a decision rule — and set one now on the merits, without looking at which side of it the observed value falls. ### Step 3 — decompose the number A pooled p95 is an average over very unequal populations. Two decompositions matter most: - **By device and network.** A 20 ms pooled degradation frequently means a handful of milliseconds on fast desktops and a much larger hit on low-end mobile devices or poor connections. If the harm concentrates on the users least able to absorb it, the pooled number understates the real cost. - **By page or flow.** Latency added to a rarely visited settings page is not the same cost as latency added to the checkout path, even at identical millisecond counts. Also ask about the *shape*: p95 moving while the median is flat means a tail got worse — some requests became much slower. That is a different failure than the whole distribution shifting, and often points at a specific slow path worth fixing. ### Step 4 — quantify the exchange, using your own evidence To weigh a sign-up gain against a latency cost you need a conversion rate between them, and the only trustworthy source is your own product's past latency experiments — deliberate slowdown tests or infrastructure speed-ups measured on the same population. If such evidence exists, the tradeoff becomes arithmetic: the estimated engagement or revenue cost of 20 ms versus the measured sign-up gain. If it does not, be explicit that you are making a judgement call, and avoid importing rules of thumb from other products, where the traffic mix, device profile and baseline latency are all different. ### Step 5 — ask whether the cost is necessary at all Often the most useful move is to reject the framing. Added latency is frequently incidental: a synchronous call that could be deferred, an extra render-blocking asset, a payload that could be paginated. Sending it back with "the feature is good, the implementation costs 20 ms, remove the 20 ms" gets the win without the trade. Alternatives worth naming: ship to a subset while the fix lands, or ship with a monitoring alert on the guardrail so regression is caught in production. ### Step 6 — decide and record The defensible outcomes: 1. **Ship** — degradation inside budget, verified real, no segment concentration. 2. **Ship with a fix** — the gain is real and the latency cost is removable; land the optimisation first or immediately after. 3. **Ship as an explicit tradeoff** — breach accepted by the named owner, written down with the amount of budget consumed, so the next team knows how much room is left. 4. **Do not ship** — the breach is large, concentrated on vulnerable segments, or on a critical path. Whichever it is, record the consumed budget. The systemic risk with latency is a ratchet: every experiment degrades performance slightly, each degradation is individually defensible, and the product gets steadily slower with no single decision to blame. Tracking cumulative consumption against a standing budget is what stops that drift, and it is the part of the answer that marks out someone who has lived through it.

  • The 20ms breach is small and the sign-up win is large. Why not just move the latency threshold?
    Because a threshold that moves once results are known is not a guardrail — it is a formality that always yields to whichever metric is winning. Moving it also destroys the record: nobody can tell later which launches consumed the performance budget. If the threshold is genuinely wrong, change it as a standing policy decision, applied prospectively to future experiments, not retroactively to this one.
  • Pooled p95 latency is up 20ms. What decomposition would change your decision?
    By device class and network quality first: if fast desktops absorb 5ms while low-end mobile takes 80ms, the pooled figure badly understates the harm on the users least able to tolerate it. Then by page or flow, since latency on a critical path costs more than the same milliseconds on a rarely visited page. Concentration in either dimension pushes toward blocking.
  • How do you stop many individually acceptable latency degradations from making the product slow?
    Treat performance as a budget that launches consume rather than a per-test pass/fail. Record how much each shipped experiment used, hold a standing total, and require replenishment work when the budget runs low. Without cumulative accounting every degradation is defensible in isolation and nobody is accountable for the aggregate slowdown.
  • What would make you suspect the 20ms is a measurement artefact rather than a real regression?
    A new code path where the timer boundaries differ between arms, a degradation concentrated in the first days after deploy that fades as caches warm, an unusually wide interval on a skewed percentile estimate, or different bot and outlier filtering between arms. Any of these means fix the measurement before treating the number as a real cost.

saying these in an interview costs you the question

  • Ships because the primary win is bigger than the guardrail loss
  • Raises the latency threshold after seeing the result
  • Never checks the interval around the 20ms estimate
  • Reads pooled p95 without segmenting by device
  • Treats each degradation in isolation, ignoring cumulative drift

context