skip to content

Why isn't a non-significant guardrail p-value enough to conclude an A/B test did no harm?

level: middleimportance: should knowfreq 46%

answer

  1. absence of evidence is not evidence of absence
  2. guardrails ride on the primary metric's power
  3. flip the burden of proof
  4. state a tolerated degradation margin
  5. interval must sit inside the margin

basics

~20 s

A non-significant result means the data does not rule out zero difference; it never proves the difference is zero. Claiming no harm needs a stated degradation margin and a guardrail interval sitting entirely inside it.

solid answer

~50 s

A p-value answers one question: how surprising is this data if the true difference were exactly zero. A large p-value means the data is compatible with zero — but it is equally compatible with a degradation big enough to matter, especially on guardrails, which are often noisier and less powered than the primary metric. So "not significant" is absence of evidence, not evidence of absence. The correct frame is non-inferiority: pick a margin before the test, say a 1% relative degradation you are willing to tolerate, then require the confidence interval on the guardrail difference to lie entirely inside that margin. For a lower-is-better guardrail such as latency, that means the interval's upper bound must sit below +1%. If the interval is wider than the margin, the experiment simply has not answered the question, and the honest report is "undetermined", not "safe".

go deeper

for a junior

Be ready to state that a large p-value means the data is compatible with no difference, not that no difference exists, and that a wide interval can hide real harm.

for a middle

Explain the non-inferiority setup: a margin fixed in advance, the null placed in the harmful direction, and the interval rule — lower bound above the negative margin for higher-is-better metrics, upper bound below it otherwise.

for a senior

Show you check guardrail power at design time, report intervals rather than pass/fail flags, and are willing to return 'undetermined' instead of letting an under-powered check wave a launch through.

for a principal

Own the policy angle: who sets margins, why they are fixed before results are seen, and how a default-pass guardrail regime lets many individually non-significant degradations accumulate across launches.

### What a p-value does and does not say A p-value is the probability of observing data at least as extreme as what you saw, *assuming the null hypothesis is true*. For a standard guardrail check the null is "the true difference between treatment and control is zero". A large p-value therefore says only that the observed data would not be surprising under a zero difference. It does not say the difference is zero, and it gives no probability that the null is true. The asymmetry is structural. Rejecting the null is a positive claim backed by a controlled error rate. Failing to reject is the *default* outcome — it is what you get from a true null, and also what you get from a small sample, a noisy metric, or a short run. Reporting "p = 0.4, so the guardrail is fine" conflates those situations. ### Why guardrails are the worst place to make this mistake Experiments are powered for the primary metric. Guardrails then ride along on whatever sample that provides, and they are frequently harder to detect on: - **Rare events.** Crashes, refunds and cancellations have low base rates and high relative variance, so their intervals are wide. - **Heavy or skewed distributions.** Latency percentiles and revenue-adjacent guardrails carry more variance per user than a conversion rate. - **Sub-population effects.** A guardrail may degrade badly for low-end devices while the pooled average barely moves. The result is a systematic bias in the decision process: the primary metric is measured sensitively enough to cross its bar, while the guardrail is measured too coarsely to raise an alarm. Ship on that pairing repeatedly and you accumulate real harm, each instance individually "not significant". ### The non-inferiority frame Non-inferiority testing inverts the burden of proof. Instead of asking "can we detect harm?", it asks "can we rule out harm larger than a stated amount?". The machinery: 1. **Choose a margin before the test.** It is a substantive judgement, not a statistical one: the largest degradation the organisation is willing to accept in exchange for a win elsewhere. For example, at most a 1% relative degradation, or at most 10 ms added to p95 latency. 2. **State the null in the harmful direction.** For a higher-is-better guardrail the null is "the true effect is a degradation of at least the margin", i.e. `T - C <= -margin`. For a lower-is-better guardrail such as latency it is `T - C >= +margin`. 3. **Reject that null to claim non-inferiority.** Operationally this is a one-sided test, and the equivalent interval rule is easier to communicate: build the confidence interval for `T - C` and check that it lies entirely on the acceptable side of the margin. Higher-is-better: the *lower* bound must be above `-margin`. Lower-is-better: the *upper* bound must be below `+margin`. Note what this does to the direction of caution. Under the usual zero-null test, a low-powered experiment tends to pass the guardrail by default. Under a non-inferiority test, a low-powered experiment *fails* by default, because a wide interval spills past the margin. That is the correct default for a safety check: uncertainty counts against shipping, not for it. ### Reading the four cases Take a lower-is-better guardrail with a margin of +1% and a confidence interval on the relative difference: - Interval `[-0.4%, +0.3%]` — entirely inside the margin. Non-inferiority established; ship on this guardrail. - Interval `[+0.2%, +0.6%]` — a real degradation, but the whole interval is inside the tolerated margin. Statistically significant harm that is nonetheless acceptable by the pre-agreed rule. This case is where the margin earns its keep, and where teams without one argue. - Interval `[-2.0%, +2.5%]` — not significant, and useless. The data cannot distinguish "harmless" from "twice the tolerated harm". Run longer, reduce variance, or report undetermined. - Interval `[+1.4%, +2.2%]` — clearly worse than the margin. The guardrail is breached. The second and third cases are the ones that separate a good answer from a rote one: significance and acceptability are independent axes, and a non-significant result can be the least informative outcome on the board. ### Practical consequences Set margins when the experiment is designed, alongside the primary metric's minimum detectable effect, and check up front whether the planned sample can produce an interval narrower than the margin — if not, the guardrail is decorative before the test even runs. Prefer intervals over bare p-values in the scorecard so the magnitude question is unavoidable. And when the interval is too wide, say so plainly: "undetermined" is a legitimate experimental result and a far better basis for a launch conversation than a p-value that quietly encodes nothing more than a small sample.

  • How would you pick the non-inferiority margin for a latency guardrail?
    From the product cost of the degradation, not from statistical convenience. Use what past latency experiments in your own product showed about how much engagement a given slowdown costs, the perceptual thresholds where users notice, and the standing latency budget the change has to fit inside. Then sanity-check feasibility: if the achievable interval is wider than the margin, the check cannot run as designed.
  • The guardrail interval is significantly worse than zero but entirely inside the tolerated margin. Do you ship?
    Yes, if the margin was agreed before the test — that is exactly the case it was written for. The degradation is real and known, and the pre-agreed rule says it is an acceptable price. The failure mode to avoid is renegotiating the margin after seeing the number, in either direction: raising it to permit a win, or discovering new objections to block one.
  • Under a non-inferiority frame, what happens when a guardrail is badly under-powered?
    The test fails to establish non-inferiority, because the interval is wider than the margin and cannot be contained by it. That is the intended behaviour: uncertainty counts against shipping. Under a plain zero-null test the same weak data would pass silently, which is precisely the failure mode non-inferiority is designed to remove.

A smoke alarm with dead batteries never goes off. Silence from an under-powered guardrail is the same reassurance: it means nothing was detected, not that nothing happened.

saying these in an interview costs you the question

  • Says a large p-value proves the metrics are unchanged
  • Treats failing to reject the null as accepting it
  • Reports guardrails as pass/fail with no interval
  • Picks the tolerated margin after seeing the result
  • Ignores that guardrails inherit the primary metric's sample size

context