skip to content

An A/B test shows +0.05% lift, CI [+0.01%, +0.09%], under a 0.5% cost bar — ship it?

level: seniorimportance: should knowfreq 46%

answer

  1. excludes zero, still below the bar
  2. zero is not the decision boundary
  3. cost sets the break-even lift
  4. precision makes the no conclusive
  5. override only for reasons off this metric

basics

~20 s

No. The interval excludes zero, so the lift is real, but its entire range sits below the 0.5% the feature must return to cover its cost. That is a conclusive don't-ship, not a marginal call.

solid answer

~50 s

This is one of the cleanest results an experiment can produce, and the answer is still no. The interval `[+0.01%, +0.09%]` excludes zero, so a genuine positive effect is established; it also lies entirely below the 0.5% break-even lift the feature needs to pay for its build and ongoing cost. Precision here works in the decision's favour: rather than leaving the outcome open, the test has ruled out the possibility that the effect is large enough to matter. So report it as a confirmed but sub-threshold win and decline to ship on the metric alone. The decision only reopens for reasons outside this metric — the feature also removes operating cost, unblocks a platform migration, or serves a strategic bet whose payoff is deliberately not in this readout — and those reasons should be argued explicitly, not smuggled in as "but it was significant".

go deeper

for a junior

Be ready to say what excluding zero does and does not establish: the effect is real, its size is between +0.01% and +0.09%, and nothing about whether that size is worth having.

for a middle

Explain why very large samples make tiny effects significant, and why the decision rule therefore compares the interval against a cost threshold rather than against zero.

for a senior

Demonstrate that you treat this as a conclusive no rather than a marginal call, and that you can name the specific overrides — reduced operating cost, a wrong primary metric, a no-harm check for a migration — and argue them on their own terms.

for a principal

Own the break-even calculation as a standing input to experiment planning, so bars are set before launch and the organisation stops accumulating confirmed, sub-threshold features that each add permanent cost.

## The readout - Lift: **+0.05%** relative, 95% CI **[+0.01%, +0.09%]**. - Break-even bar: **0.5%** — the lift at which the feature's incremental revenue covers what it costs to build, run and maintain. The interval excludes zero. In the vocabulary of a readout meeting, the test "won". It is also, on the number that governs the decision, a decisive loss: every value the data are compatible with is roughly an order of magnitude below the bar. ## Two thresholds, only one of which is zero An interval is read against whatever boundary changes the decision. Zero is the boundary for "does this do anything at all". It is almost never the boundary for "should we ship this". The second boundary is set by cost, and it is a real number someone has to produce: - One-off build cost amortised over the feature's expected life. - Ongoing cost: infrastructure, a paid dependency, latency budget, the support and on-call surface it adds. - Opportunity cost: the surface area it occupies and the next change it makes harder. Divide the annualised cost by the revenue the metric represents and you get a break-even lift. Here that is 0.5%. The comparison then has three outcomes, and all three are useful: 1. **Interval entirely above the bar** — ship, and the size of the win is bounded from below. 2. **Interval entirely below the bar** — the case here. Do not ship, and the conclusion is firm: the data have excluded a lift worth paying for. 3. **Interval straddling the bar** — the awkward one. The effect may or may not pay for itself, and the experiment has not resolved the question that matters even though it cleared significance. Case 2 is frequently misreported as a win because significance testing trained everyone to look at zero. A candidate who spots that the *precision* is what makes this a confident no — not a hesitant maybe — is showing the judgment the question is testing. ## Why tiny effects reach significance On a high-traffic surface the standard error of the lift shrinks roughly with the square root of the sample size, so a large enough experiment will resolve differences far smaller than anything commercially interesting. Significance is a statement about precision relative to zero, and precision is something you buy with traffic. That is exactly why the threshold in the decision rule has to be the cost bar rather than zero: otherwise a mature product ends up shipping a long tail of confirmed, worthless changes, each of which adds permanent cost. ## When the decision legitimately reopens Declining on the metric is not the same as declining forever. Reasons to override, each of which must be stated in its own terms: - **The feature reduces cost.** If it removes a paid dependency or cuts serving cost, the break-even lift is lower — possibly negative, in which case a small confirmed gain is pure upside. Recompute the bar rather than arguing about the lift. - **The primary metric is not the whole payoff.** A change that barely moves conversion but measurably reduces support contacts is being judged on the wrong number. This is a reason to fix the readout metric, not to ignore the interval. - **It is a stepping stone.** A migration or platform change whose value is in what it unblocks should have been framed that way before launch, with the experiment run to establish *no harm* rather than to establish a win. That framing is the honest one, and the interval here does support it: a bounded, tiny positive effect means the change is safe. - **Cumulative small wins are the strategy.** Some organisations deliberately accumulate sub-threshold gains where marginal cost is genuinely near zero. Legitimate — provided the near-zero marginal cost is real and someone tracks the aggregate, rather than each team asserting it separately. ## The write-up Report it as: confirmed positive effect, magnitude bounded between +0.01% and +0.09%, break-even bar 0.5%, recommendation not to ship on this metric. The precision is the headline, because it turns an open question into a closed one. A readout that says only "significant, p < 0.05" invites six weeks of argument that the interval settles in a sentence. ## The mirror-image mistake The complement is just as common: killing a feature whose interval sits entirely *above* the bar because the lift "looks small". If the bar is 0.5% and the interval is `[+0.6%, +1.1%]`, the change pays for itself across its whole plausible range, and the smallness of the number is irrelevant. Both errors come from reading the estimate against intuition instead of reading the interval against the threshold.

  • What if the interval were [+0.2%, +0.9%] against the same 0.5% bar?
    Then the experiment is inconclusive on the decision that matters: the lift may or may not cover its cost. Significance is irrelevant here. Either buy more precision to resolve which side of 0.5% the effect falls on, or tighten the cost estimate, since the bar is an estimate too and may be the cheaper thing to sharpen.
  • Who should set the break-even lift, and when?
    Whoever owns the cost line, together with the experiment owner, and before launch. Setting it afterwards invites a bar chosen to match the result. Writing it into the test plan alongside the MDE also surfaces the cases where the required lift is implausibly large, which is a reason not to build the feature at all.
  • Does a confirmed but tiny lift justify shipping if the feature is already built?
    Only if the remaining cost is genuinely near zero. Build cost is sunk, but serving, maintenance and complexity costs continue, so recompute the bar on ongoing cost alone. If the interval still sits below that reduced bar, the sunk build is not an argument for shipping.

A scale that confirms a parcel weighs between 1 and 9 grams has not failed you because the postage threshold is 500 grams — it has settled the question.

saying these in an interview costs you the question

  • Shipping because the interval excludes zero
  • Treating a small p-value as a business case
  • Setting the break-even bar after seeing the result
  • Calling a precise sub-threshold result a weak test
  • Counting sunk build cost as a reason to launch

context