skip to content

Why report a confidence interval on A/B lift instead of just a p-value?

level: middleimportance: must knowfreq 72%

answer

  1. one number versus a range
  2. magnitude and units survive
  3. significance says nothing about size
  4. compare against the planned MDE
  5. wide interval exposes a weak test

basics

~20 s

A p-value only says how compatible the data are with zero effect. A confidence interval reports the lift itself with its precision, so you can compare the plausible range against the minimum detectable effect and the launch bar.

solid answer

~50 s

A p-value compresses an experiment into one number about a single hypothesis: how surprising the observed difference would be if the true lift were exactly zero. It carries no magnitude and no units, so `p = 0.03` fits a lift of +0.1% or +12%, and `p = 0.4` fits both a genuinely flat result and a hopelessly imprecise one. A confidence interval on the lift reports the estimate together with the range of effects the data are compatible with at the stated confidence level, in the metric's own units. That range is what a launch decision actually consumes: you check it against the minimum detectable effect the test was planned around, and against the lift the feature has to clear to pay for its build and running cost. Reporting the interval also makes a low-powered test visibly low-powered instead of letting it masquerade as a null result.

go deeper

for a junior

Be ready to say what each output is: the p-value speaks only to the no-difference hypothesis, the interval reports the lift's plausible range in the metric's own units. Quote both when you present a result.

for a middle

Explain why two readouts with the same p-value can license opposite decisions, and how the interval's width, not the p-value, exposes an under-powered test. Show the estimate-interval-threshold reporting format.

for a senior

Demonstrate that you decide against a threshold: hold the interval up to the planned MDE or the feature's break-even lift and say which of ship, drop, or extend that comparison implies. Insist on the interval in review even when it is unflattering.

for a principal

Own the reporting standard. Decide what every experiment readout must carry, on which metric and which scale, and defend why the organisation reads intervals rather than negotiating significance thresholds feature by feature.

## The two numbers, precisely An A/B readout starts from an **estimate**: the observed difference between treatment and control on the decision metric, usually expressed as a lift, either in absolute units (+0.6 percentage points of conversion) or relative to control (+2.1%). That number alone is not a result, because a different random split of the same users would have produced a different number. Everything that follows is about attaching uncertainty to it. A **p-value** is the probability, computed under the assumption that the true difference is exactly zero, of observing a test statistic at least as extreme as the one you got. It answers one question — how compatible are these data with the no-difference hypothesis — and answers it on a scale that has no connection to the size of the business effect. A **confidence interval** on the lift is a range of effect sizes, computed from the same data, that are compatible with what you observed at the stated confidence level. It is reported in the units of the lift itself: `[-2.0%, +2.8%]`, `[+0.01%, +0.09%]`, `[+3.1%, +9.4%]`. ## What the p-value cannot tell you Three readouts illustrate the gap. 1. A huge test on a high-traffic surface returns `p = 0.001` with a lift of +0.04%. Highly significant, commercially irrelevant if the feature costs anything to run. 2. A small test returns `p = 0.30` with a lift of +1.5% and an interval running from -4% to +7%. Not significant, and also not informative: effects far larger than anything the roadmap assumed remain fully on the table. 3. A well-powered test returns `p = 0.30` with a lift of +0.1% and an interval from -0.3% to +0.5%. Also not significant — but this one has genuinely bounded the effect, and licenses a very different conclusion from case 2. Cases 2 and 3 carry nearly the same p-value and opposite decision content. Nothing in the p-value distinguishes them. The interval does it immediately: one is wide, one is tight. ## What the interval feeds into An experiment is planned around a **minimum detectable effect (MDE)** — the smallest lift the test was sized to detect reliably, chosen because it is the smallest lift anyone would act on. The MDE is a planning input, and at readout the interval is what you hold it up against: - If the whole interval sits **above** the MDE, the test is a conclusive win at a size worth shipping. - If the whole interval sits **below** the MDE (and above the harm bound you care about), the test conclusively rules out an effect big enough to matter — a decisive *don't ship*, not a failure. - If the interval **spans** the MDE, the experiment has not separated the outcomes that would drive different decisions, whatever the p-value says. The same comparison works against a cost threshold rather than the MDE: a feature with an ongoing infrastructure or maintenance cost has a break-even lift, and the interval either clears it, sits under it, or straddles it. ## Reporting practice Write the readout as estimate, interval, and threshold together — "+2.4% (95% CI +0.8% to +4.0%) against an MDE of 1%" — rather than "significant, p = 0.01". Three habits follow from that: - **State the metric and the scale.** Relative and absolute lift differ by a factor of the baseline rate, and a stakeholder reading "+2%" will assume whichever is more flattering. - **Keep the interval on the quantity being decided.** If the decision is about revenue per user, report the interval on revenue per user, not on a proxy that happens to be less noisy. - **Never drop the interval when it is embarrassing.** A very wide interval is information: it says the test could not answer the question, which is a finding about the experiment design, not a result to bury behind "no significant difference". ## The residual role of the p-value None of this makes p-values useless. They are a compact summary for a pre-registered decision rule, and reviewers are used to them. The point is that the p-value is the *smaller* of the two outputs, and shipping decisions are made in the units of the interval. A readout that shows only the p-value has thrown away the part a launch committee needs.

  • If the interval already tells you more, why do teams still pre-register a significance threshold?
    Because a pre-committed rule stops the readout from being renegotiated after the fact. Fixing the confidence level and the decision threshold before launch means the interval is compared against a bar nobody chose once the numbers were visible. The interval supplies the magnitude; the pre-registered rule supplies the discipline.
  • Should the interval be reported on relative or absolute lift?
    Report whichever the decision is denominated in, and label it. Absolute lift in percentage points is what capacity and revenue models consume; relative lift is easier to compare across surfaces with different baselines. Reporting one while stakeholders assume the other is a common source of inflated expectations.
  • What does a narrower confidence level, say 80% instead of 95%, change about the readout?
    It narrows the interval, so more results appear to exclude zero and more marginal features clear the bar. That is a deliberate trade of a higher wrong-ship rate for faster decisions, and it should be set as policy before the test, not chosen at readout because 95% was inconvenient.

A p-value is a smoke alarm: it beeps or it does not. An interval is a thermometer reading with a tolerance — it tells you how hot, and how sure.

saying these in an interview costs you the question

  • Treating a small p-value as evidence of a large effect
  • Reporting significance without the lift's magnitude
  • Assuming a non-significant result means the effect is zero
  • Quoting a relative lift as if it were percentage points
  • Switching the confidence level after seeing the numbers

context