skip to content

Readout & Decision Rules

Reading a finished experiment: the effect with an interval, practical versus statistical significance, what a flat result licenses, and how to look at segments without inventing findings.

on this pageshow

explore

questions

15

Why report a confidence interval on A/B lift instead of just a p-value?

level: middleimportance: must knowfreq 72%

answer

  1. one number versus a range
  2. magnitude and units survive
  3. significance says nothing about size
  4. compare against the planned MDE
  5. wide interval exposes a weak test

basics

~20 s

A p-value only says how compatible the data are with zero effect. A confidence interval reports the lift itself with its precision, so you can compare the plausible range against the minimum detectable effect and the launch bar.

solid answer

~50 s

A p-value compresses an experiment into one number about a single hypothesis: how surprising the observed difference would be if the true lift were exactly zero. It carries no magnitude and no units, so `p = 0.03` fits a lift of +0.1% or +12%, and `p = 0.4` fits both a genuinely flat result and a hopelessly imprecise one. A confidence interval on the lift reports the estimate together with the range of effects the data are compatible with at the stated confidence level, in the metric's own units. That range is what a launch decision actually consumes: you check it against the minimum detectable effect the test was planned around, and against the lift the feature has to clear to pay for its build and running cost. Reporting the interval also makes a low-powered test visibly low-powered instead of letting it masquerade as a null result.

go deeper

for a junior

Be ready to say what each output is: the p-value speaks only to the no-difference hypothesis, the interval reports the lift's plausible range in the metric's own units. Quote both when you present a result.

for a middle

Explain why two readouts with the same p-value can license opposite decisions, and how the interval's width, not the p-value, exposes an under-powered test. Show the estimate-interval-threshold reporting format.

for a senior

Demonstrate that you decide against a threshold: hold the interval up to the planned MDE or the feature's break-even lift and say which of ship, drop, or extend that comparison implies. Insist on the interval in review even when it is unflattering.

for a principal

Own the reporting standard. Decide what every experiment readout must carry, on which metric and which scale, and defend why the organisation reads intervals rather than negotiating significance thresholds feature by feature.

## The two numbers, precisely An A/B readout starts from an **estimate**: the observed difference between treatment and control on the decision metric, usually expressed as a lift, either in absolute units (+0.6 percentage points of conversion) or relative to control (+2.1%). That number alone is not a result, because a different random split of the same users would have produced a different number. Everything that follows is about attaching uncertainty to it. A **p-value** is the probability, computed under the assumption that the true difference is exactly zero, of observing a test statistic at least as extreme as the one you got. It answers one question — how compatible are these data with the no-difference hypothesis — and answers it on a scale that has no connection to the size of the business effect. A **confidence interval** on the lift is a range of effect sizes, computed from the same data, that are compatible with what you observed at the stated confidence level. It is reported in the units of the lift itself: `[-2.0%, +2.8%]`, `[+0.01%, +0.09%]`, `[+3.1%, +9.4%]`. ## What the p-value cannot tell you Three readouts illustrate the gap. 1. A huge test on a high-traffic surface returns `p = 0.001` with a lift of +0.04%. Highly significant, commercially irrelevant if the feature costs anything to run. 2. A small test returns `p = 0.30` with a lift of +1.5% and an interval running from -4% to +7%. Not significant, and also not informative: effects far larger than anything the roadmap assumed remain fully on the table. 3. A well-powered test returns `p = 0.30` with a lift of +0.1% and an interval from -0.3% to +0.5%. Also not significant — but this one has genuinely bounded the effect, and licenses a very different conclusion from case 2. Cases 2 and 3 carry nearly the same p-value and opposite decision content. Nothing in the p-value distinguishes them. The interval does it immediately: one is wide, one is tight. ## What the interval feeds into An experiment is planned around a **minimum detectable effect (MDE)** — the smallest lift the test was sized to detect reliably, chosen because it is the smallest lift anyone would act on. The MDE is a planning input, and at readout the interval is what you hold it up against: - If the whole interval sits **above** the MDE, the test is a conclusive win at a size worth shipping. - If the whole interval sits **below** the MDE (and above the harm bound you care about), the test conclusively rules out an effect big enough to matter — a decisive *don't ship*, not a failure. - If the interval **spans** the MDE, the experiment has not separated the outcomes that would drive different decisions, whatever the p-value says. The same comparison works against a cost threshold rather than the MDE: a feature with an ongoing infrastructure or maintenance cost has a break-even lift, and the interval either clears it, sits under it, or straddles it. ## Reporting practice Write the readout as estimate, interval, and threshold together — "+2.4% (95% CI +0.8% to +4.0%) against an MDE of 1%" — rather than "significant, p = 0.01". Three habits follow from that: - **State the metric and the scale.** Relative and absolute lift differ by a factor of the baseline rate, and a stakeholder reading "+2%" will assume whichever is more flattering. - **Keep the interval on the quantity being decided.** If the decision is about revenue per user, report the interval on revenue per user, not on a proxy that happens to be less noisy. - **Never drop the interval when it is embarrassing.** A very wide interval is information: it says the test could not answer the question, which is a finding about the experiment design, not a result to bury behind "no significant difference". ## The residual role of the p-value None of this makes p-values useless. They are a compact summary for a pre-registered decision rule, and reviewers are used to them. The point is that the p-value is the *smaller* of the two outputs, and shipping decisions are made in the units of the interval. A readout that shows only the p-value has thrown away the part a launch committee needs.

  • If the interval already tells you more, why do teams still pre-register a significance threshold?
    Because a pre-committed rule stops the readout from being renegotiated after the fact. Fixing the confidence level and the decision threshold before launch means the interval is compared against a bar nobody chose once the numbers were visible. The interval supplies the magnitude; the pre-registered rule supplies the discipline.
  • Should the interval be reported on relative or absolute lift?
    Report whichever the decision is denominated in, and label it. Absolute lift in percentage points is what capacity and revenue models consume; relative lift is easier to compare across surfaces with different baselines. Reporting one while stakeholders assume the other is a common source of inflated expectations.
  • What does a narrower confidence level, say 80% instead of 95%, change about the readout?
    It narrows the interval, so more results appear to exclude zero and more marginal features clear the bar. That is a deliberate trade of a higher wrong-ship rate for faster decisions, and it should be set as policy before the test, not chosen at readout because 95% was inconvenient.

A p-value is a smoke alarm: it beeps or it does not. An interval is a thermometer reading with a tolerance — it tells you how hot, and how sure.

saying these in an interview costs you the question

  • Treating a small p-value as evidence of a large effect
  • Reporting significance without the lift's magnitude
  • Assuming a non-significant result means the effect is zero
  • Quoting a relative lift as if it were percentage points
  • Switching the confidence level after seeing the numbers

context

open as a page

Why does slicing a flat A/B test result into many segments so often produce a false winner?

level: middleimportance: must knowfreq 72%

basics

~20 s

Every segment cut is another hypothesis test. At a 5% false-positive rate, about twenty cuts of a genuinely flat experiment throw up roughly one significant segment by chance alone, so the winner you find is usually noise.

open as a page

What is dilution in an A/B test where only 3% of assigned users ever see the change?

level: middleimportance: must knowfreq 72%

basics

~20 s

Dilution is the shrinking of a measured effect when most assigned users never encounter the change. At a 3% trigger rate an effect among exposed users appears about thirty times smaller in the all-up read, while the noise stays.

open as a page

In an A/B readout, how does an intent-to-treat analysis differ from a triggered-only analysis?

level: middleimportance: must knowfreq 60%

basics

~20 s

Intent-to-treat counts every assigned user and answers what shipping does to the whole population. A triggered analysis keeps only users who met the trigger condition in both arms, answering what the change does to those who encounter it.

open as a page

An A/B test reads +0.4% lift, 95% CI [-2.0%, +2.8%], planned MDE 2% — is that flat?

level: seniorimportance: must knowfreq 58%

basics

~20 s

No — it is inconclusive, not flat. The interval reaches +2.8%, above the 2% minimum detectable effect the test was planned around, so a lift worth shipping has not been ruled out, and neither has a 2% loss.

open as a page

How should you report an A/B test with p = 0.3 without claiming 'no difference'?

level: juniorimportance: should knowfreq 55%

basics

~10 s

Report the estimated lift with its confidence interval and say what the interval rules out. A p-value of 0.3 means the data are compatible with no difference, not that the difference is zero.

open as a page

In an A/B test, what is the difference between a pre-registered segment cut and a post-hoc one?

level: juniorimportance: should knowfreq 58%

basics

~20 s

A pre-registered cut is fixed in the design doc before launch, so its error rate holds. A post-hoc cut is chosen after seeing results, so its p-value was selected and cannot be read at face value.

open as a page

In an A/B test, what does it mean for a user to trigger into the experiment?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Triggering means the user actually reached the condition where the tested change can act on them, such as opening the settings page a new feature lives on. Assigned users who never reach it are never exposed.

open as a page

An A/B test shows +0.05% lift, CI [+0.01%, +0.09%], under a 0.5% cost bar — ship it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

No. The interval excludes zero, so the lift is real, but its entire range sits below the 0.5% the feature must return to cover its cost. That is a conclusive don't-ship, not a marginal call.

open as a page

How do you tell a real subgroup effect in an A/B test from post-hoc story fitting?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Demand evidence the story cannot supply: a mechanism stated before the data, a formal test that the effect differs between segments rather than two separate within-segment tests, consistency across neighbouring cells, and replication in a confirmatory experiment.

open as a page

An A/B test shows a +40% conversion lift in one locale and flat elsewhere — what do you do?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Treat it as a suspected defect before treating it as a finding. Twyman's law says any figure that looks unusually interesting is usually wrong, so check instrumentation, currency and units, bot traffic and the assignment split in that locale first.

open as a page

How do you log the trigger condition in the control arm when the feature exists only in treatment?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Evaluate the same trigger condition inside the control code path and record that the user would have qualified, while changing nothing they see. Without that counterfactual record there is no comparable control group for a triggered readout.

open as a page

How should a launch committee set interval-based ship rules for low-traffic A/B tests?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Decide up front what a wide interval licenses. On low-traffic surfaces, pre-commit to a rule such as ship if the interval excludes meaningful harm and the change is cheap, and never log an uninformative test as evidence of no effect.

open as a page

What policy should govern segment reporting in A/B test readouts across an experimentation team?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Make the all-up metric the decision, allow a short pre-declared segment list per experiment, keep exploratory cuts in a labelled section with the cut count disclosed, and require a powered confirmatory run before any segment result changes a launch.

open as a page

Your triggered analysis shows +9% lift on the 3% of users who see the feature — what goes in the launch memo?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

Both numbers plus the bridge between them: roughly +9% for the users who meet the trigger, roughly +0.3% across the whole population, and the 3% trigger rate that connects the two. Quoting either figure alone misleads a reader.

open as a page