skip to content

An A/B test reads +0.4% lift, 95% CI [-2.0%, +2.8%], planned MDE 2% — is that flat?

level: seniorimportance: must knowfreq 58%

answer

  1. check the upper end, not just zero
  2. compare endpoints to the MDE
  3. contains zero and contains the MDE
  4. precision came in worse than planned
  5. uninformative is a finding, not a null

basics

~20 s

No — it is inconclusive, not flat. The interval reaches +2.8%, above the 2% minimum detectable effect the test was planned around, so a lift worth shipping has not been ruled out, and neither has a 2% loss.

solid answer

~50 s

The interval contains zero, so there is no significant win, but it also contains +2.8%, comfortably past the 2% MDE the experiment was sized for — and it reaches -2.0% on the other side. The data are compatible with a launch-worthy gain, with nothing, and with a real harm. Calling that flat asserts something the experiment never established. The right label is inconclusive, and the first diagnostic step is why the observed precision came in worse than planned: did the test stop early, was the metric more variable than the power calculation assumed, or did much of the traffic land outside the intended audience? If the surface can supply more traffic without a seasonal confound, extend to the precision the plan required. If it cannot, record the result as uninformative and make the call on cost and strategy rather than dressing it up as evidence of no effect.

go deeper

for a junior

Be ready to notice that an interval containing zero is not significant, and to read both endpoints out loud. The upper end of +2.8% is the detail most candidates skip.

for a middle

Explain the difference between an interval that rules out anything worth acting on and one that merely fails to exclude zero, and relate the interval's width back to the MDE the test was powered for.

for a senior

Show the diagnosis: name the plausible reasons precision fell short of plan — short sample, underestimated variance, diluted audience, unit mismatch — and pick a defensible next step rather than shipping or killing on a coin flip.

for a principal

Own how the organisation records uninformative experiments so they are not later cited as evidence of no effect, and decide when a low-traffic surface should not be A/B tested at all.

## The readout - Point estimate: **+0.4%** relative lift on the primary metric. - 95% confidence interval: **[-2.0%, +2.8%]**. - Planned minimum detectable effect: **2%** — the smallest lift the team decided in advance was worth acting on, and the size the test was powered to detect. Zero sits inside the interval, so the result is not significant at the 5% level. The tempting shorthand is "the feature did nothing". That shorthand is wrong here, and the reason is visible without any further computation. ## Flat versus inconclusive These are two different findings that share a p-value. **Conclusively flat** means the experiment measured the lift precisely enough to rule out anything you would have acted on. Concretely: the entire interval lies inside the band of effects you consider not worth a decision — upper bound below the MDE, lower bound above whatever regression you would treat as harm. An interval of `[-0.3%, +0.5%]` against a 2% MDE is flat. It says: whatever this feature does, it is smaller than the smallest thing we care about. **Inconclusive** means the interval spans the decision boundary. Here the upper end, +2.8%, is 40% larger than the MDE. If the true lift were exactly the 2% the roadmap was built on, this readout would be an entirely unremarkable outcome. The lower end, -2.0%, is a mirror-image problem: a loss of the same magnitude as the win being chased is equally unexcluded. The experiment has not distinguished the three worlds — good, nothing, bad — that would each lead somewhere different. The general rule, worth being able to state cleanly: **compare the interval's endpoints to the thresholds that change the decision, not to zero.** Zero is rarely the boundary anyone acts on. ## Why the precision missed the plan A test powered for a 2% MDE should return an interval roughly of that half-width. This one has a half-width of about 2.4 percentage points of relative lift — wider than planned. That gap is itself a finding, and it has a short list of causes worth checking before any decision: - **Short sample.** The test was stopped early, or the traffic forecast that fed the power calculation was optimistic. Check actual exposures against the planned number. - **Variance underestimated.** The power calculation assumed a metric variance from historical data; if the metric is skewed or driven by a heavy tail (revenue per user is the usual offender), realised variance runs higher and every interval widens. - **Dilution of the audience.** If a large share of assigned users never reached the changed surface, they contribute noise and no signal, and the measured effect shrinks toward zero while the interval stays wide. - **Metric or unit mismatch.** Randomising by user but analysing by session, or the reverse, distorts the standard error. - **Imbalance or instrumentation loss.** A split that did not come out as designed, or logging that dropped events in one arm, both degrade precision. ## What to do **Extend, if extending is honest.** More exposures shrink the interval. This is legitimate only if the extension was planned as an option with an appropriate stopping rule, or the test is simply restarted with the correct sample size — peeking repeatedly and stopping the moment an interval clears zero inflates the wrong-decision rate. **Reduce variance instead of buying traffic.** Where the surface cannot deliver more users, precision can sometimes be recovered by analysing a less noisy primary metric, trimming or capping a heavy tail by a pre-declared rule, or moving to a triggered population if the trigger is properly logged. **Or accept that this question will not be answered by this test.** Low-traffic surfaces often cannot support a 2% MDE at all. Then the honest output is: the experiment was uninformative, and the decision falls back to cost, strategy, and qualitative evidence. That is a defensible call. What is not defensible is laundering it as "we tested it and it did nothing". ## The reporting trap Inconclusive tests get written up as null results because null results are tidy. Six months later somebody cites "we tested that, it was flat" as a reason not to revisit the idea, and the organisation has acquired a false belief from an experiment that established nothing. Record inconclusive readouts with the interval attached and the achieved precision stated, so the next reader can see what the test could and could not have detected.

  • What interval would have licensed you to call this test genuinely flat?
    One whose whole range sits inside the band of effects nobody would act on — upper bound below the 2% MDE and lower bound above the harm you would care about, say [-0.3%, +0.5%]. That result rules out a launch-worthy lift, which is a real conclusion rather than an absence of one.
  • The test hit its planned sample but the interval is still wider than the MDE. What happened?
    The power calculation's variance assumption was too low, or the analysis unit differs from the randomisation unit. Skewed metrics such as revenue per user routinely come in noisier than a historical estimate suggests. Re-derive the required sample from the realised variance before promising the same MDE again.
  • Can you just run the test longer until the interval excludes zero?
    Not as an unplanned reaction to the readout. Repeatedly extending and stopping at the first interval that clears zero inflates the false-win rate well beyond the nominal level. Either extend to a pre-declared sample size, or restart the test correctly powered for the MDE you need.
  • How should this result be written into the experiment log?
    As inconclusive, with the estimate, the interval, the achieved precision and the MDE it failed to reach. Stating what the test could have detected stops a future reader from citing it as evidence the idea does not work.

saying these in an interview costs you the question

  • Calling any non-significant interval a flat result
  • Ignoring that the interval reaches past the MDE
  • Extending the test until the interval finally excludes zero
  • Reporting no effect when precision was never achieved
  • Reading the +0.4% point estimate as a small real win

context