How should you report an A/B test with p = 0.3 without claiming 'no difference'?
answer
- absence of evidence is not evidence of absence
- report the range, not the verdict
- same p-value, opposite conclusions
- check whether the bounds are small
- say what the interval excludes
basics
~10 sReport the estimated lift with its confidence interval and say what the interval rules out. A p-value of 0.3 means the data are compatible with no difference, not that the difference is zero.
solid answer
~40 sA p-value of 0.3 says the observed difference would not be surprising if the true lift were zero. It does not say the true lift is zero, and it carries no information about how big the effect could still be. So write the readout as the estimate plus its interval: "+0.6% lift, 95% CI -1.1% to +2.3%". Then say what that range excludes — here, anything better than about +2.3% or worse than about -1.1%. If those bounds sit inside the band nobody would act on, you have a genuine no-meaningful-effect finding and can say so in exactly those words. If they do not, the honest label is inconclusive, and the write-up should record the precision the test achieved so a future reader knows what it could and could not have detected.
go deeper
Be ready to state that a large p-value means the data are compatible with no difference, not that the difference is zero, and to write a readout as estimate plus interval.
Explain how two results with the same p-value can license opposite conclusions, and describe the pre-declared band that makes a no-meaningful-effect statement defensible.
Show that you record achieved precision in the experiment log so an inconclusive result is not later cited as proof the idea fails, and that you resist rewriting a null into a subgroup win.
Own the write-up standard: what a null readout must contain, who decides the band of effects too small to act on, and how the organisation stops null results hardening into unexamined institutional beliefs.
## What p = 0.3 actually claims The p-value is the probability, computed assuming the true difference between the arms is exactly zero, of seeing a test statistic at least as extreme as the one observed. At 0.3, a result like this one is entirely ordinary under a no-difference world, so there is no basis for rejecting that hypothesis. The logical step that fails is the next one. Failing to find evidence against a hypothesis is not evidence for it. The test asked "are these data surprising if the effect is zero?" and got "no". It never asked "is the effect zero?", and the same data would have been unsurprising under many non-zero effects too. ## The reporting sentence that does work Write the estimate, the interval, and what the interval rules out: > Treatment showed **+0.6%** relative lift, 95% CI **-1.1% to +2.3%**. The data are compatible with anything from a 1.1% regression to a 2.3% gain; effects outside that range are not well supported. Everything a reader needs is present: the direction and size of what was seen, the precision, and the bound. Compare the two shorthands people reach for instead: - **"No difference"** — asserts a fact the test cannot establish, and quietly converts an absent finding into a permanent one. - **"Not significant"** — technically correct, decision-useless. It cannot distinguish a tight interval around zero from one that spans the whole roadmap. ## When you may say the effect is not meaningful There is a legitimate version of the no-effect claim, and it depends entirely on the interval, not the p-value. Decide in advance the band of effects too small to act on. If the whole interval falls inside that band, you can report that the experiment found no effect worth acting on, and that statement is supported. Two readouts make the difference concrete, both at p = 0.3: - `+0.1%, 95% CI [-0.4%, +0.6%]` against a band of plus or minus 1% — a real finding. Anything large enough to matter has been excluded. - `+1.5%, 95% CI [-4%, +7%]` against the same band — nothing has been established. A 7% win and a 4% loss are both fully compatible with these data. The p-values are the same and the conclusions are opposite. This is the single most useful thing to be able to explain about a non-significant result. ## Practical habits - **Lead with the interval, not the p-value.** If a template forces a p-value, put it after the interval. - **State the band you are judging against.** "No effect worth acting on" is meaningless until "worth acting on" has a number, and that number should come from the test plan rather than from the readout. - **Never write the point estimate alone.** "+0.6%" with no interval reads as a small win and will be repeated as one. - **Do not shop for a subgroup.** Slicing a null result until something clears significance manufactures findings; that whole practice belongs to a different discipline of pre-registration and correction, not to the readout of this metric. - **Record achieved precision in the log.** Six months on, the only defence against "we tested that and it did nothing" is a written interval showing what the test could have detected. ## Why this matters beyond pedantry Organisations accumulate beliefs from experiment write-ups. A null result written as "no difference" becomes institutional knowledge that the idea does not work, and the idea does not get retried — even when the original test was too small to have detected a large effect. Writing the interval instead costs one extra clause and keeps the record honest about what was and was not learned.
- When is it legitimate to state that an experiment found no meaningful effect?When the whole interval falls inside a band of effects declared too small to act on before the test ran. Then the data have excluded everything worth a decision. If the interval reaches past that band in either direction, the correct label is inconclusive rather than no effect.
- Two tests both report p = 0.3. Can they support different conclusions?Yes, and that is the point of reporting intervals. A tight interval such as [-0.4%, +0.6%] bounds the effect below anything actionable; a wide one such as [-4%, +7%] establishes nothing at all. The p-values are identical and the decisions are opposite.
- Should the point estimate be reported at all when the result is not significant?Yes, but never on its own. The estimate is the best single guess and readers will ask for it, so report it attached to its interval. A bare +0.6% in a summary slide gets repeated as a small win long after the uncertainty has been forgotten.
A metal detector that stays quiet has not proved the field is empty — unless you know it would have beeped for anything worth digging up.
saying these in an interview costs you the question
- Saying the test proved there is no difference
- Reporting only 'not significant' with no interval
- Quoting the point estimate as a small confirmed win
- Treating p = 0.3 as weak evidence for the null
- Hunting for a slice that finally reaches significance