A two-sample test returns p = 0.20. Does that show there is no effect?
answer
- the test only points one direction
- unsurprising under the null, and under others too
- look at the range, not the single number
- wide interval means uninformative, not negative
- absence of evidence versus evidence of absence
basics
~20 sNo. A p-value of 0.20 says the data are compatible with the null of no difference, but they are usually compatible with a range of real differences too. Absence of evidence against a hypothesis is not evidence that it holds.
solid answer
~50 sA large p-value only says the data are not surprising under the null; it never confirms the null. The test is built to be one-directional — it can make a null look implausible, but it has no mechanism for making one look true, because the same data are typically also unsurprising under many nearby non-null values. The correct move is to look at the interval estimate around the observed difference: if it runs from a meaningful negative effect to a meaningful positive one, the honest summary is "this study cannot tell", not "no effect". If instead the whole interval sits inside a range everyone agrees is negligible, you can make a positive claim of practical equivalence — but that comes from the interval, not from the p-value. Report the estimate with its uncertainty and describe the result as inconclusive rather than negative.
go deeper
Learn the phrase and its meaning: a large p-value means the data are compatible with no difference, not that no difference exists. Never write 'there is no effect' from a p-value alone.
Explain the asymmetry — a test can undermine the null but has no machinery to confirm it — and show that the same data are usually consistent with a range of non-zero differences too.
Demonstrate the constructive move: read the interval, separate an uninformative study from one that precisely measured a negligible effect, and write the conclusion so a reader cannot misuse it.
Own how negative results are handled across the org — require an interval and a pre-declared margin before any parity, safety or no-regression claim, so 'not significant' can never quietly become 'proven safe'.
## The asymmetry at the heart of testing A hypothesis test is deliberately one-directional. It sets up a null — say, no difference between two groups — computes how surprising the observed data would be if that null held, and reports that surprise as a p-value. When the p-value is small, the data are hard to reconcile with the null and you have an argument against it. When the p-value is large, all you have learned is that the data are **not surprising** under the null. That is a much weaker statement than it sounds, because the same data are usually also not surprising under a whole neighbourhood of *other* hypotheses. A difference of 0 is compatible with the data; so, very often, is a difference of 2%, or 5%, or whatever the interval reaches. The test only ever asked about one of those values, so it cannot single it out as the truth. The slogan is: **absence of evidence is not evidence of absence.** p = 0.20 is an absence of evidence against the null. It is not evidence for it. ## What to do instead: read the interval The practical fix is to stop reasoning from the p-value alone and look at the **interval estimate** around the observed difference — the range of values for the true difference that the data are consistent with. Three shapes, three very different conclusions: **1. Wide interval spanning meaningful values in both directions.** For example, the difference in conversion could plausibly be anywhere from -3 points to +4 points. The p-value is large, but the study is simply **uninformative**: it has not ruled out an important improvement or an important harm. The honest report is "inconclusive" and the next question is what it would take to get a usable answer. **2. Narrow interval sitting entirely inside the range nobody cares about.** For example, the difference is somewhere between -0.1 and +0.1 points, and the team agreed beforehand that anything under 0.5 points is irrelevant. Here you *can* make a positive claim — practical equivalence — but note where it came from: the narrowness of the interval relative to a pre-agreed threshold of importance, not the size of the p-value. This is the only route to a defensible "no meaningful difference" statement. **3. Narrow interval sitting slightly off zero but inside the irrelevant range.** Same conclusion as (2). Whether the interval happens to include exactly zero is not the interesting fact; whether it includes anything worth acting on is. Notice that shapes (1) and (2) can both produce p = 0.20. The p-value alone cannot distinguish "we learned nothing" from "we learned the effect is negligible", and those two conclusions lead to completely different decisions. ## Language discipline How a large p-value is written down determines how it is used downstream. Compare: - **"There is no difference between the variants."** A claim about the world the data do not support. - **"The test was not significant, so we are keeping the current version."** Better, but it hides whether the study was informative. - **"The estimated difference was +0.4 points, with the data consistent with anything from -1.2 to +2.0 points. The study cannot distinguish a meaningful gain from a meaningful loss, so we are keeping the current version and will need a larger sample to decide."** This is what a senior analyst writes: estimate, uncertainty, and an explicit statement of what remains unresolved. ## Where this bites in practice - **Safety and regression checks.** "We tested for a latency regression and it wasn't significant" is often used to close a ticket. If the interval reaches +80 ms, the check has not demonstrated safety at all. - **Killing initiatives.** A promising change is abandoned on a large p-value from a small trial, when the interval was consistent with a substantial win. - **Claiming parity.** Two suppliers, two models, two flows declared "the same" because a test did not reject. Equivalence needs to be argued from an interval and a pre-agreed threshold, and it needs to be planned before the data arrive, not decided afterwards. ## The interview answer Say no, and say why: the test is asymmetric and can only ever provide evidence against the null. Then immediately go to the interval, distinguish "uninformative" from "precisely estimated as negligible", and note that a genuine claim of no meaningful difference requires stating in advance what magnitude would count as meaningful. That progression — refuse the fallacy, then supply the constructive alternative — is what separates a senior answer from a recited slogan.
- What would let you make a defensible claim that two variants perform equivalently?A pre-agreed margin of what counts as a negligible difference, plus an interval estimate that lies entirely inside that margin. The claim then rests on having measured the difference precisely enough to exclude anything that matters, which is a positive finding — unlike a large p-value, which merely fails to exclude zero.
- Two studies both report p = 0.20. Why might one be far more useful than the other?Because the p-value hides precision. One may have an interval from -0.1 to +0.2 points, ruling out anything worth acting on; the other an interval from -5 to +8 points, ruling out nothing. Identical p-values, opposite practical conclusions — which is why the interval and the estimate have to be reported.
- How would you phrase a non-significant regression check in a release note?State the estimated change and the range the data are consistent with, then say explicitly what has and has not been excluded. For example: latency changed by +3 ms, consistent with anywhere from -10 ms to +16 ms, so a regression above 16 ms is ruled out but a smaller one is not. That tells the reader what the check actually bought.
A metal detector that stays silent has not proved the field is empty; it has proved nothing was found with that detector, at that sensitivity, on that pass.
saying these in an interview costs you the question
- Says a non-significant result proves there is no effect
- Reports only the p-value with no estimate or interval
- Claims two variants are equivalent because the test did not reject
- Confuses an uninformative study with a negative finding
- Treats p = 0.20 as an 80% chance the effect is real