skip to content

A two-sample test returns p = 0.20. Does that show there is no effect?

level: seniorimportance: should knowfreq 55%

answer

  1. the test only points one direction
  2. unsurprising under the null, and under others too
  3. look at the range, not the single number
  4. wide interval means uninformative, not negative
  5. absence of evidence versus evidence of absence

basics

~20 s

No. A p-value of 0.20 says the data are compatible with the null of no difference, but they are usually compatible with a range of real differences too. Absence of evidence against a hypothesis is not evidence that it holds.

solid answer

~50 s

A large p-value only says the data are not surprising under the null; it never confirms the null. The test is built to be one-directional — it can make a null look implausible, but it has no mechanism for making one look true, because the same data are typically also unsurprising under many nearby non-null values. The correct move is to look at the interval estimate around the observed difference: if it runs from a meaningful negative effect to a meaningful positive one, the honest summary is "this study cannot tell", not "no effect". If instead the whole interval sits inside a range everyone agrees is negligible, you can make a positive claim of practical equivalence — but that comes from the interval, not from the p-value. Report the estimate with its uncertainty and describe the result as inconclusive rather than negative.

go deeper

for a junior

Learn the phrase and its meaning: a large p-value means the data are compatible with no difference, not that no difference exists. Never write 'there is no effect' from a p-value alone.

for a middle

Explain the asymmetry — a test can undermine the null but has no machinery to confirm it — and show that the same data are usually consistent with a range of non-zero differences too.

for a senior

Demonstrate the constructive move: read the interval, separate an uninformative study from one that precisely measured a negligible effect, and write the conclusion so a reader cannot misuse it.

for a principal

Own how negative results are handled across the org — require an interval and a pre-declared margin before any parity, safety or no-regression claim, so 'not significant' can never quietly become 'proven safe'.

## The asymmetry at the heart of testing A hypothesis test is deliberately one-directional. It sets up a null — say, no difference between two groups — computes how surprising the observed data would be if that null held, and reports that surprise as a p-value. When the p-value is small, the data are hard to reconcile with the null and you have an argument against it. When the p-value is large, all you have learned is that the data are **not surprising** under the null. That is a much weaker statement than it sounds, because the same data are usually also not surprising under a whole neighbourhood of *other* hypotheses. A difference of 0 is compatible with the data; so, very often, is a difference of 2%, or 5%, or whatever the interval reaches. The test only ever asked about one of those values, so it cannot single it out as the truth. The slogan is: **absence of evidence is not evidence of absence.** p = 0.20 is an absence of evidence against the null. It is not evidence for it. ## What to do instead: read the interval The practical fix is to stop reasoning from the p-value alone and look at the **interval estimate** around the observed difference — the range of values for the true difference that the data are consistent with. Three shapes, three very different conclusions: **1. Wide interval spanning meaningful values in both directions.** For example, the difference in conversion could plausibly be anywhere from -3 points to +4 points. The p-value is large, but the study is simply **uninformative**: it has not ruled out an important improvement or an important harm. The honest report is "inconclusive" and the next question is what it would take to get a usable answer. **2. Narrow interval sitting entirely inside the range nobody cares about.** For example, the difference is somewhere between -0.1 and +0.1 points, and the team agreed beforehand that anything under 0.5 points is irrelevant. Here you *can* make a positive claim — practical equivalence — but note where it came from: the narrowness of the interval relative to a pre-agreed threshold of importance, not the size of the p-value. This is the only route to a defensible "no meaningful difference" statement. **3. Narrow interval sitting slightly off zero but inside the irrelevant range.** Same conclusion as (2). Whether the interval happens to include exactly zero is not the interesting fact; whether it includes anything worth acting on is. Notice that shapes (1) and (2) can both produce p = 0.20. The p-value alone cannot distinguish "we learned nothing" from "we learned the effect is negligible", and those two conclusions lead to completely different decisions. ## Language discipline How a large p-value is written down determines how it is used downstream. Compare: - **"There is no difference between the variants."** A claim about the world the data do not support. - **"The test was not significant, so we are keeping the current version."** Better, but it hides whether the study was informative. - **"The estimated difference was +0.4 points, with the data consistent with anything from -1.2 to +2.0 points. The study cannot distinguish a meaningful gain from a meaningful loss, so we are keeping the current version and will need a larger sample to decide."** This is what a senior analyst writes: estimate, uncertainty, and an explicit statement of what remains unresolved. ## Where this bites in practice - **Safety and regression checks.** "We tested for a latency regression and it wasn't significant" is often used to close a ticket. If the interval reaches +80 ms, the check has not demonstrated safety at all. - **Killing initiatives.** A promising change is abandoned on a large p-value from a small trial, when the interval was consistent with a substantial win. - **Claiming parity.** Two suppliers, two models, two flows declared "the same" because a test did not reject. Equivalence needs to be argued from an interval and a pre-agreed threshold, and it needs to be planned before the data arrive, not decided afterwards. ## The interview answer Say no, and say why: the test is asymmetric and can only ever provide evidence against the null. Then immediately go to the interval, distinguish "uninformative" from "precisely estimated as negligible", and note that a genuine claim of no meaningful difference requires stating in advance what magnitude would count as meaningful. That progression — refuse the fallacy, then supply the constructive alternative — is what separates a senior answer from a recited slogan.

  • What would let you make a defensible claim that two variants perform equivalently?
    A pre-agreed margin of what counts as a negligible difference, plus an interval estimate that lies entirely inside that margin. The claim then rests on having measured the difference precisely enough to exclude anything that matters, which is a positive finding — unlike a large p-value, which merely fails to exclude zero.
  • Two studies both report p = 0.20. Why might one be far more useful than the other?
    Because the p-value hides precision. One may have an interval from -0.1 to +0.2 points, ruling out anything worth acting on; the other an interval from -5 to +8 points, ruling out nothing. Identical p-values, opposite practical conclusions — which is why the interval and the estimate have to be reported.
  • How would you phrase a non-significant regression check in a release note?
    State the estimated change and the range the data are consistent with, then say explicitly what has and has not been excluded. For example: latency changed by +3 ms, consistent with anywhere from -10 ms to +16 ms, so a regression above 16 ms is ruled out but a smaller one is not. That tells the reader what the check actually bought.

A metal detector that stays silent has not proved the field is empty; it has proved nothing was found with that detector, at that sensitivity, on that pass.

saying these in an interview costs you the question

  • Says a non-significant result proves there is no effect
  • Reports only the p-value with no estimate or interval
  • Claims two variants are equivalent because the test did not reject
  • Confuses an uninformative study with a negative finding
  • Treats p = 0.20 as an 80% chance the effect is real

context