skip to content

A 95% confidence interval for a treatment effect includes zero - does that prove there is no effect?

level: seniorimportance: should knowfreq 48%

answer

  1. absence of evidence is not evidence of absence
  2. read the endpoints, not just zero
  3. which effect sizes are still compatible
  4. a wide interval versus a tight one
  5. pre-specify what counts as negligible

basics

~20 s

No. Containing zero only means zero is among the values compatible with the data. Read the endpoints: a wide interval also contains large effects, so nothing is ruled out, while a tight interval around zero genuinely rules out anything big.

solid answer

~40 s

Containing zero means zero was not ruled out, which is not the same as ruling everything else out. The information is in the endpoints. An interval of `[-9.0, 10.2]` percentage points contains zero and also contains a large harm and a large benefit, so it supports no conclusion beyond "we learned very little." An interval of `[-0.1, 0.2]` on the same scale also contains zero but excludes every effect anyone would call meaningful, and that is a genuinely informative near-null result. To claim no meaningful effect as a positive finding rather than a shrug, decide in advance the smallest effect that would matter and check that the entire interval falls inside that negligible range - that is an equivalence claim, and it is the only version of "no effect" the data can support.

go deeper

for a junior

Be ready to say that an interval containing zero means zero was not ruled out, not that the effect is zero, and to look at how wide the interval is before summarising it.

for a middle

Explain the contrast between a wide interval containing zero and a tight one: identical verdicts, opposite information. Be able to name which effect sizes each interval leaves compatible with the data.

for a senior

Show you handle this in practice - reframing an inconclusive result honestly for stakeholders, and knowing that a positive claim of no meaningful difference needs a pre-specified negligible range the whole interval sits inside.

for a principal

Own the standard for how null-looking results are allowed to be reported and acted on, including who sets the smallest effect worth caring about and how that number is protected from being renegotiated after the fact.

## Absence of evidence, evidence of absence An interval that straddles zero is one of the most consistently over-read outputs in applied statistics. It gets summarised as "no effect," then as "we proved the change does nothing," and by the third retelling it is a settled fact. What actually happened is narrower: zero survived, along with every other value between the endpoints. The interval is a **compatibility set**. Read it as "given these data and this model, the parameter values in this range are the ones not ruled out at this level." Zero being a member of that set makes it one candidate among many. Whether the result is informative depends entirely on who else is in the set. ## Read the endpoints, not the membership of zero Suppose the outcome is a conversion rate in percentage points and stakeholders care about anything above 2 points. **Case one: `[-9.0, 10.2]`.** Zero is inside. So is a 9-point harm and a 10-point benefit. The data are compatible with a disaster, with nothing, and with a triumph. The correct summary is not "no effect" but "this measurement cannot distinguish between outcomes that would lead to completely different decisions." **Case two: `[-0.1, 0.2]`.** Zero is inside again, and so is nothing else of consequence: every value in the interval is far below the 2-point threshold that matters. This is a real, useful finding. It does not prove the effect is exactly zero - no finite dataset does - but it does say any effect is too small to care about. The two cases produce the identical binary verdict and carry opposite information. That is why reporting the verdict alone destroys the analysis, and why an interviewer asking this question is watching for whether you look at the width before you speak. ## Turning "no meaningful effect" into a claim you can defend The standard move is to state, **before** seeing the results, a range of effects you would consider practically negligible - often written as a smallest effect of interest in each direction. The claim of equivalence is then supported when the whole interval lies inside that negligible range. This is deliberately harder than noting that zero is inside, and it should be: a positive claim of no meaningful difference deserves a positive test, not the leftovers of a failed one. Declaring the threshold in advance also stops the threshold from being negotiated afterwards to match whatever interval arrived, which is how equivalence claims quietly turn into rationalisations. ## Other reasons to be careful - **The coverage caveat still applies.** A 95% procedure misses the true value about 5% of the time. One interval containing zero is not a verdict from nature; it is one draw from a procedure with a known long-run error rate. - **The estimate is not zero.** An interval of `[-0.1, 3.0]` has a point estimate somewhere well above zero and is quite compatible with a meaningful effect. Reporting it as "no effect" throws away the fact that the bulk of the compatible range is positive. - **Repetition does not settle it.** A single near-null result on its own says what it says; combining evidence properly is a different exercise from declaring the question closed. - **Zero may not be the interesting reference.** For a change that costs something to keep, the question may be whether the effect clears a cost threshold, and the interval answers that just as directly by asking whether the threshold is inside it. ## How to phrase the finding A defensible sentence names the estimate, the interval, and the range the data exclude. For example: "the estimated effect is 0.6 points with a 95% interval from -9.0 to 10.2, so these data are compatible with anything from a substantial harm to a substantial benefit and do not distinguish between them." Or, in the tight case: "the estimated effect is 0.05 points with a 95% interval from -0.1 to 0.2, so any effect is well under the 2-point threshold that would matter." Both sentences take one line, neither says "no effect," and both let a reader make their own decision. ## What the interviewer is checking That you do not reflexively translate "interval contains zero" into "no effect"; that you reach for the endpoints and the practical threshold instead; and that you know a positive claim of no meaningful difference requires a pre-specified negligible range rather than an inconclusive result recycled as a conclusion.

  • How would you word a finding with a 95% interval of -9.0 to 10.2 percentage points for a stakeholder summary?
    Something like: the estimated effect is small, but the 95% interval runs from a 9-point harm to a 10-point benefit, so these data are compatible with outcomes that would lead to opposite decisions and do not distinguish between them. That sentence gives the estimate, the range, and the honest limitation in one line, and it avoids the phrase no effect, which the data do not support.
  • The interval is tight and excludes zero, but the effect is 0.2 points. Is that an important result?
    Statistically distinguishable from zero, yes; important, only if 0.2 points clears the threshold that matters for the decision. Precision makes small effects detectable, so significance and practical relevance separate as data accumulate. The right response is to compare the whole interval against the smallest effect worth acting on, and to report the magnitude in the outcome's units rather than leaning on the exclusion of zero.
  • How do you stop a pre-specified negligible range from being adjusted after the results arrive?
    Write it down before the analysis, with the reasoning that ties it to a real cost or benefit, and have it reviewed by whoever owns that decision. Record it where it is visible next to the result. If the threshold genuinely needs revisiting, do it explicitly and say so in the write-up, rather than quietly picking the number that makes the interval fit the desired conclusion.

saying these in an interview costs you the question

  • Reports an interval containing zero as proof of no effect
  • Checks only whether zero is inside and ignores the endpoints
  • Treats a very wide inconclusive interval as a null result
  • Claims equivalence without a pre-specified negligible range
  • Forgets that the procedure still misses the truth some of the time

context