Why do we say we fail to reject the null hypothesis instead of accepting it?
answer
- the guarantee only runs one way
- absence of evidence, evidence of absence
- many nearby effects also fit the data
- a weak study fails to reject anything
basics
~20 sA test controls only the risk of wrongly rejecting a true null; it gives no guarantee that a null which survives is true. A non-significant result is a verdict of not proven, not evidence of no effect.
solid answer
~40 sThe procedure is built to control one risk only — how often a true null is wrongly rejected — and offers no matching guarantee in the other direction. A non-significant result says the data are compatible with `H0`, but they are usually just as compatible with a range of small effects in `H1` that this sample was too small or too noisy to separate from zero. Saying "we accept `H0`" treats absence of evidence as evidence of absence. "Fail to reject" keeps the asymmetry visible: rejection is a positive conclusion the design supports, non-rejection is "not proven". If you genuinely need to argue a difference is negligible, that requires a design built around a pre-stated tolerance, not a non-significant result read backwards.
go deeper
Memorise the phrasing and the reason behind it: rejecting is a conclusion, not rejecting is not. Never write that a study accepted or proved the null hypothesis.
Explain the mechanism — the procedure bounds only the rate of wrongly rejecting a true null — and why a non-significant result is compatible with a whole range of small non-null effects.
Demonstrate the habit of reporting the compatible range of effects alongside the verdict, and of calling an underpowered non-result uninformative instead of reassuring.
Own how this language leaves your organisation: a non-significant check turning into "no impact" on a summary slide is a governance problem, and the fix is a reporting standard, not a reminder.
## The asymmetry is structural, not stylistic The wording looks like pedantry until you see where it comes from. A hypothesis test is a procedure with exactly one guarantee attached: **if the null is true, the chance it is wrongly rejected is at most `alpha`**, typically 5%. That guarantee is built into the critical value. Nothing in the construction bounds how often a *false* null survives — that depends on how big the real effect is and how much data you collected, neither of which the test controls. So the two outcomes carry very different weight: - **Reject `H0`** — backed by a designed error rate. The data would be unusual if the null were true. - **Fail to reject `H0`** — backed by nothing in particular. The data were not unusual enough under the null. That is all it says. "Accept the null" borrows the credibility of the first outcome and applies it to the second. ## Why non-rejection is so weak The deeper problem is that `H0` is a point claim while the data are compatible with a whole neighbourhood of values. Suppose a bottling line is tested against `H0: mu = 500` and the result is non-significant. The data are compatible with 500. They are usually also compatible with 499.4, with 501.2, and with everything in between — values that live in `H1` and may well matter operationally. The test simply could not tell them apart from 500 at the precision available. This gets worse the weaker the study is. A tiny sample, or a very noisy measurement, produces a wide range of compatible values and therefore fails to reject almost any null you hand it. Under that logic "accepting the null" would reward the sloppiest experiments with the strongest-sounding conclusions — the exact inversion of what evidence should do. **A non-significant result from a weak design is uninformative, not reassuring.** ## Absence of evidence is not evidence of absence The phrase is worn but it is precisely right here. "We looked and did not find it" only counts as evidence that it is not there if you looked somewhere it would have been visible. A test that could not have detected a meaningful effect even if one existed provides no information about whether one exists. This is why the honest follow-up to a non-significant result is not "so there is no effect" but "what effects can we still not rule out?" — a question about the precision of your estimate rather than about the verdict of the test. ## What you can legitimately report After a non-significant result, a defensible write-up says: 1. The data did not provide sufficient evidence against `H0` at the stated significance level. 2. Here is the estimated effect and the range of effect sizes the data remain compatible with. 3. Here is whether that range excludes the effects we would have cared about. Point 3 is where the real information lives. If the compatible range is narrow and excludes everything operationally meaningful, you have learned something substantive — "no effect larger than this" — and you learned it from the precision of the estimate, not from the test's verdict. If the range is wide and includes effects that would change decisions, the correct conclusion is "this study was not informative", and the correct next step is more data, not a claim of no effect. ## When you actually need to prove sameness Sometimes the claim you must establish genuinely is "these are not meaningfully different". The standard test cannot deliver it by failing to reject, because failing to reject is what an underpowered study does automatically. Establishing sameness requires flipping the setup so that the claim to be demonstrated is what the data must support, against a threshold of what counts as meaningfully different — a threshold that has to be stated in advance on substantive grounds. The key interview point is simply that this is a **different design**, decided before data collection, not a reinterpretation of a non-significant result after the fact. ## The language to use Acceptable: "we fail to reject `H0`", "the data are consistent with `H0`", "we found no evidence of a difference at this level", "any difference is smaller than we could detect here". Not acceptable: "we accept `H0`", "we proved there is no difference", "the groups are the same", "the effect is zero". The second list is not just loose phrasing. Each one converts a non-result into a claim the procedure never made, and once it enters a summary slide, nobody downstream can tell it apart from an actual finding. ## Interview framing A strong answer names the one-sided guarantee, explains that non-rejection is compatible with a range of non-null values, notes that weak designs fail to reject by default, and closes with what you would report instead. Candidates who stop at "you can never prove a negative" have the slogan but not the mechanism.
- What can you legitimately report after a non-significant result?Report the verdict and the estimate together: the data provided no evidence against the null at the stated level, and here is the range of effect sizes still compatible with what you observed. If that range excludes everything operationally meaningful, say so — that is a real finding. If it includes effects that would change a decision, the honest conclusion is that the study was uninformative.
- Does a very large sample make non-rejection stronger evidence for the null?More informative, but still not proof. With a large sample the range of effects compatible with the data is narrow, so non-rejection rules out large effects and you can say "nothing bigger than this". That statement comes from the precision of the estimate, not from the test's verdict, and it never establishes exact equality — an arbitrarily small real effect remains possible.
- Your stakeholder reads a non-significant guardrail check as proof the change was harmless. How do you respond?Separate the two claims. The test showed no detectable harm at the precision available; whether that rules out harm depends on how wide the compatible range of effects is. Show them the range next to the level of harm they would consider unacceptable. If the range covers it, the correct message is that the check could not answer the question, not that the change is safe.
saying these in an interview costs you the question
- Writes "we accept the null hypothesis" in a report.
- Concludes the two groups are identical in the population.
- Treats a non-significant result as proof of no effect.
- Ignores that a small or noisy sample fails to reject almost anything.
- Reports the verdict without any estimate of the effect.