skip to content

How does a 95% confidence interval relate to a two-sided test at the 5% level?

level: middleimportance: should knowfreq 58%

answer

  1. two views of the same computation
  2. which candidate values survive the test
  3. the interval is the non-rejected set
  4. excludes the value if and only if it rejects

basics

~20 s

They are two readings of one computation. A 95% confidence interval is the set of null values a two-sided test at the 5% level would not reject, so the interval excludes a value exactly when the test rejects that value.

solid answer

~40 s

A 95% interval and a two-sided 5% test are dual: the interval is precisely the collection of candidate parameter values that the test fails to reject at that level. So if a 95% interval for a difference in means is `[0.4, 2.1]`, the null of zero difference is rejected at 5%, because zero is outside; if the interval were `[-0.3, 1.8]`, the same test would not reject. The correspondence is exact when the interval and the test are built from the same statistic and the levels match - a 95% interval pairs with alpha = 0.05, not with alpha = 0.01. The interval is the more useful report of the two: it delivers the same accept/reject verdict while also showing the magnitude and the whole range of values still compatible with the data.

go deeper

for a junior

Be ready to use the shortcut correctly: check whether the null value sits inside the 95% interval and say whether a two-sided 5% test rejects. Getting the direction right is what is being tested.

for a middle

Explain the mechanics behind the shortcut - that the interval is the set of null values surviving the test, that it answers many hypotheses at once, and that the level and the alpha must match.

for a senior

Demonstrate judgment about reporting: know when interval and test can disagree because they were built from different statistics, and insist on a matched pair rather than arguing over which output wins.

for a principal

Own the reporting standard. Argue for estimates with intervals instead of binary verdicts across the organisation, and be ready to explain what that changes about how borderline results get discussed and acted on.

## One computation, two presentations Hypothesis tests and confidence intervals are usually taught in separate chapters, which hides the fact that they are the same machinery viewed from different ends. The link is called **duality**, and stating it cleanly is a standard interview probe because it shows whether a candidate has an internal model or two memorised recipes. The statement: **a 100(1 - alpha)% confidence interval consists of exactly those candidate parameter values that a two-sided test at level alpha would fail to reject.** Equivalently, the interval excludes a particular value if and only if the two-sided test of that value as the null rejects at level alpha. That gives the everyday shortcut. Does a 95% interval for a treatment effect contain zero? If no, a two-sided test of "the effect is zero" rejects at 5%. If yes, it does not reject. You never had to run the test separately - the interval already answered it. ## Reading it in both directions **Interval to test.** Given an interval for a difference of `[0.4, 2.1]`: zero is outside, so a two-sided test of the zero-difference null rejects at 5%. But zero is not the only null you can read off. A hypothesised difference of 1.0 lies inside, so a test of "the difference is 1.0" would *not* reject at 5%. The interval answers infinitely many tests at once. **Test to interval.** Fix a level and ask "which candidate values survive the test?" Sweep every possible value of the parameter, test each one as a null, and keep the survivors. The set you collect is the confidence interval. This construction is called *inverting the test*, and it is how intervals are derived for parameters where no neat closed form exists. ## The level has to match The pairing is between a 95% interval and alpha = 0.05, a 99% interval and alpha = 0.01, a 90% interval and alpha = 0.10. Mixing them is a common error: a value sitting outside a 90% interval does not license a rejection at 1%. Higher confidence means a wider interval means a more conservative test, and this monotone relationship is directly readable. If a 90% interval excludes zero while a 99% interval includes it, the two-sided p-value for the zero null lies somewhere between 0.01 and 0.10. ## Where the correspondence is only approximate Duality is exact when the test and the interval are built from the same underlying statistic. In practice a team sometimes reports an interval derived one way and a p-value computed from a differently constructed statistic; the two are then only approximately equivalent, and in borderline cases they can disagree - an interval that just barely excludes a value while the reported test just barely fails to reject. That is not a paradox, it is two different procedures, and the fix is to report a matched pair rather than to argue about which one is right. One-sided tests pair with one-sided intervals - a bound on one side and infinity on the other - so if you are quoting a two-sided interval, keep the test two-sided too. ## Why the interval is the better report Since the interval carries the test's verdict for free, everything else it carries is a bonus: - **Magnitude.** "Rejected" says a difference exists; `[0.4, 2.1]` says how big it plausibly is, in the outcome's own units. - **Precision.** A tight interval and an extremely wide one can both exclude zero, and they warrant very different confidence in the size of the effect. - **The full compatible set.** Reviewers can check their own hypothesis against the interval, rather than having to ask you to rerun a test. - **Fewer binary cliffs.** Reporting the range makes it harder to treat a result just inside a threshold as categorically different from one just outside. This is why so much reporting guidance now asks for the estimate and its interval rather than the bare verdict. ## Answering well State the duality as an if-and-only-if, demonstrate it on a concrete interval by reading off two different nulls, note that the level must match, and close with the practical reason it matters: you already have the test, so report the range and let the reader test whatever value they care about.

  • Why do reporting guidelines increasingly prefer the interval over a bare reject or fail-to-reject verdict?
    Because the interval contains the verdict and more. You can read the accept/reject decision off it directly, and you additionally get the magnitude in real units, the precision, and the full set of parameter values still compatible with the data. A reader can check their own hypothesis against it. The verdict alone collapses all of that into one bit and invites treating a result just past a threshold as categorically different from one just short of it.
  • A 90% interval excludes zero but the 99% interval includes it. What does that tell you about the two-sided p-value?
    It pins the p-value for the zero null between 0.01 and 0.10. Excluding zero at 90% means the test rejects at alpha = 0.10, so p is below 0.10; including zero at 99% means it does not reject at alpha = 0.01, so p is above 0.01. Nested intervals at different levels bracket the p-value without your computing it separately.
  • Does a value falling inside the interval mean that value is the true parameter?
    No. Inside means only that the value is not rejected at that level, which is a statement about compatibility, not proof. A wide interval fails to reject a great many mutually contradictory values at once, and they obviously cannot all be true. Not rejected is the weakest possible endorsement, and treating it as confirmation is the classic misuse of the duality.

saying these in an interview costs you the question

  • Treats the interval and the test as unrelated procedures
  • Says an interval cannot tell you anything about significance
  • Pairs a 95% interval with a 1% test level
  • Reads a value inside the interval as proven rather than not rejected
  • Compares a two-sided interval against a one-sided test

context