Two 95% confidence intervals overlap - does that prove the two group means do not differ significantly?
answer
- eyeballing two bars is not a test
- test the difference, not each mean
- variances add before the square root
- a gap of 2.8 SEs versus 3.9 SEs
basics
~20 sNo. Overlapping intervals can still accompany a significant difference, because the difference has its own smaller standard error: variances add, then you take the square root. Test the difference directly instead of comparing two intervals by eye.
solid answer
~40 sOverlap proves nothing. The right object is an interval on the difference itself, whose standard error is `sqrt(SE_A^2 + SE_B^2)` - smaller than `SE_A + SE_B`, which is effectively what eyeballing two separate intervals uses. Concretely: group A has mean 10.0 with SE 1.0, giving roughly `[8.0, 12.0]`, and group B has mean 13.0 with SE 1.0, giving roughly `[11.0, 15.0]`. They visibly overlap, yet the difference is 3.0 with SE `sqrt(2) = 1.41`, so the interval on the difference is about `[0.2, 5.8]` and the two-sided test rejects at 5%. The implication only runs one way: for independent estimates, non-overlap does imply a significant difference, so it is a conservative signal, while overlap is simply uninformative.
go deeper
Be ready to say plainly that overlapping intervals do not settle the comparison, and that the right move is to compute an interval for the difference between the groups.
Explain why: variances add for a difference of independent estimates, so its standard error is the root of the sum of squares rather than the sum. Being able to work a small numeric example is expected here.
Show the operating judgment - catching this pattern in someone else's dashboard, knowing that non-overlap is a valid one-way signal, and adjusting when the estimates are paired or otherwise correlated.
Own the visualisation and reporting standard that prevents the error at scale: which contrast gets the headline interval, what error bars on a chart are allowed to mean, and how those defaults are enforced across teams.
## The habit and why it is wrong Put two estimates side by side with their 95% error bars and the eye immediately wants to adjudicate: bars touch, so no real difference; bars clear each other, so a real difference. The second half of that instinct is roughly safe. The first half is a genuine error that changes conclusions. The reason is that comparing two groups is a question about **one** quantity - the difference - and that quantity has its own uncertainty, which is not obtained by lining up the two separate uncertainties. ## The arithmetic For two independent estimates, variances of a difference add: ``` Var(A - B) = Var(A) + Var(B) SE_diff = sqrt(SE_A^2 + SE_B^2) ``` The visual overlap test implicitly compares the gap between the estimates against `SE_A + SE_B` (scaled by the critical value), and the sum of two positive numbers is always larger than the square root of the sum of their squares. With equal standard errors `s`, `SE_diff = 1.41 s` while the sum is `2 s`. The eyeball test is therefore systematically too strict about declaring a difference. Write it in units of `s`, still with equal standard errors and 95% intervals on each estimate: - The two intervals stop overlapping only once the gap exceeds `2 x 1.96 x s = 3.92 s`. - The two-sided test on the difference already rejects once the gap exceeds `1.96 x 1.41 x s = 2.77 s`. Everything between `2.77 s` and `3.92 s` is the ambiguity zone: visibly overlapping bars, a difference that is nonetheless significant at 5%. That band is wide, and real comparisons land in it constantly. ## A worked case Group A: mean 10.0, standard error 1.0, so a 95% interval of about `[8.0, 12.0]`. Group B: mean 13.0, standard error 1.0, so a 95% interval of about `[11.0, 15.0]`. The intervals share the stretch from 11.0 to 12.0 - clear visual overlap. Now do the comparison properly. The difference is `13.0 - 10.0 = 3.0`, its standard error is `sqrt(1 + 1) = 1.41`, the test statistic is `3.0 / 1.41 = 2.12`, which exceeds 1.96, and the 95% interval on the difference is `3.0 +/- 1.96 x 1.41`, that is roughly `[0.2, 5.8]`. Zero is excluded, so the difference is significant at 5% despite the overlap. ## The asymmetry The two directions are not mirror images: - **Non-overlap implies significance** (for independent estimates at the same level). If the gap exceeds `1.96 (SE_A + SE_B)` it certainly exceeds `1.96 sqrt(SE_A^2 + SE_B^2)`, so the test rejects. Seeing separated bars is a conservative, valid signal. - **Overlap implies nothing.** It is consistent with a significant difference and with no detectable difference alike, so it should never be reported as evidence of similarity. A known consequence: if you insist on judging two independent estimates by eye, error bars drawn at roughly 84% rather than 95% make non-overlap correspond approximately to significance at 5% when the standard errors are similar. It is a plotting convention, not a substitute for the actual comparison, and it must be labelled or it will be misread as a 95% interval. ## When the estimates are not independent Everything above assumes the two estimates are independent. If they are not - repeated measures on the same units, before and after on one cohort, two metrics from overlapping populations - the covariance term enters: ``` Var(A - B) = Var(A) + Var(B) - 2 Cov(A, B) ``` Positive correlation shrinks the variance of the difference, sometimes dramatically, so a paired comparison can be far more precise than either marginal interval suggests and the overlap heuristic becomes even more misleading. Negative correlation pushes the other way. In either case the marginal intervals simply do not carry the information needed, and only a difference computed on the correct pairing does. ## What to do instead Report an estimate and an interval **for the contrast you care about**. If the comparison is the point of the analysis, the difference is the headline number and the per-group intervals are context. When you must plot the groups separately, add the difference and its interval in the caption or a companion panel, and never let a reader conclude "no difference" from touching bars. ## Interview framing Give the one-way implication, produce the variance-adds argument, and if you can, quote the `2.77 s` versus `3.92 s` band - it converts a vague warning into a precise claim, and it shows you can carry the algebra rather than repeat a slogan.
- What should you report instead of asking a reader to compare two intervals by eye?The estimate and interval for the difference itself, in the outcome's units. That single object answers the question being asked, carries the matched test's verdict, and shows how large the gap plausibly is. Per-group intervals are useful context for describing each group, but they are not the comparison. If the plot shows groups separately, put the difference and its interval alongside it.
- How does the picture change when the two estimates are correlated, such as before and after on the same units?The variance of the difference becomes Var(A) + Var(B) - 2 Cov(A, B). Positive correlation, which is typical for repeated measures on the same subjects, shrinks that variance, so the paired comparison can be much more precise than either marginal interval hints. The overlap heuristic gets worse, not better, because the marginal intervals no longer contain the information the comparison depends on.
- Is the reverse direction safe - if two 95% intervals clearly do not overlap, is the difference significant at 5%?Yes for independent estimates, and it is conservative. Non-overlap means the gap exceeds the critical value times the sum of the standard errors, which is always at least the critical value times their root-sum-of-squares, so the matched test rejects. It is a valid one-way signal - it just misses many significant differences, which is why it should not be run backwards.
Two people's arrival-time ranges can overlap while the gap between them is still solidly measurable: knowing when each arrived to within a few minutes tells you the interval between their arrivals more precisely than the two ranges, laid side by side, appear to allow.
saying these in an interview costs you the question
- Concludes no significant difference because the two intervals overlap
- Adds the two standard errors instead of adding the variances
- Applies the overlap rule to paired or otherwise correlated estimates
- Treats overlap and non-overlap as equally informative signals
- Compares group intervals by eye instead of reporting the difference