Why do significant results from underpowered studies overstate the true effect?
answer
- the surviving results are a selected set
- only tail draws clear the threshold
- worse the lower the power
- conditioning on significance selects upward noise
- the average significant estimate exceeds the truth
basics
~20 sIn a low-power study only unusually large estimates clear the significance threshold, so the ones that do come from the extreme tail of the sampling distribution. Averaged over such studies, the surviving estimates exaggerate the true effect, sometimes several-fold.
solid answer
~40 sPower is low when the true effect is small relative to the noise. In that regime the estimate must land far out in the tail of its sampling distribution to clear the significance threshold at all, so conditioning on `p < alpha` selects exactly the runs where noise pushed the estimate upward. The expected magnitude of a significant estimate therefore exceeds the true effect, and the exaggeration grows as power falls: at very low power a significant estimate can be twice the truth or more, and can even carry the wrong sign. This is a property of one study's own sampling distribution, not of any downstream filtering. It is why small-n literatures produced strong-looking findings that shrank sharply when repeated at scale.
go deeper
Remember the headline: a significant result from a small, noisy study is likely to overstate how big the effect really is. Do not read a large reported effect as confirmation of a large true effect.
Explain the mechanism in terms of the sampling distribution: at low power only estimates far out in the tail cross the significance threshold, so the significant subset is selected on favourable noise.
Demonstrate that it changes your practice — quoting intervals rather than points, estimating what power the design had for a plausible effect, and refusing to plan follow-up work off an exaggerated number.
Own the organisational consequence. Set expectations that early small-sample wins will shrink, build confirmation steps into how findings graduate to decisions, and stop inflated estimates from hardening into planning assumptions.
## The selection mechanism Suppose the true effect is real but small, and your study is noisy. The estimate you compute is the true effect plus sampling error. To reach significance, the estimate must exceed a fixed threshold set by alpha and the estimate's own uncertainty. When power is low, the true effect sits well below that threshold — which means the only way a run reaches significance is for the sampling error to be large and in the favourable direction. So the set of significant results is not a random sample of what the study could produce. It is the upper tail. The average magnitude among significant estimates is therefore strictly greater than the true effect. The gap between the two is often called the exaggeration factor, and it is a deterministic consequence of conditioning on significance — no misconduct, no selective filing, no questionable analysis choices required. ## How bad it gets The exaggeration is governed by power. - At high power, say 90%, most runs clear the threshold and the significant subset is close to the full distribution. Inflation is mild. - At 50% power, roughly half the runs are excluded and the significant estimates skew noticeably high. - At 10-20% power, only the most extreme runs survive, and the average significant estimate can be a multiple of the truth. - At extreme low power, a further failure appears: some estimates clear the threshold in the *wrong direction*. A study can report a confidently significant effect whose sign is opposite to reality. Note the perverse pairing. Low power means most real effects are missed — and it also means the few that are caught are misleading about magnitude. An underpowered design is not merely a weaker version of a good design; it is a design whose successes cannot be trusted quantitatively. ## Where this showed up historically The clearest illustrations come from literatures built on very small samples. Early candidate-gene association studies, run on dozens or a few hundred subjects, reported large effects for single variants that turned out, once studies with tens of thousands of subjects arrived, to be far smaller or absent. Small-sample neuroscience studies produced striking effect estimates that shrank markedly when the same designs were run with substantially more subjects. In both cases the original analyses were often perfectly competent; the problem was that at those sample sizes only an exaggerated estimate could have reached significance in the first place. ## What it changes about how you read a result **Significance is not a magnitude guarantee.** A p-value below 0.05 tells you the data were unlikely under the null. It tells you nothing about how close your estimate is to the truth, and in a low-power setting it actively predicts that the estimate is too big. **Look at the interval, not the point.** A wide confidence interval that just excludes zero is the signature of a threshold-clearing tail draw. If the interval's lower bound is far below anything practically meaningful, the study has not established a meaningful effect even though it is significant. **Ask what power the study had for a plausible effect.** Not for the effect it found — for the effect you actually believe in. If that number is 20%, treat the point estimate as an upper bound and the finding as hypothesis-generating. **Do not feed the estimate forward.** Using an inflated estimate as the assumed effect for the next study's power calculation is self-perpetuating: the next study is planned to detect an effect that does not exist at that size, so it is underpowered for the real one and, if it reaches significance, exaggerates again. **Prefer pooling over picking.** Combining several studies, including the ones that came out null, undoes some of the tail selection because the null runs carry the downward information that the significant ones lack. ## The senior signal What an interviewer is listening for is whether you can explain the mechanism — conditioning on a threshold selects the tail — without reaching for misconduct as the explanation, and whether the insight changes your behaviour: shrinking your expectation of a small significant study, quoting intervals, and refusing to plan the next piece of work off an inflated number.
- Does this exaggeration happen inside a single study, or only across many studies?Inside a single study. The selection is on that study's own sampling distribution: to be significant at low power, the estimate had to land in the tail. Any process that also discards null results compounds the problem, but it is not required — the inflation is already present in the one study you are reading.
- How should you report a significant finding from a small, noisy study?Lead with the interval, not the point estimate, and state its width honestly. Say what power the design had for a plausible effect size, note that a significant estimate at that power is expected to be inflated, and frame the result as motivating a larger confirmatory study rather than as a magnitude anyone should plan against.
- At very low power, can a significant estimate carry the wrong sign?Yes. When the true effect is small relative to the noise, the sampling distribution puts non-trivial mass beyond the threshold on the opposite side. Some significant results therefore point the wrong way. The probability is small in absolute terms but can be a meaningful share of all significant results when power is in single digits.
- Why is it dangerous to reuse an inflated estimate in the next power calculation?Because the next study is then sized to detect an effect larger than the real one, which leaves it underpowered for the truth. It will usually return a null, and when it does reach significance it will exaggerate again. Power the follow-up off the smallest effect that would change a decision instead.
If you only keep the days a noisy thermometer reads above 30 degrees, the days you keep will average well above the true temperature — not because the thermometer is biased, but because you selected on its errors.
saying these in an interview costs you the question
- Treats significance as evidence the estimate is accurate
- Says a small study finding significance is unusually strong evidence
- Thinks a smaller p-value implies a larger true effect
- Blames only selective reporting, missing the within-study selection
- Feeds the inflated estimate into the next power calculation