skip to content

Why is post-hoc power computed from an observed effect uninformative?

level: seniorimportance: should knowfreq 38%

answer

  1. the input is the study's own estimate
  2. no information beyond what the p-value gave
  3. a monotone function of the p-value
  4. nulls always look underpowered this way
  5. report the interval instead

basics

~20 s

Observed power is computed from the effect the study just estimated, so for a fixed test it is a one-to-one function of the p-value and adds no new information. A null result always yields low observed power.

solid answer

~40 s

Post-hoc or observed power plugs the study's own estimate back in as the assumed true effect. Because the estimate and the p-value are two views of the same statistic, observed power is a strictly decreasing function of the p-value for that test, so a large p always maps to a small observed power. The sentence 'the result was null and observed power was only 30%, which explains it' is therefore circular. For a two-sided z test at alpha = 0.05, a p-value of exactly 0.05 gives observed power of about 50%. What is informative afterwards is the estimate with its confidence interval, plus a power figure computed against a prespecified effect of practical interest.

go deeper

for a junior

Remember the rule of thumb: do not compute power after the fact from the effect your own study estimated. If a result is not significant, report the estimate and its confidence interval instead.

for a middle

Explain why it is circular — the observed estimate and the p-value are two views of the same statistic, so recomputed power is just the p-value in different units and always looks low after a null.

for a senior

Show what you do instead in real analyses: quote the interval, judge whether it excludes effects that matter, and run a design analysis against a prespecified effect of interest when someone asks whether the study could have found anything.

for a principal

Own the reporting standard. Rule observed power out of your team's write-ups, require intervals and prespecified effects of interest, and be able to explain the reasoning to stakeholders who ask for a power number after the fact.

## What observed power is A power calculation needs an assumed true effect. Observed (post-hoc) power supplies that assumption from the finished study: take the effect estimate you just computed, treat it as if it were the truth, and ask what fraction of hypothetical repeat studies would reach significance. It sounds like a diagnostic. It is not, because the input is the very quantity the test already summarised. ## The circularity For a given test, alpha and sample size, the p-value and the effect estimate are two encodings of the same test statistic. Feeding the estimate into a power formula therefore produces a number that is a deterministic, monotone decreasing function of the p-value. Large p, small observed power; small p, large observed power. Nothing else can happen. The calibration point worth memorising: for a two-sided z test at alpha = 0.05, an observed p of exactly 0.05 gives observed power of about 50%. The observed estimate sits exactly at the critical value, so in the hypothetical repeat world half the draws land above it and half below. Two consequences follow. **Every non-significant result has low observed power.** Not because the design was weak — because `p > alpha` mathematically forces observed power below roughly 50% for that test. Reporting it as an explanation for the null is like explaining a failed exam by noting the score was below the pass mark. **Observed power cannot support the null.** People sometimes reason: high observed power plus a null result means the effect is truly absent. But a null result cannot produce high observed power in the first place, so the premise never occurs. ## What people are actually trying to ask The underlying question is legitimate: *was this study capable of detecting an effect I would have cared about?* The mistake is answering it with the effect the study happened to see rather than the effect that matters. The honest versions: **Report the interval.** The estimate with its confidence interval answers the question directly and without circularity. A confidence interval running from a large negative effect to a large positive one says the study was uninformative — you learned almost nothing. An interval hugging zero, with both ends inside the range of effects too small to act on, says something much stronger: this study rules out any effect worth caring about. **Do a design analysis against a prespecified effect.** Choose the smallest effect that would change a decision — from theory, prior work, or the economics of the decision, but never from this dataset — and compute what power the completed design had against it. That number is meaningful, is not a repackaged p-value, and directly answers whether the null is informative. It also tells you honestly what a replication would need. **Frame absence explicitly if that is the claim.** If the goal is to argue that no meaningful effect exists, an equivalence-style framing — asking whether the interval sits entirely inside a prespecified range of negligible effects — states the claim as a testable proposition rather than resting on a failure to reject. ## Where the confusion comes from Power is a design-stage concept: it is a property of a procedure plus an assumed truth, computed before data exist. Once the data are in, the procedure has run and the informative quantities are the estimate and its uncertainty. The urge to recompute power afterwards comes from wanting a single number that says 'this null is meaningful' — but the interval already is that number, and it does not require pretending the sample estimate is the truth. ## Answering in an interview Say three things. First, observed power is a monotone function of the p-value, so it adds no information. Second, that is why every null result looks underpowered by this measure, which makes the explanation circular. Third, name the replacements: the confidence interval, and a prospective-style power calculation against an effect size chosen independently of the data. If you can also state the p = 0.05 gives roughly 50% observed power calibration, you have shown you understand the mechanism rather than repeating a slogan.

  • What should you report instead of observed power after a null result?
    The effect estimate with its confidence interval. If the interval spans effects from clearly harmful to clearly beneficial, the study was uninformative and you should say so. If it sits tightly around zero, inside the range of effects too small to act on, you can make the much stronger claim that no meaningful effect is present.
  • Is any retrospective power calculation legitimate?
    Yes, provided the assumed effect comes from outside the data — a prespecified minimum effect of interest, a value from prior literature, or the smallest effect that would change the decision. Computing what power the completed design had against that effect is a design analysis. It is informative precisely because its input is independent of the result.
  • Why does a non-significant test always produce low observed power?
    Because observed power is a decreasing function of the p-value for a fixed test. If p exceeds alpha, the observed estimate lies inside the acceptance region, which forces the recomputed power below roughly 50% for a two-sided test at that alpha. The low number is a restatement of non-significance, not an independent diagnosis.

saying these in an interview costs you the question

  • Explains a null result by citing power computed from the observed effect
  • Claims high observed power makes a null result strong evidence of no effect
  • Does not realise observed power is a function of the p-value
  • Uses observed power to argue the study merely needs more subjects
  • Confuses a data-derived power figure with a design analysis

context