What does it mean to say a hypothesis test has 80% power?
answer
- a detection rate, not a correctness rate
- conditional on an effect really existing
- one minus the miss probability
- quoted only at one assumed effect size
- 80% leaves a one-in-five miss
basics
~20 sPower is the probability a test rejects the null when a specified real effect exists. 80% power means that if an effect of that size is truly present, the test detects it 80% of the time.
solid answer
~50 sPower is `1 - beta`: the probability of getting a statistically significant result when the alternative is true at a stated effect size. So 80% power is a promise about detection, not about correctness — if the real effect is exactly the size you assumed, four runs in five come back significant and one in five comes back null even though the effect is there. The number is meaningless without the assumptions behind it: the alpha level, one- or two-sided, the sample size, the outcome's variability, and above all the effect size you are powering for. Power is really a curve over possible true effects, not a single number — a study with 80% power for a large effect may have 20% power for a small one. The 80% figure is a convention, not a statistical law.
go deeper
Be ready to state the definition cleanly: power is the probability of detecting an effect that is really there, and it equals one minus the miss rate. Remember that 80% power means missing one real effect in five.
Explain the mechanics: power is conditional on an assumed effect size, so it is a curve rather than a number, sitting at alpha when the true effect is zero and rising toward 1. Name every input a quoted power figure depends on.
Interviewers expect you to audit power claims in real work — challenge an implausibly large assumed effect, check whether the realised variability matched the planning assumption, and refuse to read a null from a low-power design as evidence of absence.
Own the convention itself. Argue when 80% is the wrong target for your organisation's decisions, make the implicit miss-versus-false-alarm tolerance explicit, and set a standard for how power assumptions are documented and revisited.
## The definition Power is the probability that a test produces a statistically significant result, *given* that a real effect of a particular size exists. Formally, power = `P(reject H0 | H1 true at effect size delta)`. Its complement is beta, the probability of failing to reject when that effect is real, so power = `1 - beta`. A test with 80% power has beta = 0.20. The word *given* carries the entire idea. Power is a conditional probability, and the condition is an assumption you supply, not something the data tell you. Nothing about power says how likely it is that the effect exists at all. ## Power is a curve, not a number Because power is defined at a stated effect size, a single test has a different power for every possible true effect. Trace power on the vertical axis against the true effect on the horizontal axis and you get a power curve: - At a true effect of exactly zero, the curve sits at alpha — the test still rejects at its nominal significance rate. - As the true effect grows, the curve rises monotonically. - For a large enough effect the curve flattens toward 1: the test essentially always detects it. This is why 'our study has 80% power' is an incomplete sentence. It is shorthand for 'our study has 80% power to detect an effect of size X, at alpha = 0.05, two-sided, with n = N per group, assuming the outcome's standard deviation is S'. Change any of those and the 80% changes. A study powered at 80% for a 10-point difference might have only 25% power for a 4-point difference — and 4 points might be the difference you actually care about. ## What 80% actually commits you to Eighty percent is a convention popularised as a default planning target, not a threshold with any mathematical status. Accepting it means accepting that one real effect in five, at the assumed size, will slip past you and be written up as a null. Paired with the usual alpha = 0.05, it encodes an implicit judgment that a miss is about four times more tolerable than a false alarm (beta = 0.20 versus alpha = 0.05). That ratio deserves to be a decision, not a habit: a screening step that will be confirmed later can live with 60% power, while an expensive, hard-to-reverse call may warrant 95%. ## Four things power is not 1. **Not the probability the result is correct.** Power says nothing about whether your particular conclusion is right; it is a long-run detection rate over hypothetical repeated studies. 2. **Not `1 - alpha`.** The confidence level and power are different quantities computed under different assumed truths — alpha under the null, power under the alternative. 3. **Not the probability the effect exists.** That is a question about the world (or, in Bayesian terms, a posterior), and no power calculation addresses it. 4. **Not a property of the collected data.** Power is a design-time property of the *procedure* plus an assumed truth. Once the data are in, the informative summary is the estimate and its interval, not a recomputed power figure. ## Reading a power claim critically When someone quotes a power number, ask three questions. *Powered for what effect?* If the assumed effect is far larger than anything plausible, the reassuring 80% is fiction. *Under what variability?* Power calculations use an assumed variance; if the real data are noisier, realised power is lower than planned. *At what alpha and sidedness?* A one-sided test at alpha = 0.05 has more power than a two-sided one, which is fine only if the direction was fixed before seeing data. ## Why interviewers ask Underpowered analysis is the most common quiet defect in applied statistics. A candidate who can state power as a conditional detection probability, tie it to a specific effect size, and resist the slide from 'not significant' to 'no effect' is showing the habit of mind the whole topic exists to test.
- Why is quoting a study's power without naming an effect size meaningless?Power is a function of the true effect, not a fixed property of the design. The same test has near-alpha power against a tiny effect and near-certain power against a huge one. Without the effect size — plus alpha, sidedness, sample size and assumed variability — the number 80% identifies nothing you can check or reuse.
- What does a power curve plotted against the true effect size look like?It starts at alpha when the true effect is zero, because a valid test rejects at its nominal rate under the null. It then rises monotonically as the true effect grows, steeply through the middle range, and flattens asymptotically toward 1 for effects large relative to the noise. The point where it crosses 0.80 is the effect the study is conventionally described as powered for.
- Does 80% power mean there is an 80% chance the effect is real?No. Power is computed assuming the effect is real and of a stated size; it is the detection rate under that assumption. The probability that an effect exists depends on prior plausibility and the evidence together, which no power calculation supplies. Conflating the two is the most common error on this topic.
Power is the sensitivity of a metal detector. Eighty percent power means that if a coin of a particular size is buried there, the detector beeps four passes out of five — and it says nothing about whether a coin is buried at all.
saying these in an interview costs you the question
- Says power is the probability the result is correct
- Quotes a power figure without naming the effect size
- Confuses power with the confidence level, one minus alpha
- Thinks 80% power means an 80% chance the effect exists
- Treats power as a property of the collected data