What four levers determine statistical power in a hypothesis test?
answer
- four inputs, only three are yours
- signal over noise, scaled by sample size
- one lever costs false alarms
- the underrated one is noise reduction
- halving variance equals doubling n
basics
~20 sPower rises with sample size, with the size of the true effect, with a larger alpha, and with lower outcome variance. Three of those are design choices; the true effect is not yours to set, only to assume honestly.
solid answer
~50 sThe four levers are sample size, the true effect size, the significance level alpha, and the variability of the outcome. More data raises power, with diminishing returns. A bigger true effect is easier to detect — but you cannot choose it, you can only choose which effect you are willing to power for. Relaxing alpha from 0.05 to 0.10 raises power by lowering the rejection threshold, at the price of more false alarms, so it is a decision to declare, not a knob to turn quietly. Reducing variance is the underrated lever: better measurement, a paired or within-subject design, covariate adjustment, or trimming a needlessly heterogeneous population all shrink noise. Effect and variance enter power together as a signal-to-noise ratio, which is why halving the outcome's variance buys roughly the same power as doubling the sample size.
go deeper
Be able to list the four inputs — sample size, effect size, alpha and variance — and say which direction each moves power. Knowing that more data and less noise both help is the baseline expectation.
Explain the mechanics: how enlarging the rejection region raises power, why n has diminishing returns, and why effect and variance act together as a signal-to-noise ratio rather than independently.
Show you can rescue a design in the field. Reach for variance reduction before more data, spot an optimistic assumed effect in a planning document, and refuse to move alpha after seeing results.
Frame the levers as a budget. Decide when buying variance reduction beats buying sample size, when a looser alpha is a legitimate organisational choice for screening work, and how assumed effects get justified rather than wished for.
## The four inputs Every power calculation, whatever the test, is a function of the same four quantities. **1. Sample size.** More observations make the sampling distribution of the estimate tighter, so a fixed true effect sits further from the rejection threshold in standardised terms. Power rises monotonically with n, but with sharply diminishing returns: moving from 100 to 200 observations buys far more power than moving from 1,000 to 1,100. The returns are governed by square-root scaling, which is why quadrupling n is needed to double the standardised signal, not doubling it. **2. The true effect size.** Bigger effects are easier to detect. This is the lever you do not own. You cannot make an intervention work better by writing a larger number into a planning spreadsheet — you can only decide which effect is worth being able to detect. Assuming an optimistic effect is the single most common way a study ends up nominally powered at 80% and actually powered at 30%. **3. The significance level alpha.** Raising alpha moves the rejection threshold closer to the null, so more of the alternative's sampling distribution falls in the rejection region and power goes up. Going from 0.05 to 0.10 is a real power gain and is sometimes the right call for an exploratory screen. It is never a free one, and switching alpha after seeing the data is not a power lever at all — it is a way of manufacturing a result. **4. Outcome variance.** Noise is the denominator of the signal-to-noise ratio the test is really evaluating. Anything that lowers it raises power at no cost in sample size: more precise measurement instruments, longer or repeated observation per unit to average out within-unit noise, a paired or within-subject design that differences away between-unit variation, adjusting for a strong pre-treatment covariate, restricting to a more homogeneous population, or replacing a heavy-tailed outcome with a better-behaved transformation or a more robust statistic. ## How they combine The effect and the variability never act separately: what drives power is the effect measured in units of the outcome's spread, scaled up by the square root of the sample size. Two consequences follow. First, **halving the outcome's variance is worth about as much as doubling the sample size**. Both multiply the standardised signal by the same factor. In many real settings the variance lever is far cheaper — a cleaner measurement or a within-subject design costs a design conversation, while doubling n costs money and calendar time. Second, **a doubling of n does not double power**. Power is bounded by 1 and follows an S-shaped curve, so the same extra sample can move power from 30% to 55% in one regime and from 90% to 94% in another. Always ask where on the curve you are standing before quoting a gain. ## Sidedness and test choice Two further design choices act like levers in practice. A one-sided test at the same alpha has more power than a two-sided one, but only if the direction of interest was fixed in advance; choosing sidedness after seeing which way the data went inflates the false-positive rate. And a test that matches the data's structure — a paired comparison for paired data, a test suited to the outcome's distribution — recovers power that a mismatched test throws away. ## The lever you must not pull Assuming a larger effect raises the *computed* power without changing anything about the study. This is how planning documents come to promise 80% power for effects nobody believes in. The honest move is the reverse: fix the smallest effect that would change a decision, then find out what n, variance and alpha it would take to detect it — and if the answer is out of reach, say so before running rather than after. ## Answering in an interview List all four levers, then immediately separate the ones you control from the one you do not, and name variance reduction explicitly. Most candidates say 'increase the sample size' and stop. Saying 'or halve the noise, which buys the same power for less money' is what distinguishes someone who has actually had to rescue an underpowered design.
- You cannot collect any more data. What is your best remaining route to power?Variance reduction. Improve measurement precision, average repeated observations per unit, switch to a paired or within-subject comparison, adjust for a strong pre-treatment covariate, or restrict to a more homogeneous population. If none of that is available, the honest options are to raise alpha with the extra false-positive cost stated openly, or to accept that only a larger effect is detectable and say so.
- Is halving the outcome's variance genuinely equivalent to doubling the sample size?For power purposes, yes to a good approximation. Power depends on the effect divided by the outcome's standard deviation and multiplied by the square root of n. Halving the variance divides the standard deviation by the square root of two; doubling n multiplies by the square root of two. Both scale the standardised signal by the same factor, so the resulting power is the same.
- Why does moving alpha from 0.05 to 0.10 increase power?It moves the rejection threshold closer to the null, enlarging the rejection region. More of the alternative's sampling distribution now falls inside it, so a real effect is detected more often. The same enlargement also lets more results through when nothing is there, so the false-positive rate doubles. It is a stated tradeoff, and it must be fixed before the data are seen.
Detecting an effect is like hearing a voice in a noisy room. You can move closer (more data), ask the speaker to talk louder (a bigger effect, if you could), lower your threshold for what counts as hearing something (alpha), or turn down the music (variance).
saying these in an interview costs you the question
- Names sample size as the only lever on power
- Raises the assumed effect size to hit a power target
- Forgets that alpha buys power at a false-positive cost
- Never mentions variance reduction or paired designs
- Thinks doubling the sample size doubles power