Your analysis can reach only 40% power; how do you decide whether to run it at all?
answer
- the number alone does not decide it
- check the levers before conceding
- what decision does the result feed
- pilots and pooled estimates are fine
- pre-commit to how a null reads
basics
~20 sDecide by what the result will be used for. A 40%-power run is defensible as a pilot or as one input to a pooled estimate, and indefensible if a null will be read as no effect.
solid answer
~40 sFirst try to move the levers: reduce outcome variance through better measurement or a paired design, extend the collection window, pool with a sister team, or raise the target effect you are willing to detect. If 40% is genuinely the ceiling, ask what decision the result feeds. Running is reasonable when the study is a pilot that also buys a variance estimate, when the estimate and its interval will be pooled with other evidence, or when nothing irreversible hangs on it. Running is unwise when a null will be reported as absence of effect, when a significant point estimate would be taken at face value, or when the population can only be studied once. Either way, pre-commit in writing to how each outcome will be read.
go deeper
Know that 40% power means most real effects go undetected, so a null from such a study proves very little. Flag low power to someone senior rather than interpreting the result on your own.
Be ready to work the levers before conceding: variance reduction, a longer collection window, pooling, or restating the target effect. Explain why each one moves power and what it costs.
Show judgment about use. Distinguish a defensible pilot or pooled contribution from a study whose null will be misread, and write down the interpretation rules before the data land.
Own the policy. Justify the power target from the decision's cost and reversibility, defend running below convention where it is genuinely cheaper to learn in stages, and make sure low-power findings never harden into planning assumptions.
## Reject the false binary first '80% or don't run' is not a principle, it is a habit. Eighty percent power is a planning convention that encodes a particular tolerance — a one-in-five miss rate for the assumed effect — and there is nothing in the mathematics that privileges it. The real question is whether the study, at whatever power is achievable, produces information worth its cost given how the result will be used. ## Step one: exhaust the levers Before accepting 40%, check each input honestly. - **Variance.** Usually the cheapest lever and the one nobody tried. Better instrumentation, repeated measurement per unit, a paired or within-subject comparison, adjustment for a strong pre-treatment covariate, restricting to a more homogeneous population. Halving the outcome variance is worth roughly as much as doubling the sample. - **Sample.** Longer collection windows, additional sites or teams, pooling with a parallel effort. - **Target effect.** Is 40% power quoted against an effect nobody believes, or against the smallest effect that would change a decision? These give very different diagnoses. If the design has 40% power for a tiny effect but 85% for the effect that would actually trigger action, there may be no problem at all. - **Alpha.** Moving to 0.10 is a legitimate choice for a screening step whose winners will be confirmed later — but it must be declared in advance with the higher false-alarm rate stated, never adjusted afterwards. ## Step two: ask what the result feeds If 40% survives that scrub, classify the decision. **Reasonable to run.** The study is a pilot whose main products are feasibility and a variance estimate for a properly powered successor. The estimate will enter a pooled analysis where null runs count as evidence too. The result is one input among several and the team is comfortable with a wide interval. The cost of running is near zero and the population is not consumed by being studied. **Unwise to run.** Someone will read a null as 'no effect' — at 40% power, a real effect is missed most of the time, so a null is close to uninformative. Someone will act on the point estimate if it is significant — at 40% power that estimate is expected to be inflated, so the magnitude will not replicate. The study consumes a scarce, non-repeatable population or budget that a properly powered version would need. Or the outcome is irreversible, in which case the right power target is above 80%, not below it. ## Step three: pre-commit to the interpretation If you run, write down before the data arrive what each outcome will and will not license: - Intervals lead the report; point estimates never appear alone. - A null is reported as 'this design could not distinguish effects between X and Y', with the interval quoted, not as absence of effect. - A significant result is framed as motivating confirmation, with an explicit note that the magnitude is expected to shrink. - The follow-up study is powered against a prespecified effect of interest, never against whatever this run produced. Writing this down before the result exists is what separates a defensible low-power study from one that will be misused. It is also the part a principal-level candidate is expected to volunteer without prompting. ## Setting the target as policy Zoom out and the same reasoning sets the organisation's default. Eighty percent is a middle setting. Cheap, reversible, repeatable decisions can run below it, because a miss simply gets another attempt and the cost of extra data is real. Expensive, irreversible, safety- or trust-affecting decisions should sit at 90-95%, because a miss is not recoverable and the cost of another run is small next to the cost of being wrong. State the target as a consequence of the decision's cost profile, and revisit it when that profile changes. ## What the interviewer is listening for Three things. That you treat 80% as a convention with a rationale rather than a rule. That you go looking for the levers — especially variance — before conceding the power number. And that you distinguish *running* a low-power study from *over-reading* one: the defect is almost never the run itself, it is the interpretation that gets attached to it afterwards.
- What would you pre-commit to before running a study you know has 40% power?That intervals lead every report and point estimates never stand alone; that a null will be described as the range of effects the design could not distinguish, not as absence of effect; that a significant result is labelled hypothesis-generating with the magnitude expected to shrink; and that any follow-up is powered against a prespecified effect of interest rather than this run's estimate.
- When is 80% power the wrong target?Too low when the decision is expensive, irreversible or safety-relevant — there 90 to 95 percent is warranted because a missed real effect cannot be recovered. Too high when the step is a cheap, repeatable screen whose candidates will be confirmed later; there a 60 to 70 percent screen with a confirmation stage detects more true effects per unit of budget than one large, cautious study.
- Your team wants to publish the null from the 40%-power run as evidence the intervention does nothing. What do you say?That the design misses a real effect of the assumed size roughly six times in ten, so failing to see one is weak evidence of absence. Show the confidence interval: if it still contains effects large enough to matter, the honest statement is that the study could not distinguish them. Only an interval lying entirely inside the negligible range supports a claim of no meaningful effect.
saying these in an interview costs you the question
- Reports a null from a low-power run as evidence of no effect
- Treats 80% power as a hard rule with no stated rationale
- Ignores that a significant result at 40% power will be inflated
- Concedes the power number without trying variance reduction
- Raises alpha quietly to make the power target look met