skip to content

A drug lowers systolic blood pressure by 0.4 mmHg with p < 0.001 in 120,000 patients - is that worth adopting?

level: seniorimportance: must knowfreq 58%

answer

  1. two verdicts, evidence and value
  2. precision is not importance
  3. detectable is not worth acting on
  4. set the threshold before the data
  5. compare the interval to that threshold

basics

~20 s

Almost certainly not. A 0.4 mmHg reduction is far below any clinically meaningful change; the 120,000 patients merely make so tiny an effect distinguishable from zero. The decision needs magnitude weighed against cost and harm, not a test verdict.

solid answer

~40 s

The result is statistically significant and practically negligible, and those are separate judgments. A 0.4 mmHg drop in systolic blood pressure is far smaller than the reductions that translate into meaningful cardiovascular benefit, so no p-value can make it worth prescribing. What the enormous sample buys is precision, not importance: the interval around 0.4 is narrow, which means the study has confidently established that the effect is trivial. That is genuinely useful — it rules out a worthwhile effect rather than leaving the question open. My process is to fix the smallest effect worth acting on with clinicians *before* seeing results, then compare the estimate and its interval with that threshold. I would also weigh side effects, cost and existing alternatives, all of which are magnitudes the significance verdict never touches.

go deeper

for a junior

Be ready to say that a significant result can still be too small to matter, and that you would ask for the size of the change in its original units before drawing any conclusion.

for a middle

Expect to explain why a very large sample makes even trivial effects detectable, and to distinguish a narrow interval around a tiny effect from a wide interval around an unknown one.

for a senior

Show the working process: a smallest-effect-of-interest agreed with domain owners before analysis, the interval compared against it, and costs, harms and alternatives brought into the recommendation.

for a principal

Own the policy that no result reaches a decision forum without a magnitude, an interval and a pre-registered threshold, and be ready to defend killing a statistically flawless programme on effect size alone.

## Two verdicts, not one Every quantitative result carries two separable judgments. **Statistical significance** asks whether the observed effect is distinguishable from no effect given the noise in the data. It is a statement about evidence. **Practical significance** asks whether an effect of the estimated size is large enough to change what anyone does. It is a statement about value, and it requires domain knowledge that no statistical procedure contains. The blood-pressure result separates them cleanly. A 0.4 mmHg reduction is a real, well-established effect — and clinically inert. Blood-pressure targets move in units of many mmHg; a fraction of one is below the noise of an individual measurement, let alone the threshold at which outcomes change. Prescribing a drug, with its cost, its side effects and the burden of daily adherence, in exchange for 0.4 mmHg is a bad trade regardless of how confidently the 0.4 is known. ## What the huge sample actually bought A common but incorrect reaction is that the study is somehow untrustworthy because it is so large. The opposite is true. What a very large sample delivers is **precision**: the interval around the estimate is narrow. If the estimated 0.4 mmHg comes with an interval of roughly 0.3 to 0.5, the study has done something valuable — it has ruled out a clinically meaningful effect. Compare that with a 40-patient trial reporting the same 0.4 mmHg point estimate with an interval from -6 to +7: those two results share a point estimate and share nothing else. The first closes the question; the second has not begun to answer it. So the correct framing is not "the sample was too big" but "the sample was big enough to demonstrate that the effect is too small to matter". A significance test asks whether the effect is zero, and almost no real intervention has an effect of exactly zero. With enough data, non-zero effects of any size become detectable. That is why the test alone cannot be the decision rule. ## Fixing the threshold before the data The discipline that prevents this argument from happening after the fact is to state, in advance, the **smallest effect size of interest**: the magnitude below which the intervention is not worth adopting. It is set from cost, risk and available alternatives, by the people who own those, and written into the analysis plan before results are seen. With a threshold in hand, the reading of a result becomes mechanical: - **Interval entirely below the threshold** — the effect exists but is not worth acting on. This is our blood-pressure case, and it is a clear no. - **Interval entirely above the threshold** — worth acting on, subject to the other costs. - **Interval straddling the threshold** — the study cannot answer the decision; either collect more data or decide on other grounds. Notice that in this scheme the comparison of interest is between the *interval* and a *substantive threshold*, not between a p-value and 0.05. Deciding the threshold afterwards, once the estimate is visible, destroys the protection — the number will drift to whatever justifies the conclusion already preferred. ## Everything else the verdict cannot see Even with the magnitude in hand, an adoption decision has to weigh several things a test result says nothing about. **Cost and harm.** Side effects, price, and adherence burden are magnitudes on the other side of the ledger. A trivial benefit loses to almost any non-trivial cost. **Alternatives.** The comparison that matters is against the current best option, not against nothing. A drug beating placebo by a hair while an established therapy achieves far more is not a candidate. **Heterogeneity.** An average of 0.4 mmHg can hide a subgroup with a meaningful response. That is a hypothesis for a pre-specified follow-up study, not a conclusion to mine out of the same data after the fact — post-hoc subgroup hunting produces impressive-looking effects that rarely replicate. **Outcome validity.** Blood pressure is a surrogate for what is actually cared about, which is cardiovascular events. Even a large surrogate effect must be argued through to the real outcome before it justifies adoption. ## How to report it Write the estimate first, with its units and interval, then the verdict, then the comparison against the pre-agreed threshold. "A 0.4 mmHg reduction (interval 0.3 to 0.5), reliably non-zero but an order of magnitude below the 5 mmHg we agreed would justify adoption" is a complete, decision-ready sentence. "Highly significant reduction in blood pressure (p < 0.001)" is the same result written so that it will be misread.

  • How do you decide in advance what magnitude counts as meaningful?
    Elicit it from the people who own the costs — clinicians, product owners, finance — by asking what effect would actually change their decision given the price, the risks and the alternatives. Write that number into the analysis plan before any results are seen, so it cannot drift to fit whatever the data turned out to show.
  • What would you conclude if the interval ran from -0.1 to 0.9 mmHg instead?
    That the study is inconclusive on the null but conclusive on the decision. Even the most favourable end of the interval, 0.9 mmHg, is far below anything clinically useful, so the data have ruled out a worthwhile effect regardless of whether zero is excluded. Failing to reach significance is beside the point here.
  • Does the tiny effect mean the trial was a waste?
    No. A precisely estimated null-for-practical-purposes result is a real finding: it closes a question and stops further investment in that compound. The waste would be an underpowered trial producing a wide interval that leaves everyone arguing, or a large trial whose write-up sells the verdict and hides the magnitude.

saying these in an interview costs you the question

  • Treats p < 0.001 as proof the drug matters clinically
  • Says the sample was too large so the result is untrustworthy
  • Never asks what magnitude would change the decision
  • Sets the meaningful-effect threshold after seeing the estimate
  • Reports the verdict without the estimate and its units

context