skip to content

Why report an effect size alongside the p-value from a hypothesis test?

level: juniorimportance: should knowfreq 66%

answer

  1. two different questions, one result
  2. whether versus how much
  3. a verdict carries no units
  4. decisions compare magnitudes to costs
  5. always pair estimate with an interval

basics

~20 s

A p-value only says whether an effect is distinguishable from zero; the effect size says how big it is. Decisions, cost-benefit comparisons and pooling results across studies all need the magnitude, not just the verdict.

solid answer

~40 s

A hypothesis test answers a yes/no question: is the observed pattern distinguishable from what we would expect if the null were true. It returns no magnitude, no units and no direction of practical importance. The effect size is the estimated magnitude itself — a raw difference in the outcome's own units (dollars per order, mmHg, percentage points), or a standardized version like Cohen's d that divides that difference by a standard deviation. I always report the point estimate with an interval around it, because the width tells the reader how precisely the magnitude is pinned down. That combination is what a decision actually consumes: someone has to judge whether an effect of *this* size, plausibly ranging over *this* interval, is worth the cost of acting. "Significant" alone cannot be weighed against a cost.

go deeper

for a junior

Be ready to state in one sentence that a test tells you whether an effect is distinguishable from zero while an effect size tells you how big it is, and to name one example such as a difference in means.

for a middle

Expect to explain the families of effect size — raw differences, standardized differences, ratio measures, variance explained — and to say which fits a continuous, binary or multi-group outcome.

for a senior

Show that you report the estimate with an interval and compare both against a magnitude threshold agreed before the analysis, and that you can tell a precisely-measured tiny effect from an imprecisely-measured unknown one.

for a principal

Own the reporting standard: what every analysis in the organisation must state, who sets the smallest effect worth acting on, and how you stop teams from shipping on verdicts that carry no magnitude.

## Two different questions A significance test and an effect size answer different questions about the same data. - The test asks: **is this pattern distinguishable from no effect?** Its output is a verdict-shaped quantity. - The effect size asks: **how much?** Its output is a magnitude, with units or an explicit standardization. A study can answer the first question loudly and the second one embarrassingly — a difference that is unmistakably not zero and also far too small to be worth anything. It can also fail the first while the second is interesting: an estimate of real practical size, measured too imprecisely to rule out zero. Neither situation is visible if only the test result is reported, which is why journals, regulators and experimentation platforms all now insist that an estimate accompany the verdict. ## What counts as an effect size The term covers any statistic that quantifies the magnitude of a phenomenon. The main families: **Raw (unstandardized) differences.** The difference in group means or proportions in the outcome's own units: 3 points on a test, 0.4 mmHg, 2.1 percentage points of conversion. These are the most decision-friendly numbers because a domain expert can price them directly. They are not comparable across outcomes measured on different scales. **Standardized differences.** The raw difference divided by a standard deviation — Cohen's d being the canonical example. Because the units cancel, a d computed on a reading test and a d computed on a reaction-time task sit on the same scale, which is what makes pooling across studies possible. The cost is interpretability: a stakeholder cannot price "0.3 standard deviations" without knowing what a standard deviation is worth. **Ratio measures.** For binary outcomes, the relative risk (ratio of two probabilities) and the odds ratio (ratio of two odds) express the effect multiplicatively. A risk difference in percentage points is the additive counterpart, and it is usually the one a decision needs. **Variance-explained measures.** Statistics such as eta-squared after an analysis of variance express the effect as the share of outcome variation attributable to a factor. ## Why the magnitude is the thing that travels Three practical reasons. **Decisions are cost-benefit comparisons.** Acting on a result costs something: engineering time, drug side effects, a permanently more complex product. Whether the benefit exceeds that cost is a comparison of magnitudes. A verdict has no magnitude, so it cannot enter the comparison. The disciplined version of this is to state, *before* looking at the data, the smallest effect that would change the decision — then compare the estimate and its interval to that threshold. **Estimates accumulate; verdicts do not.** Combining several studies of the same question means averaging their estimated magnitudes, weighted by precision. A collection of yes/no verdicts cannot be averaged into anything meaningful. This is why standardized effect sizes exist at all. **Effect sizes expose over-claiming.** Reporting the magnitude next to the verdict makes it immediately visible when a headline claim rests on a trivial difference. It is the cheapest available defence against a write-up that says "highly significant improvement" about a change no user could notice. ## Reporting the pair properly A complete result has three parts: the estimate, an interval around it, and the units or standardization used. The interval matters as much as the point estimate, because it distinguishes the two ways a small number can arise — a genuinely tiny effect measured precisely, versus an unknown effect measured badly. Those two look identical if only the point estimate is shown, and they call for opposite next steps: stop, versus collect more data. Finally, a note on direction of inference. A non-significant test is not evidence that the effect is zero; it is a failure to distinguish the effect from zero, and the interval will usually still contain values a decision-maker would care about. Reporting the effect size and its interval makes that distinction legible instead of leaving readers to infer "no effect" from the absence of a star.

  • What is the difference between a raw and a standardized effect size?
    A raw effect size is in the outcome's own units — 3 points, 0.4 mmHg, 2 percentage points — so a domain expert can price it directly. A standardized one divides that difference by a standard deviation, making it unitless and comparable across studies and instruments, at the cost of hiding what the units were. Report the raw number for decisions and the standardized one for comparison across measures.
  • Can a result be practically important yet fail to reach significance?
    Yes. With a small or noisy sample the interval around a substantial estimate can still straddle zero. That is a statement about precision, not about the effect being absent — the honest reading is that the data cannot distinguish a meaningful effect from none, so the next step is more data, not a conclusion of no effect.
  • How do you choose which effect size measure to report?
    Match the outcome type and the audience. Continuous outcomes take a mean difference, standardized if you need cross-study comparability. Binary outcomes take a risk difference for decisions and a ratio measure for describing relative change. Multi-group factor designs take a variance-explained measure. When in doubt, report the version stated in the units the decision-maker already thinks in.

saying these in an interview costs you the question

  • Says a very small p-value means a large effect
  • Reports significance with no estimate of magnitude
  • Ranks findings by p-value instead of by effect size
  • Calls a non-significant result proof of no effect
  • Gives a point estimate with no interval around it

context