skip to content

How do you compute Cohen's d for a 3-point mean gap when the pooled SD is 15?

level: middleimportance: must knowfreq 74%

answer

  1. distance measured in spreads
  2. divide the gap by the spread
  3. pooled, not one group's SD
  4. not the standard error in the denominator
  5. 0.2 / 0.5 / 0.8 conventions

basics

~20 s

Divide the difference in means by the pooled standard deviation: 3 / 15 = 0.2. The two groups sit one fifth of a standard deviation apart, which Cohen's conventional benchmarks call a small effect, and their distributions overlap heavily.

solid answer

~40 s

Cohen's d is the difference between two group means expressed in standard deviations: `d = (mean1 - mean2) / s_pooled`. Here that is `3 / 15 = 0.2`. The pooled standard deviation combines both samples, `s_pooled = sqrt(((n1-1)*s1^2 + (n2-1)*s2^2) / (n1+n2-2))`, so it is a variance-weighted average of the two group SDs rather than either one alone. Because the units cancel, d is comparable across outcomes measured on different scales. Cohen's rough benchmarks are 0.2 small, 0.5 medium, 0.8 large — conventions he offered as a fallback, not results derived from theory, so a field with well-measured outcomes may treat 0.2 as substantial and another may not. With small samples d is biased slightly upward; Hedges' g applies a correction that shrinks it.

go deeper

for a junior

Be ready to write the formula, plug numbers into it, and say that d is measured in standard deviations rather than in the outcome's original units.

for a middle

Expect to derive the pooled standard deviation, explain why pooling weights by degrees of freedom, and say what changes if the treatment alters the spread as well as the centre.

for a senior

Show judgment about the benchmarks: compare a d against the distribution of effects previously observed on that same outcome, report an interval around it, and know when to apply a small-sample correction.

for a principal

Own the standard for how magnitudes are reported and compared across teams — when a standardized measure is the right currency, when raw units serve the decision better, and how to stop generic benchmarks becoming policy.

## The definition Cohen's d answers "how far apart are these two groups, measured in units of their own spread?" It is defined as `d = (mean1 - mean2) / s_pooled` With a 3-point gap and a pooled standard deviation of 15, `d = 3 / 15 = 0.2`. The groups' centres sit one fifth of a standard deviation apart. The numerator is the raw effect in the outcome's own units. The denominator converts it into a unitless quantity. That is the whole idea: dividing by spread makes the number comparable across instruments that were never on the same scale. ## The pooled standard deviation The usual denominator combines both samples: `s_pooled = sqrt( ((n1 - 1) * s1^2 + (n2 - 1) * s2^2) / (n1 + n2 - 2) )` Each group's variance is weighted by its degrees of freedom, the weighted variances are averaged, and the square root is taken. Two consequences worth stating in an interview: - It is a pooled *variance* average, not an average of the two standard deviations. Averaging SDs directly gives a different, wrong number when the groups differ in size or spread. - Pooling presumes the two groups have roughly similar variance. If a treatment changes the spread as well as the centre — a common outcome when an intervention helps some people a great deal and others not at all — pooling blends two genuinely different spreads. Glass's delta sidesteps this by dividing by the control group's SD only, so the denominator is a property of the untreated population and is not disturbed by the treatment. A frequent error is dividing by the **standard error** instead of the standard deviation. The standard error shrinks as the sample grows, so that quantity grows without bound with more data — it is essentially the test statistic, not an effect size. An effect size must not systematically change when you collect more observations; that invariance is what makes it an estimate of a population quantity. ## Interpreting 0.2 Cohen proposed benchmarks for researchers with no field-specific reference point: **0.2 small, 0.5 medium, 0.8 large**. He was explicit that these were conventions offered for want of anything better, not thresholds derived from theory. A more concrete reading of d = 0.2 is distributional. If both groups are roughly normal with equal spread, a d of 0.2 means the distributions overlap substantially: a randomly chosen member of the higher group exceeds a randomly chosen member of the lower group only slightly more often than half the time. About 56% of the time when d = 0.2 — well short of anything a casual observer would notice in an individual case. But "small" is not "unimportant". Two contexts flip the reading: - **Scale.** A d of 0.2 applied to a very large population, on a cheap intervention, can be worth a great deal in aggregate even though no single person notices it. - **Field norms.** Where outcomes are noisy and hard to move, 0.2 may be near the ceiling of what any intervention achieves; where outcomes are precisely measured and strongly manipulable, 0.5 may be unremarkable. Always benchmark against the distribution of effects previously observed on the same outcome, not against the generic conventions. ## Sign, direction and small-sample bias The sign of d follows the order of subtraction, so state which group is which; report the absolute magnitude plus an explicit direction rather than leaving a bare negative number to be misread. Because the pooled SD is itself estimated, d is biased slightly away from zero in small samples — it tends to overstate the population standardized difference. **Hedges' g** multiplies d by a correction factor slightly below one, with the shrinkage larger when the combined sample is small and negligible for large samples. With a few dozen observations per group the correction is worth applying; with thousands it changes nothing you would report. Finally, d is an estimate like any other, so it deserves an interval. A d of 0.2 whose interval runs from 0.15 to 0.25 is a settled small effect; a d of 0.2 whose interval runs from -0.3 to 0.7 is an unsettled question, and the two demand opposite next steps even though the point estimates are identical.

  • Why divide by the pooled standard deviation rather than one group's?
    Pooling uses the information in both samples, giving a more stable denominator, and it treats the two groups symmetrically. It assumes their variances are comparable. When the treatment plausibly changes the spread, that assumption breaks, and Glass's delta — dividing by the control group's SD alone — gives a denominator the treatment cannot distort.
  • Is a Cohen's d of 0.2 always too small to act on?
    No. The benchmarks are generic conventions, not decision rules. A 0.2 standardized effect on a cheap change reaching millions of people can be very valuable, while the same 0.2 on an expensive, risky intervention is not worth the cost. Judge it against the distribution of effects previously seen on that outcome and against the cost of acting.
  • What goes wrong if you divide the mean difference by the standard error?
    You get a test-statistic-like quantity, not an effect size. The standard error shrinks roughly as one over the square root of the sample size, so that ratio grows as you collect more data even though the underlying difference is unchanged. An effect size must be stable under added observations; only the standard deviation in the denominator gives that.

saying these in an interview costs you the question

  • Divides by the standard error instead of the standard deviation
  • Reads d as a percentage difference between the groups
  • Averages the two group SDs instead of pooling variances
  • Calls d = 0.2 negligible regardless of scale or cost
  • Believes Cohen's benchmarks are derived from theory

context