skip to content

What is regression to the mean, and when should you expect to see it?

level: juniorimportance: must knowfreq 68%

answer

  1. no cause required
  2. extremes contain luck as well as skill
  3. the second draw of noise is independent
  4. expected next score is r times the first

basics

~20 s

Regression to the mean is the tendency for an extreme measurement to be followed by a less extreme one on remeasurement. Expect it whenever two measurements are correlated but not perfectly, because chance helped produce the extreme.

solid answer

~50 s

Regression to the mean is a statistical artifact: if you select cases because they scored extremely on one noisy measure, their score on a second, imperfectly correlated measure will on average sit closer to the population mean. The reason is that an extreme observed score is usually part stable signal and part lucky or unlucky noise; the signal carries over, the noise is redrawn independently. Put both measures in standard-deviation units and the expected second score is `r` times the first, so any correlation below 1 pulls the prediction toward the mean, and lower `r` pulls harder. It is symmetric rather than a force acting over time — extreme second scores also came from less extreme first scores. Galton named it 'regression toward mediocrity' after observing that unusually tall parents had tall but less tall children.

code

python · 16 lines
python
import random, statistics
random.seed(7)

pairs = []
for _ in range(20000):
    stable = random.gauss(0, 1)             # ability that persists
    first = stable + random.gauss(0, 1)     # noisy measurement 1
    second = stable + random.gauss(0, 1)    # noisy measurement 2
    pairs.append((first, second))

# select the top 10% on the FIRST measurement only
cut = sorted(p[0] for p in pairs)[int(0.9 * len(pairs))]
top = [p for p in pairs if p[0] >= cut]

print(round(statistics.mean(p[0] for p in top), 2))   # about 2.5
print(round(statistics.mean(p[1] for p in top), 2))   # about 1.2

go deeper

for a junior

Be ready to define it in one sentence and give one concrete example, such as a spectacular test score followed by a lower retake. Know that imperfect correlation between the two measurements is the trigger, not any event in between.

for a middle

Explain the mechanics out loud: an extreme observed value is stable signal plus a transient part, the transient part is redrawn independently, and in standard-deviation units the expected second value is r times the first.

for a senior

An interviewer expects you to catch it unprompted in an evaluation someone hands you, put a rough number on how much of the observed change it explains, and propose a comparison group chosen by the same extreme rule.

for a principal

Own the policy angle. Identify which of your organisation's standing evaluation habits select on an extreme, and set measurement conventions so that a predictable rebound can never be booked as impact.

## The phenomenon Regression to the mean is the tendency of an extreme measurement to be followed, on a second imperfectly correlated measurement, by a value nearer the population average. The crucial word is *selected*: the effect appears when cases enter your analysis **because** they were extreme on the first measure. Nothing has to happen to those cases in between. No treatment, no fatigue, no complacency, no jinx — the pattern is generated by the selection rule alone. ## Why it happens: stable part plus transient part Think of any noisy measurement as a sum of two pieces: ``` observed = stable component + transient component ``` The stable component is what the case reliably brings — a pupil's actual command of the material, a junction's underlying danger, an athlete's real ability. The transient component is everything that varies from occasion to occasion: which questions happened to be asked, weather, mood, measurement error, plain luck. Now ask what kind of case ends up at the very top of the observed distribution. Mostly cases with a high stable component, yes — but also, disproportionately, cases whose transient component happened to be positive on that occasion. The top of an observed distribution is *enriched with good luck*, and the bottom is enriched with bad luck. Measure again and the stable component is still there while the luck is drawn afresh, with an expected value of zero. So the group's average moves toward the centre. Same story, mirrored, at the bottom. ## The arithmetic Express each measure in standard-deviation units (a z-score: how many standard deviations above or below its own mean the value sits). If the two measures have correlation `r`, the best prediction of the second z-score given the first is ``` z2_predicted = r * z1 ``` With `r = 0.7` and a first score of `z1 = 2.0`, the expected second score is `1.4` — still well above average, but pulled 30% of the way to the mean. The size of the regression is set entirely by `r`. At `r = 1` there is none: the second measurement is a deterministic relabelling of the first, and nothing is redrawn. At `r = 0` the prediction collapses all the way to the mean. Every real repeated measurement lives between those extremes, so some regression is always present; the only question is how much. This also tells you *when to worry*. Highly reliable measures — a precise instrument, a metric aggregated over a huge number of events, a long-run average — have `r` close to 1 and regress barely at all. Noisy single observations, small counts, short windows and small subgroups have low `r` and regress a lot. ## Galton's original finding Francis Galton studied parents' and adult children's heights in the 1880s and found that unusually tall parents had children who were, on average, taller than the population mean but **shorter than their parents**; unusually short parents had children shorter than average but taller than their parents. He called it 'regression toward mediocrity in hereditary stature'. The name stuck and later gave the statistical technique of regression its name, which is a permanent source of confusion: the phenomenon and the modelling technique share a word by historical accident. ## Two things it is not **It is not a force pulling values back.** There is no mechanism. It is a statement about the conditional average of a group you selected, not about any individual, and it happens in one step, not as a continuing slide. A group selected as extreme regresses once toward the mean and then stays there in expectation; it does not keep sliding on every subsequent remeasurement. **It does not shrink the population.** Galton's data did not show humanity converging on a single height. The distribution of the next generation's heights had roughly the same spread as the previous one. Regression works from both tails at once, and fresh variation enters every period, so the marginal distribution is stable even though every selected extreme group moves inward. It is also not the gambler's fallacy — that is the false belief that independent trials must 'balance out'. Regression makes no claim that the future compensates for the past; it says the extreme you selected was partly luck to begin with. **It is symmetric in direction, not in time.** Cases with extreme *second* measurements also had, on average, less extreme *first* measurements. If the effect were a causal force operating forward in time, that backward version would make no sense. ## Where it shows up Anywhere a rule picks cases at an extreme and then remeasures them: the bottom-decile schools enrolled in an improvement programme, the worst-performing sales representatives assigned a coach, the sickest patients recruited to a trial, junctions with a record-bad accident year, an athlete featured after a spectacular hot streak. Each of those selections will show a rebound the next period whether or not the intervention does anything, which is exactly why an unadjusted before-and-after comparison on a selected extreme group is close to uninformative. The defences follow from the mechanism: remeasure with a comparison group chosen by the *same* extreme rule, use a multi-period baseline rather than the single selection period, and compute in advance how large a rebound `r` predicts so you know what a null result looks like.

  • Does regression to the mean imply that a population's spread shrinks over generations?
    No. The conditional average of a selected extreme group moves inward, but the marginal distribution does not change. Children of tall parents are on average shorter than their parents and children of short parents taller, yet the next generation's height distribution has roughly the same spread, because regression pulls from both tails at once and fresh variation is added each period.
  • When does regression to the mean not occur at all?
    Only at a correlation of exactly 1 — a deterministic relabelling with no independently redrawn component, such as recording the same heights first in centimetres and then in inches. Any correlation below 1 produces some regression, and the smaller it is, the harder the predicted second value is pulled toward the mean. At zero correlation the best prediction is simply the population mean.
  • Why is the Sports Illustrated cover jinx a regression-to-the-mean story?
    Athletes reach the cover after an unusually good stretch, which mixes genuine ability with a hot streak. The streak part is not repeatable, so the following period sits closer to the athlete's own long-run level and reads as a jinx. No causal effect of the cover is needed; the same drop appears for any selection on extreme recent performance.

Picking the tallest-looking people at a party partly picks whoever is wearing the thickest shoes. Measure them barefoot and the group is still tall, just less remarkably so.

saying these in an interview costs you the question

  • Concludes the intervention must have worked
  • Claims populations converge toward the mean over time
  • Describes it as a force pulling values back
  • Assumes it only applies forward in time
  • Confuses it with the gambler's fallacy about luck balancing out

context