With r = 0.6 between two exams, what score do you predict for a student who scored 2 SD above the mean?
answer
- shrink the deviation, do not copy it
- work in standard-deviation units
- multiply the z-score by the correlation
- 0.6 applied to two standard deviations
basics
~20 sAbout 1.2 standard deviations above the mean. In standard-deviation units the expected second score is the correlation times the first score, so 0.6 times 2 gives 1.2. The prediction is pulled toward the mean because the correlation is below 1.
solid answer
~50 sIn standard-deviation units the best prediction of the second score is `r` times the first, so `0.6 * 2 = 1.2` standard deviations above the mean. That shrinkage is regression to the mean made quantitative: the correlation tells you what fraction of an observed deviation is repeatable, and the remainder was occasion-specific noise that gets redrawn. Two caveats matter in an interview. First, `1.2` is a conditional average over all students who scored `2` SD above, not a forecast for one student — individual second scores scatter around it with residual spread `sqrt(1 - r^2) = 0.8` SD. Second, the rule is symmetric in sign and in direction: a student `2` SD below is predicted at `1.2` SD below, and a student `2` SD above on the *second* exam is predicted to have been about `1.2` SD above on the first.
go deeper
Know the shape of the answer before the arithmetic: the prediction stays above average but moves toward the mean. Being able to say 'less than 2 SD, still positive' already beats guessing 2 SD again.
Do the calculation confidently in standard-deviation units and justify why the factor is the correlation rather than anything else. Say explicitly that it is an expectation for a group, not a prediction for one student.
Turn the formula into an evaluation tool: estimate the metric's period-to-period correlation, compute the rebound it predicts for a selected group, and present that as the null a claimed effect must clear.
Frame the reliability question. Decide which metrics are stable enough to rank units on directly and which need shrinking first, and make that a standing convention rather than a per-analysis argument.
## The rule Standardise both measurements — convert each value to a z-score, meaning how many standard deviations it sits above or below its own mean. If the two standardised measures have correlation `r`, then the best linear prediction of the second given the first is ``` z2_predicted = r * z1 ``` With `r = 0.6` and `z1 = 2.0`, the answer is `1.2`. The student is still predicted to be well above average — the exams share real signal — but only 60% as far above as the first result suggested. In original units the same rule reads ``` y_predicted = mean_y + r * (sd_y / sd_x) * (x - mean_x) ``` which is just the z-score rule with the units put back. If the two exams are on the same scale with the same spread, `sd_y / sd_x = 1` and the shrinkage factor is `r` alone. ## Why the factor is exactly r Among all predictions of the form `a + b*z1`, the one minimising expected squared error has slope equal to `Cov(z1, z2) / Var(z1)`. Because both variables are standardised, `Var(z1) = 1` and `Cov(z1, z2) = r`, so the slope is `r` and the intercept is `0`. If the two measures are jointly normal this is not merely the best *linear* prediction but the exact conditional mean; without normality it remains the best linear one, which is what interviewers are after. A useful way to feel the result: write each observed score as a stable component plus an independent transient component. The correlation between two such measurements equals the share of total variance that is stable. Multiplying the observed deviation by `r` is therefore an estimate of how much of that deviation was the repeatable part. Everything else was occasion-specific and has expectation zero next time. ## The residual spread `1.2` is a centre of mass, not a destiny. For standardised variables the variance of the second score around the prediction is `1 - r^2`, so the residual standard deviation is `sqrt(1 - r^2)`. At `r = 0.6` that is `sqrt(0.64) = 0.8` SD. So students who scored `2` SD above will be spread around `1.2` SD with a typical deviation of `0.8` SD; a good number of them will beat their first result, and a good number will land near the mean. Regression to the mean is a claim about the average of the selected group, and a candidate who says 'the student will score lower' has overstated it. ## How the pull scales with r | r | predicted z2 from z1 = 2 | interpretation | |---|---|---| | 0.95 | 1.90 | highly repeatable measure, regression nearly invisible | | 0.60 | 1.20 | typical for two exams on related material | | 0.30 | 0.60 | noisy measure, most of the deviation was occasion-specific | | 0.00 | 0.00 | nothing repeatable; the mean is the best guess | This table is the practical payoff of the formula. Before you evaluate anything measured on a selected extreme group, estimate the period-to-period correlation of the metric among untreated units, then compute `r` times the group's selected deviation. That number is the rebound you should expect with no intervention at all, and it is the honest null that any claimed effect has to beat. ## Reliability and the same-instrument case When the two measurements are parallel forms of the same instrument, `r` is what measurement theory calls the reliability of the instrument. The rule then reads: the expected true-score deviation is reliability times the observed deviation. Ranking units on a single noisy reading and treating the ranking as truth systematically overstates the top and understates the bottom, which is why estimates for small units — a small school, a low-volume junction, a rep with few deals — are usually shrunk toward the overall mean before being compared. Regression to the mean is the reason shrinkage estimators exist, not an inconvenience they work around. ## Common errors Predicting `2` SD again ignores the noise entirely. Predicting `0` treats the measure as pure noise and throws away real signal. Applying `r` directly to raw units without standardising is wrong whenever the two measures have different spreads. Using `r^2` instead of `r` confuses the shrinkage factor with a variance share and shrinks far too aggressively: at `r = 0.6` it would give `0.72` rather than `1.2`. And forgetting the symmetry leads candidates to describe the effect as something that happens *to* the student over time, rather than as a property of the pair of measurements.
- What does the same rule predict for a student who scored 2 SD below the mean?1.2 SD below the mean. The rule is symmetric in sign: the predicted second z-score is r times the first whatever its direction. Extremely low first scores are also partly bad luck, so on average they move up toward the centre by the same proportion. The pull is always toward the mean and never past it.
- Is 1.2 SD what an individual student will actually score?No — it is a conditional average across all students who scored 2 SD above. Individual second scores scatter around it with residual standard deviation sqrt(1 - r^2), which at r = 0.6 is 0.8 SD. Plenty of those students will score above 1.2 and some will beat their original result. The prediction is a centre of mass, not a forecast for one person.
- How does the answer change if the correlation were 0.95 instead?The prediction becomes 1.9 SD above the mean, so the regression is almost invisible. A high correlation means the measurement is mostly repeatable signal with little occasion-specific noise. This is why precise instruments and metrics aggregated over many events regress only slightly, while single noisy readings on small units regress a great deal.
saying these in an interview costs you the question
- Predicts the same 2 SD on the second exam
- Says the student will definitely score lower
- Applies the correlation to raw units without standardising
- Uses the squared correlation as the shrinkage factor
- Thinks the score keeps shrinking on every further retake