Pearson's r = 0.9: why does that number not tell you how big the effect is in real units?
answer
- both spreads were divided out
- tight around a line, not steep
- the units disappeared in the formula
- same r, any slope you want
- slope equals r times sy over sx
basics
~20 sPearson's r is standardised, so it reports how tightly points hug a straight line, not how steep that line is. Dividing by both spreads removes the units, leaving r free to accompany any slope at all.
solid answer
~40 sr is covariance divided by both standard deviations, which strips out the units and every scale. What survives is tightness around a straight line: r = 0.9 says the points sit close to some line, not that the line is steep. Steepness in original units depends on the spreads as well, since the change in y per one-unit change in x works out to `r * (sy / sx)`. Hold r at 0.9 and vary sy or sx and you get any slope you like, so two studies both reporting 0.9 can imply wildly different real-unit effects. The practical rule is that r answers "how noisy is this relationship" while a decision needs "how much does y move per unit of x", with units and an uncertainty range attached.
go deeper
Recall that Pearson's r is unitless and bounded between -1 and 1, so it cannot express how much one variable changes when the other moves. Know that it describes scatter around a line.
Explain the mechanics: dividing the covariance by both standard deviations removes the scale, and the change in y per unit of x is r times sy over sx. Show that fixing r leaves the slope completely free.
Show how you would report to a decision-maker: the change in units with an uncertainty range as the headline, r beside it as a noise summary. Be ready to explain why two studies with equal r can imply very different practical effects.
Own what counts as an effect in the organisation's reporting. Decide when a standardised association is the right currency and when only units-and-uncertainty will do, and stop screening pipelines from ranking work by correlation strength alone.
## Standardisation removes exactly the information an effect size needs Pearson's r is defined as `r = cov(x, y) / (sx * sy)` Dividing by both sample standard deviations is what makes r comparable across variable pairs -- and it is also what deletes the scale. Every unit, every spread, every notion of "how much" is divided out. What is left is a pure geometric statement: on the standardised picture, where both axes have been rescaled to unit spread, how tightly do the points cluster around a straight line? r = 0.9 says very tightly. It says nothing whatsoever about how much y moves when x moves by one real unit. ## Tightness versus steepness The two ideas are independent, and separating them is the whole answer: - **Tightness** is scatter around the line: how much of y's variation travels with x. That is what r measures. - **Steepness** is the change in y per unit of x, in the units the variables were measured in. That is the effect size a decision-maker needs. The bridge between them uses the spreads: `change in y per unit of x = r * (sy / sx)` So the same r = 0.9 produces a tiny slope when y barely varies relative to x, and a huge one when y varies a lot. Concretely: a training programme where scores are tightly linked to hours but the whole score range spans two points, and one where they are equally tightly linked across a forty-point range, both report r = 0.9 while one is worth funding and the other is not. Running the identity the other way is just as instructive. A relationship can be enormously important and still report a modest r, because r penalises noise: if y moves a great deal per unit of x but there is also a lot of unrelated variation in y, the line is steep and the cloud is fat, giving a big effect with a small r. Strength and size are simply different questions. ## Why this bites in practice Stakeholders hear "0.9" as "nearly perfect", and the number invites a percentage reading it does not support. Three habits keep the conversation honest: 1. **Report the effect in units.** "Each additional hour is associated with about 1.4 more points" is actionable; "r = 0.9" is not. Attach an uncertainty range so the reader sees precision as well as size. 2. **Report r alongside, as a noise statistic.** It answers a different and still useful question: how much of the variation travels together, and therefore how well individual cases will track the pattern. 3. **Never compare r across samples as if it were an effect.** Because `r = slope * (sx / sy)`, two samples of the same underlying phenomenon can report different r values purely because one sample happened to cover a wider range of x or was measured with a noisier instrument. Comparing r values from different studies compares their sampling and measurement conditions as much as the relationship. ## Where r is the right number None of this makes r a bad statistic; it makes it a specific one. r is the right tool when you want a unitless, bounded, comparable summary: screening many variable pairs, checking whether two measurements of the same thing agree in their ordering of cases, or reporting how much co-movement exists without committing to units. It is the wrong tool the moment someone asks "so how much would it help?" ## A pair of sanity checks Before quoting r as if it were an effect, ask two questions. First, if I doubled the units of y -- reported grams instead of kilograms -- would my sentence change? If it would not, but the decision it supports would, you are quoting the wrong statistic. Second, can I state the relationship as "one more unit of x goes with about this much more y"? If you cannot, you have measured tightness and not magnitude. ## In an interview Say the sentence directly: r measures how tightly the points hug a line, not how steep the line is, because standardising by both spreads divides the scale out. Then give the bridge `slope = r * (sy / sx)`, point out that fixing r at 0.9 leaves the slope free, and say what you would actually report to a decision-maker -- the change in y per unit of x, in units, with an uncertainty range, and r beside it as the noise summary.
- What extra quantities do you need alongside r to state the effect in real units?The two sample standard deviations, since the change in y per unit of x is `r * (sy / sx)`. In practice you report that slope with its units and an uncertainty range, plus the means so a reader knows where on the scale the relationship was observed. r then sits beside it as the noise summary rather than standing in for the magnitude.
- Can a relationship have a large effect in real units but a small r?Yes, and it is common. If y moves substantially per unit of x but also carries a lot of unrelated variation, the line is steep while the cloud around it is fat, so r comes out modest. Small r means noisy, not unimportant. Dismissing a relationship on r alone can throw away a large effect measured under noisy conditions.
- Why is comparing r values across two different studies risky?Because `r = slope * (sx / sy)`, so r depends on how much the sampled x values varied and how noisy y's measurement was, not only on the underlying relationship. A study that happened to sample a wider spread of x, or measured y more precisely, reports a higher r for the same phenomenon. Slopes in shared units are far more comparable across studies.
r is like being told two hikers stayed exactly on the marked trail. It tells you how little they wandered, not whether the trail climbed ten metres or a thousand.
saying these in an interview costs you the question
- Reads r = 0.9 as a 90 percent effect or 90 percent agreement
- Assumes a high r implies a steep relationship
- Quotes r to stakeholders as the size of a benefit
- Compares r across studies as if it were an effect size
- Dismisses a relationship as unimportant because r is small