Study hours and exam score correlate at r = 0.8, so what does r-squared = 0.64 mean?
answer
- square the correlation coefficient
- the answer is a share of something
- spread across students, not one student's mark
- the same number both directions
- sixty-four percent of the variance is common
basics
~20 sSquaring Pearson's r gives shared variance: 0.64 means 64 percent of the variation in exam scores moves with variation in study hours, leaving 36 percent unaccounted for. It is a proportion of variance, not of scores.
solid answer
~50 sr-squared is the proportion of variance in one column that is linearly shared with the other. With `r = 0.8`, `r-squared = 0.64`, so 64 percent of the spread in exam scores lines up with the spread in study hours and 36 percent does not. Two properties matter in an interview. First, it is **symmetric**: 64 percent of the variance in study hours is equally shared with exam score, so the number carries no direction. Second, it is a proportion of **variance**, not of score points and not of students -- saying "studying accounts for 64 percent of each student's mark" is wrong. It also rescales intuition: r = 0.8 is not twice as strong as r = 0.4, it shares four times as much variance (0.64 versus 0.16). And because variance is squared distance, the squaring makes middling correlations look much weaker: r = 0.3 shares only 9 percent.
go deeper
Know that squaring the correlation gives a proportion of shared variance and be able to compute it: r = 0.8 gives 0.64, or 64 percent. Recall that the leftover 36 percent is variation the other column does not track.
Explain why the quantity is variance rather than scores, and why it is symmetric in the two columns. Be able to walk the r-to-r-squared table and show how squaring deflates mid-range correlations such as 0.3.
Show judgment about whether 0.64 is high or low for the domain, and name concrete sources of the unshared variance. Interviewers listen for whether you translate the number into what a stakeholder can act on.
Own how association strength is communicated across the organisation. Decide whether reports quote r, r-squared or an effect in real units, and stamp out dashboard copy that turns a shared-variance proportion into a claim about individuals.
## Where the square comes from Pearson's r measures how tightly n paired observations hug a straight line, on a scale from -1 to 1. Squaring it produces a number between 0 and 1 that has a variance interpretation: **r-squared is the fraction of the variance in one variable that is linearly shared with the other**. With study hours and exam score at `r = 0.8`: `r-squared = 0.8 * 0.8 = 0.64` So 64 percent of the variance in exam scores across students is variance that moves in step with hours studied, and 36 percent is variance that does not. ## Variance, not score The single most common error is applying the percentage to the wrong thing. r-squared is a share of **variance** -- the average squared distance of a value from its column mean. It is not: - a share of each student's score ("64 percent of Ana's 78 marks came from studying"), - a share of students ("64 percent of students improved"), - a probability of anything, - a claim that any individual prediction will be 64 percent accurate. Variance is a property of the spread across the whole sample, so the interpretation is inherently about the group, never about one row. ## The symmetry Because `r(x, y) = r(y, x)`, r-squared is symmetric too. The 64 percent describes shared variance between the two columns; swapping which column you call first changes nothing. That is why the honest phrasing is "shared variance" or "variance held in common" rather than any wording that implies one column is doing something to the other. Shared variance is a description of co-movement in this sample, and a third variable moving both columns would produce exactly the same number. ## Why squaring changes how strong things feel The squaring is not cosmetic. It compresses the middle of the correlation scale hard: - r = 0.9 -> 81 percent shared - r = 0.8 -> 64 percent shared - r = 0.7 -> 49 percent shared - r = 0.5 -> 25 percent shared - r = 0.3 -> 9 percent shared - r = 0.1 -> 1 percent shared Two lessons follow. First, differences near the top of the scale matter more than they look: moving from r = 0.7 to r = 0.9 nearly doubles shared variance. Second, correlations that sound respectable in prose are thin in variance terms -- an r of 0.3, often described as "moderate", accounts for 9 percent of the spread, which means 91 percent of what makes students differ is something else. This also settles the "twice as strong" question. r = 0.8 versus r = 0.4 is a factor of two in correlation but a factor of four in shared variance (0.64 versus 0.16). Whenever someone compares two correlations as ratios, square them first. ## The 36 percent The complement, `1 - r-squared = 0.36`, is the share of variance in exam scores that has nothing linear to do with study hours: prior knowledge, sleep, test anxiety, luck on the question set, measurement error in how hours were self-reported. Naming a few plausible sources of that residual variance is a strong interview move, because it shows you read the number as a description of one narrow slice of a messy phenomenon. It is also the reason a high r-squared is not automatically good news and a low one is not automatically bad. In a tightly-controlled physical measurement, 0.64 would be alarming; for self-reported study hours against exam performance in a class of humans, it would be unusually high. ## The linear qualifier Every word above hides the qualifier **linear**. r-squared derived from Pearson's r reports only the variance that co-moves along a straight line. Structure that is real but not straight is simply not counted by this number, so "36 percent unaccounted for" means "unaccounted for by a straight-line relationship in these units", not "unexplainable". ## In an interview Say the sentence precisely: "r-squared = 0.64 means 64 percent of the variance in exam scores is shared with the variance in hours studied." Then add the two guards -- it is variance rather than score, and it is symmetric so it names no direction -- and the scale note that squaring makes mid-range correlations much less impressive than they sound.
- Is a pair with r = 0.8 twice as strongly associated as a pair with r = 0.4?Not in variance terms. Squaring gives 0.64 versus 0.16, so the first pair shares four times as much variance as the second. Correlations compare as ratios only after squaring, which is why analysts quote r-squared when they want to talk about how much of the spread is common. On the raw r scale, the gap of 0.4 says very little on its own.
- What does the remaining 36 percent represent?The share of variance in exam scores that does not co-move linearly with study hours: differences in prior knowledge, sleep, exam anxiety, which questions appeared, and noise in how hours were reported. It is not a measure of error in any single prediction; it is a statement about how much of the spread across students is left once the straight-line co-movement is accounted for.
- Does r-squared = 0.64 tell you which of the two variables came first?No. r-squared is symmetric: the shared variance between hours and score is the same number whichever column you name first. It is a description of co-movement in this sample, and an unmeasured third factor driving both columns would produce the same value. Direction has to come from design or domain knowledge, never from the statistic.
saying these in an interview costs you the question
- Says 64 percent of each student's score comes from studying
- Reports r as the percentage instead of r-squared
- Treats r = 0.8 as twice as strong as r = 0.4
- Reads shared variance as proof one column drives the other
- Calls r-squared a probability or an accuracy rate
- Forgets that only straight-line co-movement is counted