skip to content

Why does SAT score correlate weakly with college GPA when measured only among admitted students?

level: seniorimportance: should knowfreq 44%

answer

  1. selection carved out the sample
  2. predictor spread shrinks, residual spread does not
  3. r-squared is a ratio of variances
  4. the slope largely survives
  5. pooled heterogeneous groups do the reverse

basics

~20 s

Admission selects on the score, so admitted students span a narrow score range. Correlation scales with how much the predictor varies relative to the leftover scatter, so truncating that spread shrinks r even though the underlying relationship is unchanged.

solid answer

~50 s

This is restriction of range. Admission decisions cut off the low end of the score distribution, so among admitted students the predictor barely varies while the scatter in outcomes stays roughly the same. Since `r^2` can be written as explained variation over total variation, and the explained part is the slope squared times the variance of the predictor, shrinking that variance drives r down mechanically. Crucially the relationship itself has not weakened: under selection on the predictor alone, the regression slope of GPA on score is largely preserved, and it is the correlation — a standardised, variance-dependent quantity — that collapses. The practical reading is that a correlation computed in a selected sample describes that sample, not the applicant population, and it systematically understates the predictor's value. The mirror-image trap is range enhancement: pooling heterogeneous groups widens the predictor's spread and inflates r without any relationship having improved.

go deeper

for a junior

Be ready to say that when a sample only covers a narrow slice of the predictor, the correlation computed there will look small even if the relationship is real.

for a middle

Explain the mechanics: r-squared is explained variance over total, the explained part scales with the predictor's variance, and selection truncates exactly that.

for a senior

Show the diagnosis on real work — ask what filter created the analysis population, report the slope in units rather than the coefficient, and state the direction of the bias.

for a principal

Own the evaluation design so selection effects are anticipated: hold out unselected cases where ethics and cost allow, and set the standard for what a selected-sample number is allowed to claim.

## The setup A university admits students partly on a test score, then asks whether the score predicts first-year grade point average. The correlation computed among enrolled students comes back modest — often far lower than expected — and someone concludes the test is close to worthless. The conclusion does not follow, because of how the sample was built. ## Why a narrow predictor range shrinks r Start from the decomposition. If the relationship is `y = a + b * x + e`, with `e` the part of y that x does not explain, then in a sample `r^2 = (b^2 * var(x)) / (b^2 * var(x) + var(e))` Read that expression as a competition between two quantities: how much the predictor moves things around, versus how much everything else does. Selection on the score truncates `var(x)` — the admitted band might span a fraction of the applicant range — while `var(e)`, the influence of motivation, health, course choice, teaching quality and luck, is essentially untouched by admission. The numerator falls, the denominator falls by less, and `r` drops. Nothing about the underlying relationship changed; only the sample's composition did. This is why range restriction is a *sample* property rather than a property of the relationship. Two analysts can compute honest correlations from the same underlying process and report very different numbers because their samples were selected differently. ## The slope survives, the correlation does not The most useful thing to say in an interview is the distinction between the two. Under selection on the predictor alone — everyone above a score cutoff is admitted, nobody below — the conditional relationship of y given x within the retained band is the same relationship it always was, so the regression slope `b` is essentially preserved. Correlation is `b * sd(x) / sd(y)`: it carries `sd(x)` inside it, so it moves with the sample's spread. That is the whole asymmetry. A slope is expressed in real units (GPA points per score point) and answers 'how much does the outcome change'; a correlation is unitless and answers 'how much of this sample's variation is accounted for'. Restriction of range attacks the second question, not the first. The caveat worth stating: this clean result holds for selection on the predictor. If selection depends on the outcome too, or on unmeasured things related to the outcome, more can go wrong than range restriction, and the slope is no longer safe. ## Corrections and their price There are classical corrections that take the restricted correlation and the ratio of restricted to unrestricted predictor spread and estimate what the correlation would have been in the full population. They are standard in personnel-selection work. They require knowing the unrestricted variance of the selection variable — usually available for a test score, since applicants who were rejected still took the test — and they assume the relationship and the residual spread are the same outside the observed band, which is an extrapolation nobody can check. Use them to argue that a modest observed correlation is consistent with a useful predictor, not to manufacture a precise number. ## The same trap, other clothes Restriction of range appears whenever the analysis population is filtered on something related to the predictor: - A hiring test evaluated only on people who were hired because they scored well on it. - A credit model scored only on applicants who were approved, so the risky end of the score distribution has no outcomes at all. - An engagement metric correlated with retention among users who already stayed thirty days. - Salary versus experience within a single job level, where the level itself was assigned by experience. In each case, the correlation is computed in a window carved out by the very variable under study. ## The mirror image The opposite error inflates rather than deflates. Pool two groups that differ in both variables — two schools, two product tiers, two time periods — and the combined scatter stretches along a line joining the group centres, which pushes the correlation up. Neither within-group relationship has to be strong for the pooled number to look impressive. So the discipline runs both ways: before quoting a correlation, ask what defined the sample, and whether that definition narrowed or widened the spread of the variable being credited. ## What to say in an interview Name the mechanism (selection truncated the predictor's variance), state the consequence in the right direction (r understates, slope roughly survives), give the remedy (report the slope in units, use unrestricted variance where available, be explicit that the number describes the selected sample), and note the mirror-image inflation. That sequence shows the judgment rather than the vocabulary.

  • Why is the regression slope less damaged by this than the correlation?
    Because selection on the predictor keeps the conditional relationship of the outcome given the predictor intact within the retained band, so the slope in real units is essentially preserved. Correlation equals the slope times the ratio of predictor spread to outcome spread, so it carries the sample's variance inside it and falls when that variance is truncated.
  • What information do you need to correct a range-restricted correlation?
    The spread of the selection variable in the unrestricted population as well as in the retained sample. For an admissions test that is usually available, because rejected applicants still took the test. The correction also assumes the relationship and the residual spread carry over to the unobserved range, which cannot be verified, so treat the result as an argument rather than a measurement.
  • How can the same mechanism make a correlation look far too strong?
    By widening rather than narrowing the predictor's spread. Pool groups that differ on both variables and the combined cloud stretches along the line joining their centres, inflating the coefficient even when the relationship inside each group is weak. Always ask what defined the sample and whether it compressed or stretched the range.
  • A credit model shows a weak score-versus-default correlation among approved applicants. What do you conclude?
    Almost nothing about the score's quality. Approval was made on the score, so the riskiest band has no outcomes recorded at all and the retained range is narrow. Report the relationship in units over the observed band, be explicit that it describes approved applicants, and treat any judgment about the rejected range as an extrapolation.

Judge a thermometer only on days between 20 and 21 degrees and it will look useless. The instrument is fine; you removed the variation it was built to track.

saying these in an interview costs you the question

  • Concludes the predictor is useless from the selected-sample r
  • Generalises a correlation from a filtered sample to everyone
  • Thinks selection lowers the slope as much as the correlation
  • Never asks how the analysis sample was chosen
  • Applies a range correction without unrestricted variance
  • Misses that pooled groups can inflate r the same way

context