skip to content

Why is Pearson's r zero when y equals x squared over a range symmetric about zero?

level: middleimportance: should knowfreq 52%

answer

  1. linear association only
  2. deviation products cancel by symmetry
  3. odd moments vanish for symmetric x
  4. zero r does not imply independence
  5. converse holds only for joint normality

basics

~20 s

Pearson's r measures straight-line association only. On a parabola centred at zero, the rising and falling halves contribute deviation products of opposite sign that cancel exactly, so r is zero even though y is perfectly determined by x.

solid answer

~50 s

Pearson's r is covariance divided by the two standard deviations, and covariance averages `(x - mean_x) * (y - mean_y)`. Take x spread symmetrically around zero and set `y = x^2`. Every large positive x has a matching negative x with the same y, so their deviation products are equal in size and opposite in sign and the sum cancels to zero. So `r = 0` for a perfectly deterministic relationship: y is completely determined by x. This is the standard proof that zero correlation does not mean independence; the implication runs one way only, since independence does force zero correlation. Points around a circle behave the same. Rank coefficients do not rescue this case either, because the relationship is not monotone - y falls, then rises. The remedy is to plot it, or use a dependence measure not restricted to straight lines.

go deeper

for a junior

Be ready to state that the coefficient captures straight-line association only, and that a perfectly predictable U-shaped relationship can still return zero.

for a middle

Explain the cancellation: with x symmetric about its mean, each positive deviation product is matched by an equal negative one, so covariance is exactly zero.

for a senior

Demonstrate the operational consequence — screening features by correlation quietly discards inverted-U signals — and name the binned or flexible-fit checks you would run instead.

for a principal

Own the standard for feature screening and metric review, so that a single association number never becomes the sole gate through which candidate signals must pass.

## What r is built to detect The sample correlation is `r = sum((x_i - xbar) * (y_i - ybar)) / sqrt(sum((x_i - xbar)^2) * sum((y_i - ybar)^2))` The numerator is a sum of **deviation products**. A pair contributes a positive term when both coordinates are on the same side of their own mean, and a negative term when they are on opposite sides. So `r` answers: do points tend to sit in the upper-right and lower-left quadrants around the means (positive), or the upper-left and lower-right (negative)? That is a question about a straight-line tendency and nothing else. ## The parabola Let x take values spread symmetrically about zero — say -3, -2, -1, 0, 1, 2, 3 — and let `y = x^2`, so y is 9, 4, 1, 0, 1, 4, 9. The mean of x is 0, so the x-deviation of each point is just x itself. The mean of y is 4. Now pair up the symmetric points. At `x = 3` the deviation product is `3 * (9 - 4) = 15`. At `x = -3` it is `-3 * (9 - 4) = -15`. They cancel. Same for the pair at plus and minus 2, and the pair at plus and minus 1. The point at zero contributes nothing. The whole numerator sums to zero, so `r = 0` exactly. More generally, `cov(x, x^2) = E[x^3] - E[x] * E[x^2]`, and for any distribution symmetric about zero both `E[x^3]` and `E[x]` are zero, so the covariance vanishes. The result is not an approximation and does not depend on the particular values chosen. Meanwhile the dependence is total: knowing x tells you y exactly, with no error at all. A coefficient of zero is reporting no *linear* association, and it is right — the best straight line through this cloud is flat. It is simply answering a different question from the one the reader had in mind. ## Zero correlation is not independence This example is the standard counterexample to a very common error. The correct logic is one-directional: - If two variables are **independent**, their correlation is zero (assuming the variances exist). - The converse is false. Zero correlation permits arbitrarily strong dependence, as the parabola shows. The one important special case where the converse does hold is the **jointly normal** case: if `(x, y)` follow a bivariate normal distribution, then zero correlation does imply independence. That special case is why the fallacy is so sticky — people internalise it in the normal setting and then carry it everywhere. Note the requirement is joint normality, not merely that each variable is separately normal. ## Other shapes with the same property Points spread evenly around a circle have `r = 0`, yet each point satisfies an exact equation relating x and y. Any relationship symmetric about the centre of the x range behaves this way: a U shape, an inverted U, a sine wave over a whole number of periods. Inverted-U relationships are common in practice — response versus dose, performance versus arousal, engagement versus notification frequency. Scanning a correlation table for such a variable returns a small number and it gets discarded as uninformative when it is one of the strongest signals present. ## Why ranks do not fix it Substituting ranks for values catches relationships that are curved but still consistently increasing or decreasing. This one is not: y falls as x rises up to zero, then rises. A monotone-association measure has the same cancellation problem, so it also returns roughly zero. The failure here is about **shape**, not about **scale**. ## What to do instead First, plot it — every version of this trap is obvious in a scatterplot and invisible in a table. When automation is required, the options are to bin the predictor and compare the outcome mean per bin (a U shows up immediately as a non-monotone profile), to fit a flexible curve and compare its fit against a straight line, or to use a dependence measure designed to detect general, not straight-line, association. And read a reported zero carefully: it means 'no straight-line trend detected in this sample', never 'these variables have nothing to do with each other'.

  • When does zero correlation actually imply independence?
    When the pair is jointly normally distributed. Under bivariate normality, zero correlation does force independence, which is why the implication feels true to people trained on normal examples. The requirement is joint normality of the pair, not each variable being normal on its own — two separately normal variables can be dependent with zero correlation.
  • Would restricting the analysis to positive x change the answer?
    Yes, dramatically. On x from 0 upward the parabola is strictly increasing, so the cancellation disappears and the correlation is strongly positive. That is worth noticing on its own terms: the coefficient depends on the range you compute it over, and a symmetric window can hide a relationship that a one-sided window shows clearly.
  • How would you detect this automatically across many candidate predictors?
    Bin the predictor and compare the outcome mean per bin. A U or inverted-U shows up immediately as a non-monotone bin profile that no single coefficient reports. Alternatively, compare the fit of a flexible curve against a straight line, or use a dependence measure that detects general association rather than straight-line trend.

Asking for the average direction of a car that drives out and back returns zero net displacement. The car certainly moved; the summary just cannot represent an out-and-back trip.

saying these in an interview costs you the question

  • Says zero correlation means the variables are independent
  • Concludes a predictor is useless from a near-zero r
  • Thinks separate normality is enough for the converse
  • Believes rank correlation rescues any curved relationship
  • Confuses no linear trend with no relationship

context