skip to content

For x = 1..10 and y = e^x, why is Spearman's rho exactly 1 while Pearson's r is only about 0.72?

level: middleimportance: should knowfreq 52%

answer

  1. same data, two different questions
  2. ordering versus straight-line fit
  3. rank differences are all zero here
  4. curvature, not noise, costs the raw-value coefficient
  5. log the y column and the line appears

basics

~20 s

The relationship is perfectly monotone but strongly curved. Spearman's rho sees only the ordering, which matches exactly, so it hits 1. Pearson's r measures closeness to a straight line, and an exponential curve is far from straight.

solid answer

~50 s

Rank the ten x values and the ten y values: because the exponential is strictly increasing, both rank sequences are 1 through 10, identical. Spearman's rho correlates those rank columns, gets a sequence perfectly correlated with itself, and returns exactly 1. Pearson's r asks a different question — how tightly do the points hug a single straight line — and `e^10` is roughly 22,000 while `e^1` is under 3, so the cloud bends sharply upward and the best straight line leaves large residuals, giving about 0.72. Neither number is wrong; they answer different questions. The lesson is that r below 1 does not mean 'weak relationship' — here the relationship is deterministic — it means 'not straight'. Log the y column and Pearson's r becomes 1 too, while rho does not budge, because logging preserves ranks.

go deeper

for a junior

Remember the one-line contrast: rank correlation measures whether the ordering is consistent, and a straight-line correlation measures whether the points fall near a line. A curve can be perfect on the first and imperfect on the second.

for a middle

Explain the mechanism: the exponential is strictly increasing so every rank difference is zero, while the huge spacing at the top end bows the point cloud away from any single line.

for a senior

Treat the gap between the two coefficients as a diagnostic you act on — inspect the scatterplot, consider whether a monotone transform yields an interpretable slope, and never resolve it by reporting the larger number.

for a principal

Set the norm that curvature and noise are different deficits even though one coefficient conflates them, so reports that quote a single correlation without the plot are incomplete by policy, not by taste.

## Two different questions The pair of numbers looks contradictory only if you assume both coefficients measure 'strength of relationship'. They do not. - **Spearman's rho** asks: *is the relationship order-preserving?* It converts each column to ranks and correlates the ranks. - **Pearson's r** asks: *how close are the points to one straight line?* It works on the raw values and their distances from the means. For `y = e^x` on `x = 1, 2, ..., 10`, the answer to the first question is a perfect yes and the answer to the second is 'not very'. ## Why rho is exactly 1 The exponential function is strictly increasing: whenever `x_i > x_j`, it follows that `e^(x_i) > e^(x_j)`. So the ordering of the y column is the ordering of the x column, observation for observation. Ranking gives `1, 2, 3, ..., 10` in both columns, every rank difference `d` is 0, and the coefficient is 1 — you can see it directly from the no-ties shortcut `1 - 6*sum(d^2)/(n*(n^2 - 1))` with `sum(d^2) = 0`. This is not special to the exponential. **Any** strictly increasing function of x — a logarithm, a cube, a step function with no repeats, an arbitrary hand-drawn rising squiggle — gives rho exactly 1. Strictly decreasing functions give exactly -1. Rank correlation is blind to the shape of the curve and sees only its direction. ## Why Pearson's r falls short of 1 The y values here are `e^1 = 2.72` up to `e^10 = 22,026`. The last two points, `e^9` and `e^10`, are separated by more than 12,000 while the first eight points are all crammed below 3,000. No single straight line can pass near all ten points: fit a line and the middle points sit above it at one end and below at the other, in the classic bowed residual pattern. Pearson's r on this configuration is about 0.72. That is genuinely high in the abstract, but it reflects the fact that a rising cloud is roughly line-shaped in the crudest sense, not that the exponential is 72 percent of a line. The crucial reading: **r = 0.72 here does not mean the relationship is noisy or partly random.** The relationship is deterministic and noiseless. Every bit of the shortfall from 1 is curvature, not scatter. A coefficient that mixes those two very different causes into one number is exactly why 'r is only 0.72, so the effect is moderate' is such a common misreading. ## The transform test Apply a monotone transform and watch what happens: - Take logs of the y column. Now `log(y) = x` exactly, the points sit on a perfect line, and Pearson's r becomes 1. The data did not change; the coordinate system did. - Spearman's rho, meanwhile, is still exactly 1, because logging preserved every rank. It was 1 before and it is 1 after. That asymmetry is the whole story. Pearson's r is a property of the data *in the units you chose*; rank correlation is a property of the ordering, which no increasing re-expression can disturb. If your units are arbitrary — an index, a score, a currency, a raw count that could equally be reported as a rate — a coefficient that changes when you re-express them is reporting partly on your choice of scale. ## What to do with the diagnosis A large gap between a rank coefficient and a raw-value coefficient on the same data is a **signal, not a nuisance**. It says: strong consistent direction, poorly captured by a straight line. The follow-up is not to pick the bigger number and report it. It is to look at the scatterplot and decide what the curvature means: - If a transform straightens it (logs for exponential growth, a square root for count-like data), the transform is often the honest model, and the straightened relationship is easier to communicate. - If nothing straightens it, say so, and report the monotone measure while being clear that no constant per-unit slope exists. - If the direction reverses somewhere — a genuine non-monotone shape — then *both* coefficients understate the structure, and neither is the summary you want. ## The trap in the other direction Be careful not to over-claim for rank correlation either. Rho of 1 says the ordering is perfectly preserved; it says nothing about how big the effect is. Two variables that agree perfectly in order but where y barely moves across the whole range of x still give rho = 1. Monotone perfection is not practical importance, and the coefficient carries no units to tell you which you have.

  • What does a large gap between a rank coefficient and a raw-value correlation tell you to do next?
    Plot the data. The gap says the relationship has a consistent direction that a straight line captures poorly, which is a diagnostic, not a tie-break. Decide whether a monotone transform straightens the pattern into something with an interpretable per-unit slope, or whether the honest summary is simply that the association is monotone but not linear.
  • Does Spearman's rho of 1 mean the effect is large?
    No. Rho of 1 says the two orderings match perfectly and nothing more. Two variables can agree perfectly in order while y moves only a hair across the entire range of x, and rho is still exactly 1. Rank coefficients discard magnitudes, so practical importance has to come from somewhere else entirely.
  • Would the same rho of 1 appear for a strictly decreasing curve?
    The magnitude would, with the sign flipped: any strictly decreasing relationship gives rho of exactly -1, because one rank column runs 1 to n while the other runs n to 1. Rank correlation reads direction and consistency of ordering, so a mirror-image curve is just as perfect a monotone relationship.

saying these in an interview costs you the question

  • Reads r of 0.72 on noiseless data as a moderate or weak effect
  • Calls one of the two coefficients simply wrong
  • Assumes rho of 1 implies a large practical effect
  • Thinks curvature and random scatter lower a correlation for the same reason
  • Picks whichever coefficient is larger and reports only that

context