For x = 1..10 and y = e^x, why is Spearman's rho exactly 1 while Pearson's r is only about 0.72?
answer
- same data, two different questions
- ordering versus straight-line fit
- rank differences are all zero here
- curvature, not noise, costs the raw-value coefficient
- log the y column and the line appears
basics
~20 sThe relationship is perfectly monotone but strongly curved. Spearman's rho sees only the ordering, which matches exactly, so it hits 1. Pearson's r measures closeness to a straight line, and an exponential curve is far from straight.
solid answer
~50 sRank the ten x values and the ten y values: because the exponential is strictly increasing, both rank sequences are 1 through 10, identical. Spearman's rho correlates those rank columns, gets a sequence perfectly correlated with itself, and returns exactly 1. Pearson's r asks a different question — how tightly do the points hug a single straight line — and `e^10` is roughly 22,000 while `e^1` is under 3, so the cloud bends sharply upward and the best straight line leaves large residuals, giving about 0.72. Neither number is wrong; they answer different questions. The lesson is that r below 1 does not mean 'weak relationship' — here the relationship is deterministic — it means 'not straight'. Log the y column and Pearson's r becomes 1 too, while rho does not budge, because logging preserves ranks.
go deeper
Remember the one-line contrast: rank correlation measures whether the ordering is consistent, and a straight-line correlation measures whether the points fall near a line. A curve can be perfect on the first and imperfect on the second.
Explain the mechanism: the exponential is strictly increasing so every rank difference is zero, while the huge spacing at the top end bows the point cloud away from any single line.
Treat the gap between the two coefficients as a diagnostic you act on — inspect the scatterplot, consider whether a monotone transform yields an interpretable slope, and never resolve it by reporting the larger number.
Set the norm that curvature and noise are different deficits even though one coefficient conflates them, so reports that quote a single correlation without the plot are incomplete by policy, not by taste.
## Two different questions The pair of numbers looks contradictory only if you assume both coefficients measure 'strength of relationship'. They do not. - **Spearman's rho** asks: *is the relationship order-preserving?* It converts each column to ranks and correlates the ranks. - **Pearson's r** asks: *how close are the points to one straight line?* It works on the raw values and their distances from the means. For `y = e^x` on `x = 1, 2, ..., 10`, the answer to the first question is a perfect yes and the answer to the second is 'not very'. ## Why rho is exactly 1 The exponential function is strictly increasing: whenever `x_i > x_j`, it follows that `e^(x_i) > e^(x_j)`. So the ordering of the y column is the ordering of the x column, observation for observation. Ranking gives `1, 2, 3, ..., 10` in both columns, every rank difference `d` is 0, and the coefficient is 1 — you can see it directly from the no-ties shortcut `1 - 6*sum(d^2)/(n*(n^2 - 1))` with `sum(d^2) = 0`. This is not special to the exponential. **Any** strictly increasing function of x — a logarithm, a cube, a step function with no repeats, an arbitrary hand-drawn rising squiggle — gives rho exactly 1. Strictly decreasing functions give exactly -1. Rank correlation is blind to the shape of the curve and sees only its direction. ## Why Pearson's r falls short of 1 The y values here are `e^1 = 2.72` up to `e^10 = 22,026`. The last two points, `e^9` and `e^10`, are separated by more than 12,000 while the first eight points are all crammed below 3,000. No single straight line can pass near all ten points: fit a line and the middle points sit above it at one end and below at the other, in the classic bowed residual pattern. Pearson's r on this configuration is about 0.72. That is genuinely high in the abstract, but it reflects the fact that a rising cloud is roughly line-shaped in the crudest sense, not that the exponential is 72 percent of a line. The crucial reading: **r = 0.72 here does not mean the relationship is noisy or partly random.** The relationship is deterministic and noiseless. Every bit of the shortfall from 1 is curvature, not scatter. A coefficient that mixes those two very different causes into one number is exactly why 'r is only 0.72, so the effect is moderate' is such a common misreading. ## The transform test Apply a monotone transform and watch what happens: - Take logs of the y column. Now `log(y) = x` exactly, the points sit on a perfect line, and Pearson's r becomes 1. The data did not change; the coordinate system did. - Spearman's rho, meanwhile, is still exactly 1, because logging preserved every rank. It was 1 before and it is 1 after. That asymmetry is the whole story. Pearson's r is a property of the data *in the units you chose*; rank correlation is a property of the ordering, which no increasing re-expression can disturb. If your units are arbitrary — an index, a score, a currency, a raw count that could equally be reported as a rate — a coefficient that changes when you re-express them is reporting partly on your choice of scale. ## What to do with the diagnosis A large gap between a rank coefficient and a raw-value coefficient on the same data is a **signal, not a nuisance**. It says: strong consistent direction, poorly captured by a straight line. The follow-up is not to pick the bigger number and report it. It is to look at the scatterplot and decide what the curvature means: - If a transform straightens it (logs for exponential growth, a square root for count-like data), the transform is often the honest model, and the straightened relationship is easier to communicate. - If nothing straightens it, say so, and report the monotone measure while being clear that no constant per-unit slope exists. - If the direction reverses somewhere — a genuine non-monotone shape — then *both* coefficients understate the structure, and neither is the summary you want. ## The trap in the other direction Be careful not to over-claim for rank correlation either. Rho of 1 says the ordering is perfectly preserved; it says nothing about how big the effect is. Two variables that agree perfectly in order but where y barely moves across the whole range of x still give rho = 1. Monotone perfection is not practical importance, and the coefficient carries no units to tell you which you have.
- What does a large gap between a rank coefficient and a raw-value correlation tell you to do next?Plot the data. The gap says the relationship has a consistent direction that a straight line captures poorly, which is a diagnostic, not a tie-break. Decide whether a monotone transform straightens the pattern into something with an interpretable per-unit slope, or whether the honest summary is simply that the association is monotone but not linear.
- Does Spearman's rho of 1 mean the effect is large?No. Rho of 1 says the two orderings match perfectly and nothing more. Two variables can agree perfectly in order while y moves only a hair across the entire range of x, and rho is still exactly 1. Rank coefficients discard magnitudes, so practical importance has to come from somewhere else entirely.
- Would the same rho of 1 appear for a strictly decreasing curve?The magnitude would, with the sign flipped: any strictly decreasing relationship gives rho of exactly -1, because one rank column runs 1 to n while the other runs n to 1. Rank correlation reads direction and consistency of ordering, so a mirror-image curve is just as perfect a monotone relationship.
saying these in an interview costs you the question
- Reads r of 0.72 on noiseless data as a moderate or weak effect
- Calls one of the two coefficients simply wrong
- Assumes rho of 1 implies a large practical effect
- Thinks curvature and random scatter lower a correlation for the same reason
- Picks whichever coefficient is larger and reports only that