What does Spearman's rank correlation coefficient measure between two variables?
answer
- correlation, but on positions not values
- monotone, not straight-line
- unchanged by any increasing transform
- one wild point only moves one rank
- 1 - 6*sum(d^2) / n(n^2-1), ties excluded
basics
~20 sSpearman's rho is a correlation computed on the ranks of the data instead of the raw values. It measures how well two variables move together in a consistently increasing or decreasing pattern, not just along a straight line.
solid answer
~50 sSpearman's rho replaces each value with its rank inside its own column, then computes an ordinary linear correlation on the two rank columns. Because ranks carry only order, rho answers a weaker question: does y move consistently up (or consistently down) as x rises, by any shape of curve? It runs from -1 to +1, reaches +1 for any strictly increasing relationship and -1 for any strictly decreasing one, and is unchanged by any monotone transform of either variable — taking logs of x moves nothing. That also makes it resistant to a single wild value, which can shift by only a bounded number of rank positions, and to heavy skew. With no tied values you can compute it as `1 - 6*sum(d^2) / (n*(n^2 - 1))`, where `d` is the rank difference for each observation.
go deeper
Be ready to say in one breath that Spearman's rho is a correlation computed on ranks, that it ranges from -1 to +1, and that it captures increasing-or-decreasing patterns rather than straight lines.
Explain the mechanics: how ranks are assigned, why the d-squared shortcut needs no ties, and why any increasing transform of either column leaves the coefficient untouched.
Show judgment about when rho is the right summary — skewed marginals, ordinal inputs, arbitrary units — and be explicit that a value near zero rules out monotone trend only, never dependence.
Own the reporting standard: when a team defaults to rank correlation across dashboards you are trading interpretability in real units for robustness, and that tradeoff should be a stated convention rather than a per-analyst whim.
## The definition Spearman's rank correlation coefficient, written `rho` (or `r_s`), is the Pearson-style linear correlation computed **on ranks rather than on the original numbers**. The recipe is: 1. Sort the x column and give each observation its rank: smallest x gets rank 1, next gets 2, and so on up to n. 2. Do the same, independently, for the y column. 3. Compute the ordinary linear correlation between the two rank columns. That is the whole definition. Everything else people say about rho follows from it. ## A worked mini example Suppose five support accounts have satisfaction scores and ticket counts: | account | satisfaction | rank(x) | tickets | rank(y) | d = rank(x) - rank(y) | |---|---|---|---|---|---| | A | 2 | 1 | 41 | 5 | -4 | | B | 3 | 2 | 30 | 4 | -2 | | C | 5 | 3 | 22 | 3 | 0 | | D | 8 | 4 | 9 | 2 | 2 | | E | 9 | 5 | 4 | 1 | 4 | Here `sum(d^2) = 16 + 4 + 0 + 4 + 16 = 40`, `n = 5`, so `rho = 1 - 6*40 / (5*24) = 1 - 240/120 = -1`. The ranks are perfectly reversed, so rho is exactly -1 even though the raw numbers are nowhere near a straight line. The shortcut formula `1 - 6*sum(d^2)/(n*(n^2 - 1))` is only valid when there are **no ties**; with ties you must average the tied positions into midranks and run a real correlation on those midranks. ## Why the rank step matters **Monotone, not linear.** A correlation on raw values asks how close the cloud is to a straight line. Spearman asks a strictly weaker question: is the relationship *order-preserving*? Any strictly increasing curve — a steep exponential, a flattening logarithm, a staircase — has rho exactly +1, because rank(x) and rank(y) are the identical sequence. **Invariance.** Because ranks depend only on order, rho is unchanged by any strictly increasing transformation applied to either variable: change currency, take logs, square a positive column, switch from seconds to log-seconds. The number does not move at all. That is a genuinely useful property when the units are arbitrary or the scale was chosen by convention. **Robustness.** In the raw-value world one absurd observation can drag a summary a long way. In rank space that same observation can only occupy a rank between 1 and n, so its influence is bounded no matter how extreme it is numerically. The same argument covers heavy right-skew: a long tail compresses into ordinary rank spacing. **Ordinal inputs.** Rho needs only an ordering, so it is legitimate on ordinal data such as 1-to-5 satisfaction ratings, where the distance between 'satisfied' and 'very satisfied' is not meaningfully a fixed quantity. Ordinal scales usually generate many tied values, which is where midranks become mandatory rather than optional. ## What rho does not do **It is not a general dependence detector.** A clean U-shape — y falls, bottoms out, then rises symmetrically — is a strong deterministic relationship with Spearman rho near 0, because the increasing half and the decreasing half cancel. Rho near 0 means *no consistent monotone trend*, not *no relationship*. Plot the data; a single number never replaces the scatterplot. **It is not a proportion of anything.** Rho is a correlation in rank space, so intermediate values like 0.6 have no clean 'share of variation' reading in the original units. Treat it as a monotonicity score: 1 is perfect agreement of orderings, 0 is no consistent direction, -1 is perfect reversal. **It is not scale-free of ties.** Heavy tying shrinks the achievable range and changes how the coefficient behaves, which is why the midrank version, not the d-squared shortcut, is the one to implement. ## When to reach for it Use Spearman's rho when the relationship is expected to be monotone but not straight, when at least one variable is ordinal, when the marginal distributions are badly skewed or contain extreme values, or when the measurement units are arbitrary and you want a summary that survives re-expressing them. Prefer a raw-value correlation when the straight-line relationship *is* the thing you care about — for instance when the slope in real units is the deliverable, since ranks throw the units away. The other standard rank coefficient, Kendall's tau, answers a related question by counting pairs of observations that agree or disagree in ordering. Both are monotone-association measures; they are on different scales and should not be compared to each other numerically.
- Does a Spearman rho near 0 mean the two variables are unrelated?No. It means there is no consistent monotone trend. A symmetric U-shape is a strong, fully deterministic relationship whose rising half and falling half cancel, leaving rho near 0. Rank correlation only detects order-preserving structure, so a near-zero value should send you to the scatterplot, not to a conclusion of independence.
- Why does Spearman's rho stay identical after you log-transform x?Because the logarithm is strictly increasing on positive values, it preserves the ordering of every observation, so every rank in the x column is unchanged. Rho is computed only from ranks, so the input to the calculation is bit-for-bit identical and the output cannot move. The same holds for any strictly increasing transform of either variable.
- Is Spearman's rho appropriate for a 1-to-5 customer satisfaction rating?Yes — rho needs only an ordering, and a 1-to-5 rating supplies one, so it is a sensible choice against something like a ticket count. The practical catch is ties: hundreds of responses spread across five levels produce huge blocks of equal values, so use averaged midranks and a real correlation on them rather than the d-squared shortcut.
It is like judging two race commentators on whether they called the finishing order the same way, ignoring the stopwatch times they quoted.
saying these in an interview costs you the question
- Claims Spearman detects any relationship, including U-shaped ones
- Describes rho as measuring straight-line fit on the raw values
- Uses the d-squared shortcut on data with tied values
- Reads rho as the share of variation explained
- Thinks one extreme value distorts rho as much as a raw-value correlation