skip to content

How do heavy ties on a 1-5 satisfaction scale affect Spearman's rho and Kendall's tau?

level: seniorimportance: should knowfreq 38%

answer

  1. equal values need a ranking rule
  2. average the positions in each tied block
  3. the d-squared shortcut assumes distinct values
  4. tied pairs are neither concordant nor discordant
  5. tau-b shrinks the denominator, tau-a does not

basics

~20 s

Ties break the simple formulas. Spearman's rho needs averaged midranks for each tied block and a full correlation on them. Kendall's tau needs the tie-corrected tau-b, because tau-a cannot reach 1 once many pairs are tied.

solid answer

~50 s

Thousands of responses on a five-point scale produce five enormous blocks of identical values, so most pairs are tied. For Spearman's rho, assign every member of a tied block the average of the ranks that block occupies (a **midrank**) and then compute a real correlation on those midranks — the `1 - 6*sum(d^2)/(n*(n^2-1))` shortcut assumes distinct values and is simply wrong here. For Kendall's tau, tied pairs are neither concordant nor discordant, so `tau_a = (C - D)/(n*(n-1)/2)` keeps them in the denominator and can no longer reach 1 even under perfect agreement. Use **tau-b**, `(C - D) / sqrt((C + D + T_x)*(C + D + T_y))`, where `T_x` and `T_y` count pairs tied on only one variable. Report which variant you used; a tau-a and a tau-b on the same ordinal data are not comparable numbers.

go deeper

for a junior

Know that repeated values need a rule, that the rule is to average the positions a tied block occupies, and that the quick rank-difference formula does not apply once values repeat.

for a middle

Explain both corrections concretely: midranks feeding an ordinary correlation for rho, and the tie-adjusted denominator of tau-b for Kendall, including why tau-a understates on coarse scales.

for a senior

Demonstrate the operational discipline — name the variant in the report, quote the tie load, refuse comparisons across different tie structures, and check what your computation does with equal values.

for a principal

Own the standard for ordinal reporting: fix one tie-corrected variant across the organisation, require the tie load beside every coefficient, and treat variant-shopping for a larger number as a reviewable practice.

## Where the ties come from A 1-to-5 customer satisfaction rating scored against per-account support-ticket counts is the archetypal case. The rating column has exactly five distinct values, so with 2,000 responses the average tied block holds hundreds of rows. The ticket-count column ties heavily too — many accounts filed 0, 1 or 2 tickets. Between them, the great majority of the `n*(n-1)/2` pairs are tied on at least one variable. Any rank coefficient computed as if values were distinct is measuring something else. ## Spearman's rho with ties: midranks Ranking requires a rule for equal values. The convention is the **midrank**: every member of a tied block receives the average of the rank positions the block occupies. If four observations tie for positions 7, 8, 9 and 10, each gets `(7 + 8 + 9 + 10)/4 = 8.5`. The next distinct value resumes at 11, so the ranks still sum to `n*(n+1)/2`. Two consequences follow. **The shortcut formula dies.** `rho = 1 - 6*sum(d^2)/(n*(n^2 - 1))` is an algebraic simplification that holds only when both rank columns are a permutation of `1..n`. With midranks they are not, the simplification no longer follows, and the number it produces is not Spearman's rho. Compute the ordinary linear correlation of the two midrank columns instead — that is the definition, and it stays correct under ties. **The attainable range shrinks.** With five levels on x and, say, forty distinct ticket counts on y, the two midrank columns cannot be perfect images of each other no matter how strong the association: hundreds of x-midranks are literally the same number while the corresponding y-midranks vary. So rho has a ceiling below 1 that is imposed by the tie structure, not by the strength of the relationship. Do not read a rho of 0.55 on a five-point scale as 'about half as strong as it could be'. ## Kendall's tau with ties: tau-a, tau-b, tau-c A pair tied on x, on y, or on both is neither concordant nor discordant. Tau-a still divides by all `n*(n-1)/2` pairs, so those tied pairs contribute nothing to the numerator while inflating the denominator. Under *perfect* agreement between a five-point rating and an outcome, tau-a can land near 0.3 simply because most pairs were tied — a number a reader will misinterpret as weak association. **Tau-b** fixes this by removing ties from the denominator, separately for each variable: `tau_b = (C - D) / sqrt((C + D + T_x) * (C + D + T_y))` where `T_x` counts pairs tied on x but not on y, and `T_y` counts pairs tied on y but not on x. Pairs tied on both variables are excluded from both factors. With no ties, `T_x = T_y = 0` and tau-b reduces to tau-a, so the two agree exactly when ties are absent. Tau-b attains ±1 only when the tie structures of the two variables line up — informally, when the cross-tabulation is square and all the mass sits on a diagonal. When one variable has five levels and the other has forty, tau-b still has a ceiling below 1. **Tau-c** exists for exactly this rectangular case, rescaling by the smaller number of levels so the coefficient can approach ±1 on a lopsided table. ## What this means in practice 1. **State the variant.** 'Kendall's tau = 0.31' is ambiguous on tied data; 'tau-b = 0.31' is not. The same applies to whether rho was computed on midranks. 2. **Never compare across variants or across tie structures.** A tau-b from a five-point scale and a tau-b from a continuous measurement are not on a common footing, because the ceilings differ. 3. **Report the tie load.** The share of pairs tied on at least one variable is a one-line diagnostic that tells the reader how much headroom the coefficient had. 4. **Consider whether the coefficient is the right summary at all.** With only five levels on one side, a cross-tabulation of the five rating levels against ticket-count bands often communicates more than any single number — you can see whether the monotone trend is uniform or driven entirely by the bottom level. 5. **Watch the tie-breaking rule in whatever computes it.** Assigning tied values sequential ranks in input order rather than midranks silently manufactures ordering information that is not in the data, and it makes the result depend on row order. ## The underlying principle Rank coefficients trade magnitude information for robustness. Ties are the case where the ordering information itself is partially missing, so both the estimate and its achievable range degrade. The mature answer is not to hunt for a variant that produces a bigger number, but to pick the tie-corrected form, name it, and be explicit that the ceiling is set by the coarseness of the scale.

  • Why does tau-a fall short of 1 even when a coarse rating agrees perfectly with an outcome?
    Because tied pairs contribute nothing to `C - D` but stay in the `n*(n-1)/2` denominator. On a five-point scale most pairs are tied on the rating, so the numerator is capped far below the denominator regardless of how well the untied pairs agree. Tau-b removes those ties from the denominator and restores a usable range.
  • What goes wrong if tied values are given sequential ranks in input order instead of midranks?
    You invent ordering that the data does not contain, and the coefficient starts depending on row order — re-sorting the file changes the answer. Midranks are the neutral choice: every member of a tied block gets the same value, the average of the positions the block occupies, so no artificial precedence is created.
  • Does a Spearman rho of 0.55 on a five-point scale mean the association is roughly half of maximum?
    No. Heavy ties impose a ceiling below 1 that comes from the coarseness of the scale, not from the strength of the relationship, so 0.55 may be close to the maximum attainable on that tie structure. Report the tie load alongside the coefficient so readers can judge the available headroom.

saying these in an interview costs you the question

  • Applies the d-squared shortcut to a five-point rating scale
  • Reports Kendall's tau without saying which variant
  • Reads a low tau-a on tied data as weak association
  • Assigns tied values sequential ranks by row order
  • Compares a coefficient from a coarse scale against one from continuous data

context