How is Kendall's tau computed from concordant and discordant pairs?
answer
- look at every pair, not every point
- same ordering versus opposite ordering
- n(n-1)/2 pairs in the denominator
- probability of agreement minus disagreement
- (C - D) over total pairs, ties excluded
basics
~20 sKendall's tau compares every pair of observations. A pair is concordant when both variables order it the same way, discordant when they disagree. Tau is concordant minus discordant pairs, divided by the total number of pairs.
solid answer
~50 sTake all `n*(n-1)/2` pairs of observations. A pair is **concordant** when the observation that is higher on x is also higher on y, and **discordant** when the orderings disagree. With C concordant and D discordant pairs and no ties, `tau = (C - D) / (n*(n-1)/2)`. Say two judges rank eight contestants: there are 28 pairs, and if the judges agree on the relative order of 22 of them and disagree on 6, then `tau = (22 - 6)/28 = 0.57`. The interpretation is unusually direct — tau is the probability that a randomly chosen pair is ordered the same way by both variables minus the probability it is ordered oppositely, so 0.57 means agreement exceeds disagreement by 57 percentage points. Tau runs from -1 to +1, is unchanged by any increasing transform, and needs a tie-corrected form once values repeat.
go deeper
Know the vocabulary: a pair is concordant when both variables order it the same way, discordant when they disagree, and tau is built from those two counts over all pairs.
Be able to derive the coefficient on a small example — count pairs, classify them, divide by n*(n-1)/2 — and state the agreement-minus-disagreement probability reading out loud.
Show you know when the pair-counting definition needs the tie-corrected form, why the naive computation is quadratic, and how to explain tau to a non-technical stakeholder without hand-waving.
Own the choice of coefficient as a reporting standard: tau buys a defensible plain-language interpretation, and switching measures between reports to flatter a result is the failure mode to legislate against.
## The pair-counting idea Kendall's tau measures monotone association by asking one question about every pair of observations: **do the two variables agree about which member of the pair is larger?** For observations `i` and `j` with values `(x_i, y_i)` and `(x_j, y_j)`: - The pair is **concordant** if `(x_i - x_j)` and `(y_i - y_j)` have the same sign — the one that is higher on x is also higher on y. - The pair is **discordant** if those differences have opposite signs. - The pair is **tied** if either difference is zero. There are `n*(n-1)/2` pairs in a dataset of size n. Count the concordant ones as C and the discordant ones as D. With no ties anywhere, `tau_a = (C - D) / (n*(n-1)/2)` That is Kendall's tau in its simplest form. ## Worked example: two judges, eight contestants Eight contestants are ranked by two judges. The number of pairs is `8*7/2 = 28`. Walk the pairs and classify each: for the pair (contestant 3, contestant 7), if judge A puts 3 above 7 and judge B also puts 3 above 7, that pair is concordant; if judge B reverses them, it is discordant. Suppose 22 pairs come out concordant and 6 discordant. Then `tau = (22 - 6) / 28 = 16/28 = 0.57` The two judges agree on the ordering of about 79% of pairs (22/28) and disagree on about 21% (6/28); tau is the difference, `0.79 - 0.21 = 0.57`. ## Why the interpretation is the selling point Most association measures are hard to explain to a non-specialist beyond 'higher is stronger'. Tau is not: it is literally **P(concordant) - P(discordant)** for a randomly drawn pair. You can state it as a betting proposition — pick two contestants at random, and the judges are 57 percentage points more likely to agree on their order than to disagree. That plain-language reading is why tau is often preferred when the audience is not statistical, and why it is the natural coefficient when the underlying data really is a set of orderings rather than measurements. A few consequences fall straight out of the definition: - **Range.** With no ties, perfect agreement means D = 0 and C is every pair, so tau = +1. Perfect reversal means C = 0 and tau = -1. Independent orderings put C and D near each other, so tau sits near 0. - **Invariance.** Concordance depends only on the *signs* of the differences, so any strictly increasing transform of x or y (logs, unit changes, monetary conversion) leaves every pair's classification unchanged and therefore leaves tau unchanged. - **Bounded influence.** A single observation belongs to only `n - 1` of the pairs, so no matter how numerically extreme it is, it can change tau by at most about `4/n`. That is the precise sense in which tau resists outliers. ## Ties, and the variants Once values repeat, pairs tied on x or tied on y are neither concordant nor discordant, and `tau_a` cannot reach 1 any more — the tied pairs sit permanently in the numerator's dead zone while the denominator still counts them. The standard fix is **tau-b**, which shrinks the denominator to exclude ties on each variable separately: `tau_b = (C - D) / sqrt((C + D + T_x) * (C + D + T_y))` where `T_x` is the number of pairs tied on x but not on y, and `T_y` the number tied on y but not on x. Pairs tied on both are dropped entirely. When there are no ties, `T_x = T_y = 0` and tau-b collapses back to tau-a. A further variant, **tau-c**, adjusts for tables where the two variables have very different numbers of distinct levels. ## Practical notes The definition is a double loop over pairs, which costs on the order of `n^2` comparisons; counting inversions with a merge-sort-style pass brings it down to `n log n`, which matters once n reaches the hundreds of thousands. For small n the naive version is perfectly fine and is what you would sketch on a whiteboard. Tau and the other standard rank coefficient, Spearman's rho, both measure monotone association and both hit +1 on any strictly increasing relationship, but they are on different scales and their intermediate values are not comparable. Pick one for a given report and stay with it rather than quoting whichever came out larger. Finally, tau is a *descriptive* number: it summarises the ordering agreement in the data you have. Judging whether an observed tau is more than sampling noise would take a separate testing procedure, and that is a different question entirely from computing the coefficient.
- How many pairs are compared when computing Kendall's tau on 8 observations?28. Every unordered pair counts once, so the total is `n*(n-1)/2 = 8*7/2 = 28`. That denominator is what turns the raw concordant-minus-discordant count into a coefficient bounded by -1 and +1, and it is also why the naive computation grows with the square of the sample size.
- What does a Kendall's tau of 0 tell you about the two variables?That concordant and discordant pairs are balanced: pick two observations at random and the variables are equally likely to agree or disagree about which is larger. That rules out a consistent monotone trend, but not dependence — a symmetric U-shaped relationship produces exactly this balance while being fully deterministic.
- Why is Kendall's tau considered resistant to an extreme observation?Because concordance depends only on the sign of each difference, not its size. Moving one observation to an absurd value can flip at most the `n - 1` pairs it belongs to, bounding its effect on tau at roughly `4/n`. The magnitude of the outlier is irrelevant once its position in the ordering is fixed.
It scores two judges the way you would settle an argument: go through every pair of contestants and count how often they agree on who placed higher.
saying these in an interview costs you the question
- Divides by n instead of the number of pairs
- Counts concordant pairs only and calls that tau
- Says tau and rho are interchangeable numbers on the same data
- Uses the untied formula when many values repeat
- Treats a large tau as evidence that x causes y