What does the adjusted Rand index correct for that the raw Rand index does not?
answer
- count pairs, not labels
- most pairs are apart in both
- unrelated labelings still score high
- subtract the expectation, rescale the gap
- zero for random, one for identical
basics
~20 sThe adjustment removes the agreement two labelings reach by chance. The raw Rand index counts agreeing point pairs and stays high even for unrelated labelings; the adjusted version rescales it to about 0 for random, 1 for identical.
solid answer
~50 sThe Rand index counts, over all pairs of points, how often two labelings agree — putting the pair together in both, or apart in both — divided by the number of pairs. The problem is that most pairs are apart in both, so the score is inflated by trivial agreement: two independent random labelings of 1,000 points into 10 equal groups score a raw Rand index of about 0.82. The adjusted Rand index subtracts the value expected under a null model that holds both partitions' cluster sizes fixed and permutes the assignments, then divides by the gap to the maximum: `ARI = (index - expected) / (max - expected)`. Independent labelings land near 0, identical partitions at exactly 1, and a below-chance result goes negative. Because it counts pairs, cluster ids never have to be matched to class ids.
code
python · 21 linesfrom collections import Counter
from math import comb
def rand_and_ari(truth, guess):
n = len(truth)
cells = Counter(zip(truth, guess))
rows = Counter(truth)
cols = Counter(guess)
index = sum(comb(v, 2) for v in cells.values())
A = sum(comb(v, 2) for v in rows.values())
B = sum(comb(v, 2) for v in cols.values())
total = comb(n, 2)
expected = A * B / total
best = (A + B) / 2
rand = (total + 2 * index - A - B) / total
ari = (index - expected) / (best - expected)
return round(rand, 3), round(ari, 3)
truth = [0, 0, 0, 1, 1, 1]
print(rand_and_ari(truth, [0, 0, 1, 1, 2, 2])) # (0.667, 0.242)
print(rand_and_ari(truth, [0, 1, 0, 1, 0, 1])) # (0.467, -0.111)go deeper
Be ready to say what the index compares — pairs of points, together or apart in both partitions — and that the adjusted version reads 0 for unrelated labelings and 1 for identical ones.
Explain the inflation concretely: most pairs are apart in both, so the raw index runs high. Then give the correction as subtract the expected value under fixed cluster sizes and rescale by the gap to the maximum.
Show you can interpret a value in context — which label set, what the plausible ceiling is given label noise — and diagnose a negative result as a data-alignment bug rather than a weak clustering.
Own the choice of measure and the moment it is made. Fix whether the team reports the pair-counting or the information-theoretic chance-corrected score before results are seen, so nobody shops for the metric that flatters the run.
## Pair counting Both indices compare two partitions of the same `N` points — typically a clustering and a known label set — by looking at every one of the `C(N,2)` pairs of points and asking what the two partitions say about that pair. Four cases: - `a`: together in both partitions. - `b`: together in the clustering, apart in the labels. - `c`: apart in the clustering, together in the labels. - `d`: apart in both. `Rand = (a + d) / C(N,2)`. It is symmetric, bounded in [0,1], and needs no correspondence between cluster ids and class ids — that is the reason pair counting is used at all. Note that `c` is the term purity lacks: splitting one true class across clusters lands in `c` and is penalised here. ## Why the raw index is inflated With many groups, most random pairs are apart in both partitions, so `d` dominates and drags the score up regardless of structure. Two independent random labelings of 1,000 points into 10 equal groups: the chance of a pair being together in one labeling is about 0.1, so `a` is about 0.01 of pairs and `d` about 0.81, giving a Rand index near 0.82. Nothing was learned, and the score reads like strong agreement. Worse, the inflation depends on the number of clusters, so raw Rand values from partitions with different k are not comparable. ## The correction Write `index = sum_ij C(n_ij, 2)` over the contingency table of clusters against classes — this is `a`, the pairs together in both. Let `A = sum_i C(a_i, 2)` over row sums and `B = sum_j C(b_j, 2)` over column sums. Under the permutation null model — cluster sizes and class sizes held fixed, assignments shuffled — the expected value of `index` is `A * B / C(N,2)`, and its maximum is `(A + B)/2`. Then: `ARI = (index - expected) / ((A + B)/2 - expected)` Properties that follow: - **Identical partitions score exactly 1.** Nothing else does. - **Independent labelings score about 0** in expectation, for any k and any N. This is what makes ARI comparable across different numbers of clusters — the property purity and raw Rand both lack. - **Negative values are possible**, down to about -0.5 in practice. A negative ARI means the partition separates same-class points *more* than shuffling would; on real data it usually signals a bug, such as misaligned label and prediction arrays, rather than a bad-but-honest clustering. - **Invariant to relabeling.** Renaming cluster 3 to cluster 7 changes nothing, because only co-membership is counted. No Hungarian-style matching step is needed. ## Interpreting a value There is no universal threshold. Rough working intuition on real data: near 0 means no relationship to the labels; 0.1 to 0.3 means a weak but real correspondence; above 0.6 is strong; above 0.9 usually means the labels are nearly derivable from the features, which is worth checking for leakage. Always report which label set the score is against — an ARI of 0.31 between a clustering of merchant records and their merchant-category codes is a statement about that taxonomy, not about the clustering in the abstract. ## Relation to the mutual-information family Normalized mutual information (NMI) measures the same agreement information-theoretically: the mutual information between the two labelings divided by a normalising combination of their entropies. It is bounded in [0,1] and also relabeling-invariant, but plain NMI is *not* corrected for chance and drifts upward as the number of clusters grows, for the same structural reason raw Rand does. Adjusted mutual information (AMI) applies the same subtract-the-expectation trick and is the chance-corrected member of that family. ARI and AMI can disagree, and the disagreement is informative. Pair counting punishes splitting hard: cutting each true class cleanly in half produces many `c` pairs and a mediocre ARI, while mutual information stays high because a pure refinement still determines the class almost completely. So when your clustering is a fine-grained refinement of the label set, expect NMI or AMI to look better than ARI. Pick the one whose error mode you care about, decide before you look at the numbers, and report the choice. ## What ARI cannot do It needs a trusted label set. When none exists — the usual case in clustering — external agreement is unavailable and you fall back on geometric indices or on resampling stability, where ARI reappears in a different role: as the agreement measure between two clusterings of overlapping subsamples.
- Can the adjusted Rand index go negative, and what should you conclude if it does?Yes. Negative means the partition puts same-class points apart more often than a random shuffle with the same cluster sizes would, so agreement is below the chance baseline. Small negatives near -0.02 are just noise. A clearly negative value on real data is almost always a bug — labels and cluster assignments misaligned, rows reordered, or a filtered subset compared against unfiltered labels — so check the join before interpreting it.
- Why does the adjusted Rand index need no matching between cluster ids and class ids?Because it scores pairs of points, not labels. The only question asked of each pair is whether the two partitions agree about keeping it together or apart, and that is unchanged when you rename cluster 3 to cluster 7 or shuffle every id. This is why pair-counting measures suit clustering, where the ids are arbitrary output of the algorithm and carry no meaning.
- When would you prefer adjusted mutual information over the adjusted Rand index?When you care about how much the clustering tells you about the labels rather than about pairwise co-membership. Pair counting penalises splitting a true class harshly, so a clean refinement — each class cut into two tidy clusters — scores mediocre on ARI while the information-theoretic measure stays high, since a refinement still determines the class. Choose by which error matters, and choose before seeing the numbers.
- What does a Rand index of 0.82 tell you on its own?Almost nothing, because it does not say what the baseline is. Two independent random labelings of 1,000 points into 10 groups already score about 0.82; with more groups, the raw index climbs further because more pairs are apart in both partitions. The raw value is also not comparable across partitions with different cluster counts, which is exactly the problem the adjustment removes.
Grading a multiple-choice test without subtracting for guessing: a candidate who answers at random still looks respectable until you set the baseline where random guessing lands.
saying these in an interview costs you the question
- Reads a raw Rand index of 0.8 as strong agreement
- Claims the adjusted Rand index cannot be negative
- Thinks cluster ids must be matched to class ids first
- Believes normalized mutual information is corrected for chance
- Uses the adjusted Rand index when no trusted label set exists