Why can cluster purity not be reported on its own as a clustering quality score?
answer
- think about what splitting a cluster does
- only the majority class is counted
- no penalty for shattering a class
- monotone as the cluster count grows
- one point per cluster scores perfectly
basics
~20 sPurity credits each cluster only for its majority true class, so splitting a cluster can never lower it. Push the number of clusters up and purity climbs toward 1.0, reaching it when every point sits alone.
solid answer
~50 sPurity is computed by taking each cluster, finding the true class that appears most often in it, adding up those majority counts across all clusters, and dividing by the number of points. Nothing in that sum punishes splitting: if you cut a cluster in two, each half still contributes at least its share of the old majority class, so purity is monotone non-decreasing as the number of clusters grows. On a 5,000-row product taxonomy, a clustering with 4 clusters might score 0.55 and one with 900 clusters 0.97, and at 5,000 singleton clusters purity is exactly 1.0 against any label set whatsoever. So a purity figure quoted without the number of clusters carries almost no information. Report it only when comparing partitions at the same k, and pair it with a chance-corrected, split-penalising measure such as the adjusted Rand index.
go deeper
Be ready to state the recipe — majority true class per cluster, summed, divided by the number of points — and the one-point-per-cluster case that makes it 1.0. Always say k when you quote a purity.
Explain why splitting can never lower purity, and name the two missing properties: no completeness term, no correction for chance. Show the single-cluster floor equals the largest class share.
Show the judgment of when purity is still useful: same-k comparisons and per-cluster diagnostics off the contingency table, always reported next to a chance-corrected measure rather than instead of one.
Own the reporting standard. Decide which external measures your team quotes by default and insist that any homogeneity-style number ships with k and a chance-corrected companion, so dashboards cannot be gamed by raising the cluster count.
## What purity actually computes Purity is an external validation measure: it scores a clustering against a set of known class labels. The recipe is: 1. Build the contingency table of clusters against true classes: cell `n_ij` is the number of points in cluster `i` that carry true class `j`. 2. For each cluster `i`, take `max_j n_ij` — the size of the single most common true class inside it. 3. Sum those maxima over all clusters and divide by `N`, the number of points. `purity = (1/N) * sum_i max_j n_ij` A tiny example: 10 points, true classes A A A A A A B B B B. A clustering produces cluster 1 = {A A A A B} and cluster 2 = {A A B B B}. Cluster 1's majority is A with 4; cluster 2's majority is B with 3. `purity = (4 + 3)/10 = 0.7`. The reading is intuitive — 70% of points sit in a cluster whose dominant label matches their own — and that intuitiveness is exactly why the measure keeps getting reported. ## The defect: monotone in the number of clusters Take any cluster and split it into two parts. Let the original cluster's majority class be `m` with count `c_m`, distributed as `c_m1` and `c_m2` across the two parts. The new contribution is `max_j n_1j + max_j n_2j`, and since `max_j n_1j >= c_m1` and `max_j n_2j >= c_m2`, the new contribution is at least `c_m1 + c_m2 = c_m`. Splitting therefore never decreases purity, and usually increases it. Carry that to the limit. With one cluster per point, every cluster's majority count is 1, the sum is `N`, and `purity = 1.0` — against any label set, however unrelated to the data. On a 5,000-row product taxonomy that means a clustering into 5,000 clusters scores a perfect 1.0 while telling you nothing. Long before the limit, the number is already inflated: a few hundred clusters over a few thousand rows will routinely land above 0.9. The opposite extreme is just as misleading in the other direction. With a single cluster covering everything, purity equals the proportion of the largest true class. On a label set where 80% of rows share one class, the do-nothing clustering already scores 0.8, so 0.8 is not a good score there — it is the floor. ## Two properties purity lacks **It has no completeness term.** Purity asks whether each cluster is internally homogeneous; it never asks whether a true class was kept together. Shattering one true class across 50 clusters costs nothing. Measures that count point *pairs* — the Rand family — do penalise this, because two points sharing a class but sitting in different clusters register as a disagreement. The information-theoretic family expresses the same idea as a pair of numbers, homogeneity and completeness, combined into V-measure. **It is not corrected for chance.** Purity has no baseline that random labelings score. Assign 1,000 points to 50 clusters at random and purity will still come out well above the largest-class share, purely because small random clusters are homogeneous by accident. Chance-corrected measures fix a reference point: the adjusted Rand index and adjusted mutual information both sit near 0 for independent labelings, so a score is interpretable without knowing k. ## When purity is still legitimate Purity is not banned; it is unsafe *unqualified*. It is fine for: - Comparing two clusterings that produce the **same** number of clusters — the inflation is then a constant across both. - Explaining a result to a non-technical audience, alongside a defensible measure, because the sentence 'x% of rows fall in a cluster dominated by their own category' lands where 'ARI = 0.31' does not. - Diagnosing which specific clusters are mixed, by reading per-cluster purity off the contingency table rather than the single aggregate. The practical rule for an interview: never quote a purity number without stating k, and never use purity to choose between clusterings with different k. If someone shows you a purity of 0.97 and no k, the first question back is how many clusters produced it.
- Does purity mislead at the other extreme too, when there is only one cluster?Yes. With a single cluster containing every point, the majority count is just the size of the largest true class, so purity equals that class's share. On a label set where one class holds 80% of rows, the do-nothing clustering scores 0.8. That number is the floor to beat, not evidence of quality, and it is why an imbalanced label set makes purity look flattering everywhere.
- Which measure supplies the penalty for splitting that purity is missing?Pair-counting measures do it directly: the Rand family counts pairs of points that share a true class, and a pair separated into different clusters counts against you, so shattering a class is punished. The information-theoretic route names the same property completeness, pairs it with homogeneity — purity's analogue — and combines the two into V-measure, so both errors show up in one score.
- Is purity ever the right thing to report?Yes, in two situations. When comparing clusterings that all use the same number of clusters, the inflation is a shared constant and purity is a fair ranking. And per cluster, read off the contingency table, it is a useful diagnostic that tells you exactly which clusters are mixed. What is unsafe is a single aggregate purity quoted without k and compared across different k.
Grading a filing system by how uniform each drawer is: give every document its own drawer and every drawer becomes perfectly uniform, which says nothing about the filing.
saying these in an interview costs you the question
- Says a purity near 1.0 always means a strong clustering
- Compares purity across clusterings with different numbers of clusters
- Believes purity penalises splitting one true class across clusters
- Treats purity as corrected for chance like the adjusted Rand index
- Reports a purity figure without stating how many clusters produced it