What is cophenetic correlation, and how would you use it to compare two linkage methods?
answer
- the tree implies its own distances
- height where two points first meet
- correlate tree distances with the originals
- one rule wins it almost by construction
- faithfulness is not usefulness
basics
~20 sThe cophenetic distance between two observations is the height at which they first join in the tree. Cophenetic correlation is the correlation between those tree distances and the original pairwise distances; the higher-scoring tree distorts the input geometry less.
solid answer
~50 sThe cophenetic distance between two observations is the height of the merge at which they first land in the same cluster — the dendrogram's own account of how far apart they are. Cophenetic correlation is the Pearson correlation between those tree distances and the original pairwise distances, taken over all pairs. To compare linkages, build one tree with complete linkage and another with average linkage on the same distances, compute both correlations, and the higher one is the tree that distorts the input geometry less. Average linkage usually wins, because its merge heights are literally averages of the distances being compared. The catch is what it measures: faithfulness of the whole tree, not usefulness of any cut. A tree can track distances beautifully and still contain no height that yields groups anyone can act on, so I use it as a tie-breaker, never as the selection criterion.
go deeper
Know the definition: the cophenetic distance between two points is the height at which they first join in the tree, and the correlation compares those heights with the original pairwise distances.
Explain why the comparison is only meaningful between trees built on the same distances, and why the rule that averages cross-cluster distances tends to score highest almost by construction.
Show the limit of the statistic: it scores the whole tree's faithfulness, not the quality of any cut, so a chained tree can score well while being useless. Use it as a tie-breaker alongside stability and interpretability.
Own the wider point that a faithful tree over the wrong distances is still the wrong answer. Push the discussion upstream to whether the similarity being modelled matches the decision, rather than adjudicating linkage on a third decimal place.
## The cophenetic distance A dendrogram implies a distance of its own. For any two observations `i` and `j`, look up the height of the merge at which they first end up inside the same cluster; that height is the **cophenetic distance** `c(i, j)`. Every pair has one, so a tree over `n` observations induces a full `n x n` matrix of cophenetic distances alongside the original distance matrix `d(i, j)` the tree was built from. That induced distance has a strong structural property: it is ultrametric, meaning that for any three points the two largest of the three cophenetic distances are equal. This is not an accident — it is exactly what being representable as a tree means, and it is also why no dendrogram can reproduce an arbitrary distance matrix. Any tree is a lossy summary of the distances, and the interesting question is how lossy. ## The correlation **Cophenetic correlation** is the Pearson correlation between the `n(n-1)/2` original pairwise distances and the corresponding cophenetic distances. It lies in `[-1, 1]` and in practice sits high — values of 0.7 to 0.9 are routine, so the number is only meaningful in comparison, never in absolute terms. A value of 0.82 is not good or bad; it is good or bad relative to what another linkage rule achieved on the same distances. ## Comparing linkages with it The canonical use is exactly that comparison. Take one distance matrix, build a complete-linkage tree and an average-linkage tree from it, compute the cophenetic correlation of each, and you have a quantitative statement about which tree preserves the original geometry better. This is worth doing because linkage choice is otherwise argued from cluster-shape priors alone, and this gives one measurable axis. Average linkage typically scores highest, and the reason is mechanical rather than mysterious: its merge height is the mean of the cross-cluster distances involved, so heights track the raw distances closely by construction. Complete linkage systematically inflates heights (it reports the farthest pair), single linkage systematically deflates them (it reports the nearest pair), and both distortions cost correlation even when the tree's grouping is perfectly sensible. Ward's heights are not distances between points at all — they are variance increases — so its cophenetic correlation is not on a comparable footing, and comparing Ward against average linkage this way is a category error more often than an insight. ## What it does not tell you This is the part that separates a memorised definition from understanding. Cophenetic correlation scores **the whole tree's faithfulness to the input distances**. It says nothing about: - whether any cut of that tree yields useful groups, - whether the groups it yields are interpretable or actionable, - whether the structure is stable when the data is resampled, - whether the distances the tree was built from were the right distances in the first place. If the inputs encode the wrong notion of similarity, a tree that reproduces them faithfully has faithfully reproduced the wrong thing. The failure mode is concrete: single linkage on chained data can achieve a respectable correlation while producing one sprawling cluster plus a scatter of singletons — a faithful tree that is operationally worthless. Meanwhile a Ward tree with a lower score may be the one that yields the clean, balanced, business-usable groups. Selecting a linkage by cophenetic correlation alone reliably picks average linkage, which is a fine default but did not need a statistic to justify. ## How to use it honestly Treat it as a diagnostic with a narrow claim. Two sensible uses: 1. **Tie-break.** Two linkage rules both give plausible, similarly interpretable trees; take the one that distorts distances less. 2. **Warning flag.** A conspicuously low correlation says the tree is a poor summary of your distances — the data may not be hierarchical in structure at all, and the dendrogram you are about to show a stakeholder implies relationships the data does not support. What it should not do is override interpretability, stability under resampling, or the operational constraints that decide the cut. Presenting a linkage choice as justified because it scored 0.87 against 0.84 is over-reading a comparison that is not that precise.
- A tree scores 0.91 cophenetic correlation. Does that mean the clusters are good?No. It means the tree reproduces the input pairwise distances well, which is a statement about the whole hierarchy, not about any cut of it. A faithful tree can still have no height that yields interpretable, usably sized groups, and if the input distances encoded the wrong notion of similarity the tree has faithfully reproduced the wrong thing.
- Why does average linkage usually achieve the highest cophenetic correlation?Because its merge height is the mean of the cross-cluster distances being summarised, so heights track the raw distances by construction. Complete linkage reports the farthest pair and systematically inflates heights, single linkage reports the nearest pair and deflates them; both distortions cost correlation even when the grouping itself is sensible.
- Why can no dendrogram reproduce an arbitrary distance matrix exactly?Cophenetic distances are ultrametric: for any three points, the two largest of their three cophenetic distances must be equal. Ordinary distances rarely satisfy that, so representing them as a tree is necessarily lossy. Cophenetic correlation is just a measure of how much was lost.
saying these in an interview costs you the question
- Treats a high score as proof the clusters are good
- Reads an absolute value as good or bad without comparison
- Thinks it measures distance from a point to a cluster centre
- Compares Ward heights against distance-based linkages directly
- Picks the linkage rule on this statistic alone