skip to content

How do you choose where to cut a dendrogram to get a flat set of clusters?

level: middleimportance: should knowfreq 57%

answer

  1. a horizontal line across the merges
  2. n minus k merges leaves k clusters
  3. look for the long empty stretch
  4. heights do not transfer between trees
  5. unbalanced trees may need uneven cuts

basics

~20 s

A horizontal cut at some height keeps every merge below it and undoes the rest; equivalently, stopping after n minus k merges leaves k clusters. Cut inside a wide gap between consecutive merge heights, then check the groups against outside constraints.

solid answer

~50 s

A dendrogram is a nested family of partitions, so producing one flat clustering means cutting it. A horizontal cut at height `h` keeps every merge that happened below `h` and undoes everything above; equivalently, stopping the merge sequence after `n - k` merges leaves exactly `k` clusters. To pick `h` I look for a long vertical stretch in which nothing merges — a wide gap means the groups below it survived a broad range of thresholds before anything forced them together, so cutting inside that gap is where the tree is least ambiguous. I then check the cut against constraints from outside the tree: minimum usable group size, how many groups anyone downstream can act on, and whether the same cut survives refitting on a resample. Cuts need not be horizontal — on a badly unbalanced tree, cutting different branches at different heights is often honest.

go deeper

for a junior

Know that the tree is not yet a clustering and that you cut it, and be able to state the two equivalent ways: a horizontal line at a height, or stopping after n minus k merges.

for a middle

Explain why a wide vertical gap between consecutive merge heights is the tree's own evidence for a cut, and why a cut height cannot be carried across trees built with different linkage rules or different feature scaling.

for a senior

Demonstrate that you validate the cut rather than eyeball it: check cluster sizes for usability, refit on a resample to see whether the same groups return, and be willing to report that no natural cut exists when the merge heights rise smoothly.

for a principal

Own the decision that the cut is where business constraints legitimately enter a statistical procedure. Argue for a granularity the organisation can act on, and insist the report distinguishes structure the data supports from a granularity you chose.

## What a cut actually is An agglomerative run produces `n - 1` merges, each with a height. Draw a horizontal line at height `h`. Every merge whose height is below `h` is accepted; every merge above `h` is undone. The connected subtrees left below the line are the clusters. Because merge heights are non-decreasing for single, complete, average and Ward linkage, this line is well defined and the resulting groups are exactly the clusters that existed at the moment the algorithm's current-best distance first exceeded `h`. There is an equivalent count-based view. Since each merge reduces the cluster count by one, stopping after `n - k` merges leaves exactly `k` clusters. Cutting by height and cutting by count are the same operation described from two ends, so you can always obtain any `k` you want — cut anywhere between the `(n-k)`-th and `(n-k+1)`-th merge heights. ## Reading the heights for a natural cut The tree carries one intrinsic signal about where to cut: the spacing of consecutive merge heights. If merges happen at 0.4, 0.5, 0.55, 0.6 and then the next one is at 2.9, the structure present just below 2.9 held together across a very wide range of thresholds. Nothing wanted to merge for a long time, which is the definition of well-separated groups under that linkage. Cutting inside the widest gap is therefore the tree's own most defensible answer, and it is what people mean when they say a dendrogram has an obvious cut. The honest caveat is that a clean gap often does not exist. Real merge heights frequently rise smoothly with no dominant gap at all, and that is information: the data has no strongly separated group structure at that granularity, and any cut you choose is a convention rather than a discovery. Saying so is a better answer than picking the third-largest gap and presenting it as structure. A second caveat: heights are linkage-specific and unit-specific. A height of 1.8 means something entirely different under Ward than under average linkage on the same data, and rescaling the features moves every height. A cut height is never transferable between trees; only a count is. ## Constraints from outside the tree The cut is where domain requirements legitimately enter. A cut that yields 40 clusters, 33 of them containing one row each, is arithmetically fine and operationally useless. Typical outside constraints: - **Minimum group size** — clusters below it are usually merged upward or set aside as unassigned. - **Actionable count** — if a downstream process can support four treatments, a 19-cluster cut is a report nobody can use, and cutting higher is a legitimate response. - **Interpretability** — the groups must be describable in terms a domain expert recognises. - **Stability** — refit on a bootstrap resample or a random half of the rows and see whether the same cut recovers substantially the same groups. A cut that changes completely under resampling was reading noise. A worked example: build a dendrogram over 30 equity return series, using distances derived from their return correlations. Cut it low and you get 20 fragments; cut it high and everything is one market; somewhere in between there is a height at which the branches align with recognisable sectors. The check that the cut is real is not the picture but whether that alignment holds when you rebuild the tree on a different time window. ## Non-horizontal cuts and inversions A horizontal cut assumes one global threshold suits every part of the data. Trees are often unbalanced — one dense region merges at tiny heights while a diffuse region needs large ones — and forcing a single line either shatters the dense region into fragments or swallows the diffuse one whole. It is legitimate to cut branch by branch at different heights, provided you say that is what you did, since the result is still a valid partition (each chosen subtree is disjoint from the others). One structural warning: if the tree was built with centroid or median linkage, merge heights are not monotone and the dendrogram can contain inversions, where a later merge is drawn below an earlier one. A horizontal line can then cross the same branch twice and the notion of cutting at a height breaks down. With those rules, cut by merge count instead — or, better, use a linkage whose heights are monotone. ## Common mistakes Treating the number of visible branches in the plot as the number of clusters (rendering routines often truncate the tree for display); assuming a cut height carries over to a tree built with a different linkage; and reporting the cut as the answer without checking that the groups are non-trivial in size and stable under resampling.

  • Can a horizontal cut always produce exactly the number of clusters you want?
    Yes, provided the merge heights are monotone. Each merge reduces the count by one, so cutting anywhere between the (n-k)-th and (n-k+1)-th merge heights leaves exactly k clusters. Ties complicate it slightly: if several merges share a height, cutting at that height may skip past your k, so cut by merge count instead.
  • Why can a centroid-linkage dendrogram make a horizontal cut meaningless?
    Centroid and median linkage are not monotone, so a merge can occur at a lower height than one that came before it — an inversion. The picture then has crossing brackets and a horizontal line can intersect the same branch twice, so the set of subtrees below the line is not a valid partition. Cut by merge count, or use a monotone linkage.
  • The merge heights rise smoothly with no obvious gap. What do you report?
    That the data shows no strongly separated structure at this granularity under this linkage, and that any cut is therefore a convention rather than a discovery. I would still supply a cut if one is needed, chosen by the operational constraint — the number of groups the downstream process can act on — and state plainly that it is a chosen granularity, not detected structure.

saying these in an interview costs you the question

  • Assumes the plotted branch count equals the cluster count
  • Reuses a cut height across trees built with different linkages
  • Presents any gap as evidence of real separation
  • Never checks resulting cluster sizes for usability
  • Cuts a non-monotone tree horizontally without noticing inversions

context