What do the Davies-Bouldin and Calinski-Harabasz indices measure, and why can they rank two partitions oppositely?
answer
- one you minimise, one you maximise
- worst neighbour versus whole configuration
- radii over centroid distance
- between-dispersion over within-dispersion
- different aggregation, different ranking
basics
~20 sDavies-Bouldin averages, over clusters, the worst ratio of two clusters' spreads to the distance between their centroids, and lower is better. Calinski-Harabasz is a between-cluster over within-cluster dispersion ratio, and higher is better. Different aggregations can rank partitions oppositely.
solid answer
~50 sBoth are internal indices computed from centroids and spreads, but they aggregate differently. Davies-Bouldin takes, for each cluster, its most confusable neighbour — the neighbour maximising `(spread_i + spread_j) / distance(centroid_i, centroid_j)` — and averages those worst cases; it is minimised, with 0 as the floor. Calinski-Harabasz is a variance ratio: between-cluster dispersion divided by within-cluster dispersion, each scaled by its degrees of freedom, and it is maximised with no upper bound. So Davies-Bouldin is a worst-pair statistic and Calinski-Harabasz is a whole-configuration statistic. Take two partitions of the same thirteen-column beverage chemistry table at the same cluster count: one has tidy clusters but two of them nearly touching, the other has nothing touching but one diffuse, sprawling cluster. Davies-Bouldin punishes the first and Calinski-Harabasz punishes the second, so the rankings invert. Neither is wrong; they encode different notions of a bad partition.
go deeper
At minimum remember the direction of each: Davies-Bouldin is better when lower, Calinski-Harabasz when higher. Quoting one in the wrong direction is the mistake interviewers actually catch.
Be able to describe what each is built from — cluster radii over centroid distances for one, between-cluster over within-cluster dispersion for the other — and why the worst-pair maximum is what makes Davies-Bouldin behave differently.
Show you would report both and explain the disagreement in terms of the application, rather than selecting the index that flatters the result after seeing it.
Own the criterion decision before results exist: state which failure mode the business cannot tolerate, confusable segments or weak overall structure, so the evaluation is not chosen post hoc.
### Davies-Bouldin For each cluster `i`, define `S_i` as the average distance from its members to its own centroid — a radius. For each pair `i, j`, define `M_ij` as the distance between the two centroids. The similarity of the pair is `R_ij = (S_i + S_j) / M_ij`: two fat clusters whose centres are close score high, two tight clusters far apart score low. The index then takes, for each cluster `i`, the **maximum** `R_ij` over all other clusters — its single most confusable neighbour — and averages those maxima over the clusters. **Lower is better**, and 0 is the unreachable ideal of zero-radius clusters infinitely far apart. The defining property is the maximum. Davies-Bouldin does not care that eleven of your twelve cluster pairs are beautifully separated; each cluster is judged only by its worst neighbour, so a single confusable pair drags the whole index. ### Calinski-Harabasz Also called the variance ratio criterion. Within-cluster dispersion is the total squared distance from points to their own centroid; between-cluster dispersion is the total squared distance from cluster centroids to the overall mean, weighted by cluster size. The index is `CH = (between_dispersion / (k - 1)) / (within_dispersion / (n - k))` for `n` points and `k` clusters. **Higher is better** and there is no upper bound. The degrees-of-freedom divisors give it the shape of a variance-ratio statistic rather than a raw quotient. Here the defining property is that everything is pooled. Every point and every cluster contributes to both sums, and because the distances are squared, points far from their centroid contribute disproportionately. Calinski-Harabasz is a statement about the overall configuration, not about any single pair. ### Why they can invert Consider two candidate partitions of the same thirteen-column table of beverage chemistry measurements — acidity, sugar, phenolics and so on — computed at the same cluster count by two different methods. - **Partition A:** every cluster is compact, but two of them sit almost on top of each other. Within-dispersion is low overall, so Calinski-Harabasz is high. But that one nearly-touching pair produces a large `R_ij` for both clusters involved, so Davies-Bouldin looks bad. - **Partition B:** no two clusters are anywhere near each other, but one cluster is a sprawling, diffuse mass. Every cluster's worst neighbour is comfortably far, so Davies-Bouldin looks good. Yet the diffuse cluster inflates the squared within-dispersion, so Calinski-Harabasz drops. A ranks better on one index and B on the other. Neither index malfunctioned. They answer different questions: *how bad is my worst confusion?* versus *how much of the total dispersion is between clusters rather than inside them?* A second, subtler source of divergence is the exponent. Davies-Bouldin is built from average distances, Calinski-Harabasz from squared ones, so a few far-flung members hurt Calinski-Harabasz much more than they hurt Davies-Bouldin. ### Practical handling Report both, together with what they disagree about, rather than quietly picking the one that supports the answer you already prefer — an index chosen after seeing the result is not evidence. Decide in advance which failure the application cannot tolerate: if two segments being confusable makes them unusable downstream, the worst-pair criterion is the relevant one; if you care about how much structure the partition captures overall, the variance ratio is. Both are invariant to multiplying every feature by one constant — numerator and denominator scale together in each — but neither is invariant to rescaling features differently, which changes the geometry, the partition and both scores. And both are computed from centroids and straight-line spreads, so they share the assumption that clusters are compact, roughly round and comparably spread. Where that assumption fails, the two indices tend to be wrong in the same direction, and their agreement is no comfort at all.
- Which direction is better for each index, and what are their bounds?Davies-Bouldin is minimised: it is non-negative, 0 is the unattainable ideal, and lower means less confusable clusters. Calinski-Harabasz is maximised: it is non-negative and unbounded above, so there is no target value, only comparisons within one setup. A common slip is reporting a rise in Davies-Bouldin as an improvement.
- Why does Calinski-Harabasz divide by (k - 1) and (n - k)?Those are the degrees of freedom of the between-cluster and within-cluster dispersion terms, which turns a raw quotient of sums into a variance-ratio statistic. Without them the ratio would react even more strongly to how many clusters and how many points went into it, making comparisons across differently shaped configurations harder to interpret.
- Are these two indices sensitive to feature scaling?Both are ratios, so multiplying every feature by the same constant leaves them unchanged — the numerator and denominator scale together. Rescaling features differently, however, changes distances, changes which partition is found, and changes both scores. Fix the preprocessing first and compare only within it.
saying these in an interview costs you the question
- Says a higher Davies-Bouldin index means a better partition
- Treats the two indices as interchangeable quality scores
- Believes Calinski-Harabasz is bounded between 0 and 1
- Thinks Davies-Bouldin averages over all pairs rather than the worst
- Picks whichever index favours the preferred conclusion