Why does average silhouette rank a round four-way split of two interleaved spirals above the correct partition?
answer
- the metric encodes an assumption
- compact round blobs score well
- far end of your own arm
- geometry, not connectivity
- change the distance or the criterion
basics
~20 sSilhouette rewards points near their own cluster mates and far from the nearest other cluster. Along a spiral the far end of your own arm is distant while a neighbouring arm is close, so compact round chunks win.
solid answer
~50 sSilhouette measures cohesion against separation using distances between points, which encodes an assumption: a good cluster is a compact, roughly round blob. Two interleaved spirals violate it. For a point midway along one arm, the mean distance to its own cluster is large because the arm stretches far away, while the nearest points of the other spiral may be a step away, so the *correct* partition scores near zero or negative. Carving the plane into four compact round chunks gives small within-cluster distances and clear gaps, so it scores higher despite cutting across the real structure. Davies-Bouldin, Calinski-Harabasz and inertia share the bias — all are built from centroids or mean distances, and the centroid of a curved cluster may not lie inside it. So an internal index cannot fairly arbitrate between methods with different shape assumptions.
go deeper
Remember that a low silhouette does not automatically mean there is no structure. It can mean the structure is not the compact, roughly round kind the score is able to see.
Explain the mechanism concretely: on a long curved cluster the mean distance to your own members is large while a neighbouring group can be close, so cohesion loses to separation.
Show the operational consequence. Refuse to arbitrate between methods with different shape assumptions using a compactness index, and offer a concrete alternative such as a distance that follows the data or an outside criterion.
Own the evaluation policy. Decide up front what evidence counts as a clustering being good for the organisation, so a geometry statistic never becomes the default acceptance test for structures it cannot represent.
### The assumption hiding in the formula Every internal index in common use is a function of distances. Silhouette compares a point's mean distance to its own cluster against its mean distance to the nearest other cluster. Inertia sums squared distances to centroids. Davies-Bouldin divides cluster radii by centroid separation. Calinski-Harabasz ratios between-cluster to within-cluster dispersion. Different formulas, one shared premise: **a cluster is a set of points that are mutually near each other and near their own mean**. That premise is a definition of a compact, convex, roughly isotropic blob. It is not a definition of a cluster. Groups defined by *connectivity* — you belong with me because there is a chain of near neighbours between us — can be arbitrarily elongated and curved, and then mutual nearness fails badly. ### The spiral case Take two long spiral arms that wind around each other, each one a genuine group. Consider a point halfway along one arm and evaluate the correct partition, the one that assigns each arm its own cluster: - **Cohesion is terrible.** The mean distance to the rest of its own arm includes the far end of the spiral, many turns away. That single term is large, and it is an average, so distant fellow members drag it up. - **Separation is terrible too.** The other spiral's arm may be winding past only a short step away. The nearest rival cluster contains points closer than most of the point's own cluster mates. So the cohesion term exceeds or nearly matches the separation term and the score is near zero or negative — for the structurally *correct* answer. Now score a partition that ignores the spirals and carves the plane into four compact round chunks. Every point is close to its chunk mates and the chunks have visible gaps between them, so cohesion is small, separation is large, and the average silhouette is comfortably positive. The index has ranked a geometrically clean but structurally meaningless partition above the true one, and it did so consistently with its own definition. It is not a bug; it is the criterion doing exactly what it says. ### The bias is shared, not silhouette-specific Inertia is minimised by exactly the round chunks, since that is what minimising squared distance to a centroid produces. Davies-Bouldin computes a radius around each centroid, which is meaningless for a curved arm whose centroid can fall in empty space between the coils. Calinski-Harabasz squares within-cluster deviations, which punishes long arms hardest of all. Because they fail in the same direction, their *agreement* on a curved dataset is worthless as corroboration — three instruments sharing one blind spot. ### The consequence that matters in interviews Do not use an internal index to choose between algorithms that assume different cluster shapes. A centroid-based method optimises very nearly the quantity these indices reward, so the comparison is loaded before it starts: the method whose objective matches the metric wins by construction, not by merit. Judging a connectivity- or density-based partition against silhouette is scoring one contestant with the other contestant's rulebook. ### What to do instead Three honest options. 1. **Change the distance, not the index.** Silhouette is defined for any distance function. Compute it under a graph or geodesic distance that travels along the data — then the far end of your own arm is *near* in the metric that matters, and the correct partition scores well. The cost is that you must build a neighbourhood graph, and the result depends on how you build it. 2. **Use a validity criterion designed for the shape.** Density-based validation indices exist for exactly this case; DBCV, for instance, scores clusters by density connectivity rather than by compactness around a mean. 3. **Judge from outside the geometry.** Evaluate against a held-out criterion that is not a distance summary at all — the downstream decision the clustering feeds, or expert judgement on the resulting groups. When geometry and purpose disagree, purpose wins. And whichever route you take, say what the number assumes when you report it. A silhouette of 0.61 is a statement about compactness under a chosen metric, not a certificate that the clusters are real.
- Do Davies-Bouldin, Calinski-Harabasz and inertia share this bias?Yes, and for the same reason. All three are built from centroids and mean or squared distances, so they reward compact, comparably-spread, well-separated groups. For a curved arm the centroid can sit in empty space that belongs to no cluster at all. Because they fail in the same direction, their agreement on such data corroborates nothing.
- How could silhouette be made to work for elongated clusters?Change the distance function rather than the index. Silhouette is defined over any distance, so computing it on a graph or geodesic distance that travels along the data makes the far end of the same arm close, and the correct partition then scores well. The cost is building a neighbourhood graph, and the answer depends on how that graph is built.
- A colleague compares a centroid-based and a density-based clustering by average silhouette, and the centroid one wins. What do you say?That the comparison is loaded. Silhouette rewards compactness around mutual nearness, which is essentially what the centroid method optimises, so it wins by construction rather than merit. Either score both under a distance suited to the structure, use a density-aware validity criterion, or judge both by the downstream decision the clustering exists to support.
A straight ruler will always report that a coiled rope is far from itself. The rope is not wrong; the ruler is measuring the wrong kind of closeness.
saying these in an interview costs you the question
- Treats a high silhouette as proof the clusters are real
- Compares algorithms with different shape assumptions by one internal index
- Believes only silhouette carries the compact-cluster bias
- Declares data unclusterable because the internal index is low
- Ignores that the index and the algorithm optimise the same quantity