skip to content

Silhouette and Internal Indices

Silhouette contrasts a point's own-cluster distance with its nearest rival, while Davies-Bouldin and Calinski-Harabasz score compactness against separation. All of them quietly favour round blobs.

on this pageshow

questions

5

How is the silhouette coefficient computed for a single point, and what does a negative value mean?

level: middleimportance: must knowfreq 74%

answer

  1. cohesion versus separation
  2. own cluster versus nearest rival cluster
  3. difference divided by the larger term
  4. bounded between minus one and one
  5. below zero means the neighbour fits better

basics

~20 s

For a point, a is its mean distance to the other members of its own cluster and b its mean distance to the points of the nearest other cluster; s = (b - a) / max(a, b). Negative s means the point sits closer to that other cluster.

solid answer

~50 s

Silhouette is a per-point score computed from distances alone. For point i, `a(i)` is the mean distance from i to every other point in its own cluster (cohesion), and `b(i)` is the smallest, over all other clusters, of the mean distance from i to that cluster's points (separation to the nearest rival). Then `s(i) = (b(i) - a(i)) / max(a(i), b(i))`, which is bounded in [-1, 1]. Near 1 the point is much closer to its own cluster than to any other; near 0 it sits on the boundary between two clusters; below 0 it is on average nearer another cluster's points than its own, so the assignment is questionable. A point alone in its cluster is assigned 0 by convention. Averaging s(i) over a cluster gives that cluster's score, and over all points the overall silhouette.

code

python · 13 lines
python
from math import dist

# two small clusters of 2-D points; we score the first point of A
A = [(0.0, 0.0), (1.0, 0.0), (0.0, 1.0)]
B = [(4.0, 0.0), (5.0, 1.0)]

i = A[0]
a = sum(dist(i, p) for p in A[1:]) / (len(A) - 1)   # cohesion: own cluster
b = sum(dist(i, p) for p in B) / len(B)              # separation: nearest rival
s = (b - a) / max(a, b)

print(round(a, 3), round(b, 3), round(s, 3))
# 1.0 4.55 0.78  -> comfortably inside its own cluster

go deeper

for a junior

Be ready to say what the score ranges over and which direction is good: minus one to plus one, higher is better, and a value near zero means the point sits on a boundary between two clusters.

for a middle

You are expected to write the formula out and name both terms correctly, especially that the separation term uses the single nearest other cluster rather than an average over all of them.

for a senior

Show that you treat the number as evidence, not a verdict: state the distance function and the scaling it was computed under, look at the per-cluster and per-point spread, and mention the quadratic cost on large data.

for a principal

Own the question of what the team reports. Decide whether a single headline silhouette is worth publishing at all, given how easily it hides a broken cluster and how little it transfers across differently prepared feature sets.

### What the coefficient measures Silhouette is an *internal* validity measure: it judges a partition using only the data and the distances between points, with no ground-truth labels anywhere. It asks one question per point: given where you were assigned, would you have been happier next door? For a point `i` assigned to cluster `C`: - `a(i)` = the mean distance from `i` to every **other** point of `C`. This is the cohesion term: how tightly `i` sits inside its own cluster. Note it is a mean over the other members, not a distance to a centroid, so the cluster never has to be summarised by a mean vector. - `b(i)` = `min` over every other cluster `C'` of the mean distance from `i` to all points of `C'`. This is the separation term, and the minimum matters: `b(i)` is the *nearest rival* cluster, not the average of all the others. That nearest rival is the cluster `i` would most plausibly have been put in. - `s(i) = (b(i) - a(i)) / max(a(i), b(i))`. The numerator is the gain from staying home rather than moving to the best alternative. Dividing by the larger of the two keeps the result inside `[-1, 1]` whichever term dominates, and makes the score dimensionless: it is a *relative* comparison, not a distance. ### Reading the value - `s(i)` close to `+1`: `a(i)` is much smaller than `b(i)`. The point is deep inside a cluster that is well separated from its neighbour. - `s(i)` close to `0`: `a(i)` and `b(i)` are about equal. The point lies on the frontier between two clusters; the assignment is essentially a coin flip. - `s(i) < 0`: `a(i) > b(i)`. On average the point is nearer the members of another cluster than its own. This is not an outlier diagnosis, it is a *misassignment* diagnosis, and it names the cluster it would prefer. - A point that is the only member of its cluster has no `a(i)` to compute, so the convention is `s(i) = 0`. That convention matters in practice: singleton or near-singleton clusters are neither rewarded nor punished, so they quietly sit in the middle of any average you take. Rough rules of thumb circulate for the overall average — around 0.5 and above suggests reasonably distinct structure, around 0.25 and below suggests the partition is mostly imposed rather than found — but they are only rules of thumb. The number depends on the distance function, the feature scaling and the dimensionality, so an absolute threshold copied from a textbook proves nothing about your table. ### Aggregation levels There are three levels, and they carry different information. Per point, it is a diagnostic you can attach to individual rows. Per cluster (mean of the `s(i)` inside one cluster) it tells you which parts of a partition are solid and which are mush. Over the whole dataset it is a single summary — convenient to report, and the level at which the most information is destroyed, because one large, clean cluster can carry the mean over a broken one. ### What it is sensitive to Silhouette is a function of the distance matrix, so anything that changes distances changes the score. Multiplying one feature by 100 makes it dominate `a` and `b` and effectively re-clusters the space; standardising or otherwise putting features on a comparable footing is a decision you make before scoring, and the score is only comparable across runs that share that decision. Swapping Euclidean for cosine or Manhattan distance likewise gives a different, incomparable number. Always report the distance and the preprocessing next to the value. It is also expensive. In the general form it needs all pairwise distances, so time and memory grow with the square of the number of points. On large data the usual answer is to compute it on a random subsample and state the sample size, or to fall back on a centroid-based index whose cost is linear in the number of points. ### What it does not tell you A high silhouette says the partition is compact and separated *under the distance you chose*. It does not say the clusters mean anything, does not say they are stable if you resample, and it quietly assumes clusters are the kind of round, comparably-spread blobs that distances to fellow members can describe. Those are separate questions with separate tools.

  • What score does a point get if it is the only member of its cluster?
    Zero, by convention. There are no other members, so the cohesion term is undefined and the definition assigns 0 rather than leaving a hole. The practical consequence is that singleton clusters are neither rewarded nor penalised: they drag any average toward the middle and can disguise a partition that has shattered into fragments.
  • How does feature scaling change a silhouette score?
    Completely. The score is a function of distances, so a feature measured in the thousands dominates both the cohesion and the separation term while a feature in [0, 1] barely registers. Standardise or otherwise equalise the features before scoring, and report the preprocessing and the distance function alongside the number, because scores computed under different geometry are not comparable.
  • Why is silhouette expensive on large datasets, and what do practitioners do?
    The definition needs the mean distance from each point to every other point, so cost grows with the square of the number of rows in both time and memory. The usual workarounds are to score a random subsample and state its size, or to switch to a centroid-based index whose cost is linear in the number of points.

It is like asking someone at a party whether they are standing closer to their own group or to the next conversation over. A negative answer means they have already drifted into the other circle.

saying these in an interview costs you the question

  • Says the silhouette coefficient ranges from 0 to 1
  • Reads a negative score as an outlier rather than a misassignment
  • Defines b as the mean distance to all other clusters, not the nearest one
  • Compares scores computed under different scalings or distance functions
  • Reports one overall average as if every cluster scored the same

context

open as a page

Why is it wrong to compare k-means inertia between two runs built on different feature sets?

level: juniorimportance: should knowfreq 41%

basics

~20 s

Inertia is the total within-cluster sum of squared distances from points to their cluster centroid. It is an unnormalised quantity whose magnitude grows with feature count, feature scale and row count, so two runs over different feature sets are simply on different scales.

open as a page

Why does average silhouette rank a round four-way split of two interleaved spirals above the correct partition?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Silhouette rewards points near their own cluster mates and far from the nearest other cluster. Along a spiral the far end of your own arm is distant while a neighbouring arm is close, so compact round chunks win.

open as a page

On a silhouette plot the average is 0.42 but one cluster's bars are negative and another's are all short — what does that tell you?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The negative-bar cluster holds points that are on average closer to another cluster, so it overlaps a neighbour. The short-bar cluster sits on a boundary with no space of its own. A respectable-looking average of 0.42 is being carried by the healthy clusters.

open as a page

What do the Davies-Bouldin and Calinski-Harabasz indices measure, and why can they rank two partitions oppositely?

level: middleimportance: nice to knowfreq 31%

basics

~20 s

Davies-Bouldin averages, over clusters, the worst ratio of two clusters' spreads to the distance between their centroids, and lower is better. Calinski-Harabasz is a between-cluster over within-cluster dispersion ratio, and higher is better. Different aggregations can rank partitions oppositely.

open as a page