skip to content

In scikit-learn, when do you use silhouette_score versus adjusted_rand_score?

level: middleimportance: should knowfreq 32%

answer

  1. one needs the data, one needs truth
  2. first argument tells them apart
  3. corrected for chance versus not
  4. permuting cluster ids changes nothing
  5. noise labels are counted as a cluster

basics

~20 s

silhouette_score(X, labels) is internal: it scores cluster shape from the feature matrix alone, so it works without ground truth. adjusted_rand_score(labels_true, labels_pred) is external: it compares two labelings and requires true labels you usually do not have.

solid answer

~50 s

The split is internal versus external, and it shows up directly in the signatures. `silhouette_score(X, labels)` takes the **feature matrix plus the cluster assignment** and measures, per point, how much closer it sits to its own cluster than to the nearest other one, averaged over points and bounded in [-1, 1]. It needs the same `metric` you clustered with, raises if there is only one cluster or as many clusters as samples, and accepts `sample_size` to subsample because it is quadratic in points. `adjusted_rand_score(labels_true, labels_pred)` takes **two label vectors** and measures agreement over pairs of points; it is invariant to label permutation — cluster 0 versus cluster 2 does not matter — and corrected for chance, so random labelings score near 0 and can go slightly negative, unlike the uncorrected `rand_score`. Two traps: DBSCAN's -1 noise label is treated as an ordinary cluster by both, and `davies_bouldin_score` is the internal metric where lower is better.

code

python · 12 lines
python
from sklearn.cluster import KMeans
from sklearn.datasets import make_blobs
from sklearn.metrics import (adjusted_rand_score, rand_score,
                             silhouette_score, davies_bouldin_score)

X, y = make_blobs(n_samples=300, centers=3, random_state=0)
labels = KMeans(n_clusters=3, n_init=10, random_state=0).fit_predict(X)

print(silhouette_score(X, labels))
print(davies_bouldin_score(X, labels))
print(adjusted_rand_score(y, labels))
print(rand_score(y, labels))

go deeper

for a junior

Know that silhouette_score needs the feature matrix and the cluster labels while adjusted_rand_score needs true labels and predicted labels, and that cluster ids are arbitrary so accuracy makes no sense here.

for a middle

Explain what the silhouette value means per point and why it is bounded in [-1, 1], and why the Rand index needs chance correction to be comparable across different numbers of clusters.

for a senior

Show judgment about metric choice for the algorithm at hand: silhouette's convex-blob bias against density-based clustering, the -1 noise-label distortion, subsampling for quadratic cost, and why per-sample silhouettes beat the mean.

for a principal

Decide how an unsupervised result is signed off at all — which internal metric plus which downstream or human evaluation counts as evidence, since no internal score alone establishes that a partition is useful to the business.

## Internal versus external Every clustering metric in `sklearn.metrics` falls into one of two families, and confusing them is the source of most of the mistakes here. **Internal (unsupervised) metrics** judge a partition using only the data and the assignment. They answer 'is this partition geometrically tidy?'. `silhouette_score`, `calinski_harabasz_score` and `davies_bouldin_score` are the three scikit-learn ships. **External (supervised) metrics** compare a partition to a reference labeling. They answer 'does this partition agree with what we already know?'. `adjusted_rand_score`, `rand_score`, `normalized_mutual_info_score`, `adjusted_mutual_info_score`, and the `homogeneity_score` / `completeness_score` / `v_measure_score` trio live here. The practical rule follows from what you have: with ground truth, use an external metric — and then ask why you are clustering rather than classifying. Without it, you are limited to internal metrics, which is the normal case. ## silhouette_score in detail The signature is `silhouette_score(X, labels, *, metric='euclidean', sample_size=None, random_state=None, **kwds)`. For each sample it computes `a`, the mean distance to other points in its own cluster, and `b`, the mean distance to points in the nearest other cluster, then `(b - a) / max(a, b)`. Values near 1 mean the point sits comfortably inside a well-separated cluster; near 0 it sits on a boundary; negative means it is closer on average to another cluster than its own. The function returns the mean over all samples; `silhouette_samples` returns the per-point array, which is far more informative because the mean hides a single bad cluster. Three operational details matter. First, `metric` must match the geometry you clustered under — silhouette on Euclidean distance while you clustered with cosine distance measures something you did not optimise, and `metric='precomputed'` lets you pass a distance matrix instead of `X`. Second, the computation is pairwise and therefore quadratic in the number of samples; `sample_size` with `random_state` subsamples for tractability on large data. Third, it raises a `ValueError` unless the number of distinct labels is at least 2 and at most `n_samples - 1`, so a degenerate clustering that collapsed to one cluster fails loudly rather than scoring badly. A structural bias is worth stating: silhouette rewards convex, roughly equal-sized, well-separated blobs. Density-based clusterings with elongated or nested shapes score poorly even when they are correct, so it is a weak arbiter for DBSCAN-style output. ## adjusted_rand_score in detail `adjusted_rand_score(labels_true, labels_pred)` counts, over all pairs of samples, how often the two labelings agree about whether the pair belongs together, then corrects that agreement for what chance alone would produce. The correction is the whole point. The uncorrected `rand_score` is bounded in [0, 1] but drifts high simply because most random pairs are in different clusters — with many clusters it can read 0.9 for a meaningless partition. The adjusted version has expected value 0 for independent random labelings, 1 for identical partitions up to relabeling, and can dip slightly below 0 for worse-than-chance agreement. Because it works on pairs, it is invariant to permuting cluster ids and it does not require the two labelings to have the same number of clusters. That is what makes it usable at all: `KMeans` labels are arbitrary integers with no correspondence to your true class ids, so any metric that compared them element-wise — accuracy, say — would be meaningless. ## The traps **Argument order and type.** `silhouette_score` takes `X` first; `adjusted_rand_score` takes two label vectors. Passing labels where `X` is expected does not necessarily raise — a 1-D array can be reshaped or silently treated as a single feature in some pipelines — and passing `X` to an external metric raises confusingly. Watch the first positional argument. **DBSCAN noise.** Points DBSCAN cannot assign get label -1. Both metric families treat -1 as an ordinary cluster, so a scattered noise set is scored as if it were a real, terribly-shaped cluster, which drags silhouette down and distorts ARI. Filter noise points out, or report the metric on the clustered subset and the noise fraction separately. **Direction.** `silhouette_score` and `calinski_harabasz_score` are higher-is-better; `davies_bouldin_score` is lower-is-better. If you wrap the last one for any automated comparison, it needs the same negation treatment that error metrics get. **Choosing k.** Sweeping the number of clusters and picking the peak silhouette is a legitimate heuristic, but it is a heuristic: silhouette's convex-blob bias means the peak may not be the number of groups that matters to the problem, and the per-sample distribution from `silhouette_samples` is the more honest artefact to look at.

  • Why is adjusted_rand_score preferred over rand_score?
    The plain Rand index counts pair agreements without asking how many you would get by luck. With many clusters, most random pairs land in different clusters and agree trivially, so rand_score reads high for meaningless partitions. The adjusted form subtracts the expected agreement, giving roughly 0 for independent labelings and 1 for a perfect match, which makes values comparable across different numbers of clusters.
  • Your DBSCAN clustering scores a poor silhouette. What should you check before concluding it failed?
    Two things. First, the -1 noise label is treated as a real cluster, so unassigned points are scored as a diffuse cluster and drag the mean down — recompute on the non-noise subset and report the noise fraction separately. Second, silhouette rewards convex, similarly sized, well-separated blobs, so correct elongated or nested density clusters score badly by construction.
  • Which arguments does silhouette_score need beyond X and the labels?
    `metric` must match the distance used for clustering — Euclidean by default, or `'precomputed'` if you pass a distance matrix instead of the feature matrix. `sample_size` with `random_state` subsamples the data, which matters because the computation is quadratic in samples. The call also raises unless the number of distinct labels is between 2 and n_samples - 1.

saying these in an interview costs you the question

  • Passing cluster labels where silhouette_score expects the feature matrix
  • Comparing cluster ids to true labels with accuracy_score
  • Trusting rand_score without chance correction
  • Scoring DBSCAN output without excluding the -1 noise label
  • Assuming davies_bouldin_score is higher-is-better like silhouette

context