skip to content

Silhouette scores stay flat from k=2 to k=30 on document embeddings — what does that mean?

level: seniorimportance: should knowfreq 44%

answer

  1. no scale improves the contrast
  2. argmax of noise is not a choice
  3. internal indices never see the outside world
  4. hand-label a sample, then score agreement
  5. does it survive a resample?

basics

~20 s

A flat curve means no value of k gives a materially better partition than any other: the corpus behaves more like a continuum than a set of separated groups. Do not take the argmax of noise — validate against a held-out labelling or a downstream task instead.

solid answer

~50 s

Silhouette compares, for each point, how close it sits to its own cluster versus the nearest other cluster, so a flat sweep says the data offers no scale at which that contrast improves. Two things can produce it: the corpus genuinely has no clean partition — text embeddings of a broad corpus often form overlapping continua rather than islands — or the algorithm's assumptions do not match the structure, for example spherical k-means on elongated or nested groups. Absolute values run low on high-dimensional text vectors, so judge relatively, not against a textbook threshold. The response is to stop asking the internal index and get external evidence: hand-label a sample of a few hundred documents and score the clustering against it with adjusted Rand index or normalized mutual information, test stability by re-clustering bootstrap resamples and measuring agreement, and check whether a reviewer finds the clusters actionable. If nothing supports a partition, say so and switch to nearest-neighbour retrieval or a hierarchy a human can cut.

code

python · 8 lines
python
import numpy as np
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

emb = np.random.rand(1000, 64)  # stand-in for document embeddings
for k in range(2, 31):
    labels = KMeans(n_clusters=k, n_init=10, random_state=0).fit_predict(emb)
    print(k, round(silhouette_score(emb, labels, metric="cosine"), 3))

go deeper

for a junior

Know what the number means: silhouette compares how close a point sits to its own cluster versus the nearest other one, roughly between -1 and 1, and around 0 means the point is on a boundary.

for a middle

Be ready to explain the sweep and its limits — that internal indices only see the vectors and the metric you gave them, that absolute values run low on text embeddings, and that a flat curve is a statement about the data rather than a bug.

for a senior

Demonstrate the escape route: external validation against a hand-labelled sample with chance-corrected agreement measures, stability under resampling, and downstream utility. Also show you would re-run with a different clustering family before declaring the corpus structureless.

for a principal

Own the decision of what to do with an unsupported partition — reporting 'there is no natural k' rather than shipping one, choosing retrieval or a human-cut hierarchy instead, and weighing whether the real fix is a domain-adapted representation rather than more clustering hyperparameters.

## What silhouette actually measures For each point, let *a* be its mean distance to the other members of its own cluster and *b* its mean distance to the members of the nearest *other* cluster. The point's silhouette is (b − a) / max(a, b), which lands in [−1, 1]: near 1 means it sits comfortably inside its own group, near 0 means it is on the boundary between two, and negative means it is closer on average to a neighbouring cluster than to its own. The reported score is the mean over all points, so it is a compactness-versus-separation summary that assumes the distance metric you used is the right one. ## Reading a flat sweep A sweep from k=2 to k=30 that hovers in a narrow band with no meaningful peak carries a real message: at no granularity does the partition buy you separation. Two distinct causes produce the same picture. **The corpus has no partition structure.** Broad text corpora frequently form overlapping continua — support tickets shade from billing into account access into refunds, and there is no distance at which those regions detach. Silhouette is behaving correctly; it is telling you the shape of the data. **The algorithm's model does not fit.** Silhouette computed on a k-means partition inherits k-means' assumptions. Elongated, nested or wildly unequal clusters can be real yet score poorly under a mean-distance criterion. Before concluding the data is structureless, re-run the sweep with a different clustering family — hierarchical with average or Ward linkage, or a density method with a minimum-cluster-size sweep — and see whether the flatness persists. A related trap is the absolute scale. On high-dimensional text embeddings, mean silhouette values in the 0.05–0.2 range are ordinary even for partitions that human reviewers find useful, because the spread between within- and between-cluster distances is compressed. Anyone quoting "below 0.5 means bad clustering" is importing a threshold from low-dimensional tabular data. ## Why other internal indices rarely rescue you Davies–Bouldin (lower is better), Calinski–Harabasz (higher is better), inertia with the elbow heuristic and the gap statistic are all worth computing — they are cheap — but they measure variations on the same compactness/separation theme under similar geometric assumptions. When silhouette is flat they usually agree, and when they disagree you have no principled way to arbitrate. Treat unanimity as mild confirmation and disagreement as a signal that the structure is weak either way. Internal indices are weak evidence by construction: they never see anything outside the vectors and the metric you handed them. ## Getting external evidence This is where a stalled analysis actually gets unstuck. - **Held-out human labelling.** Sample 200–400 documents and have someone assign them to categories without seeing the clusters. Then compare the cluster assignments to those labels with the **adjusted Rand index** or **normalized mutual information**, both of which correct for chance agreement, and look at per-cluster purity. This answers a question no internal index can: do the machine's groups correspond to distinctions a person cares about? It can also reveal the opposite finding — that the human labels themselves do not separate in the embedding space, which is a legitimate and useful result. - **Stability selection.** Draw bootstrap or subsample replicates of the corpus, re-cluster each at the same k, and measure agreement on the overlapping points (adjusted Rand index between runs works well). Structure that is real tends to survive resampling; a k whose assignments reshuffle every replicate is not supported by the data, whatever any index says. Stability is often the single most informative check when internal indices are uninformative. - **Downstream utility.** Ultimately clusters exist to serve something — a triage queue, a taxonomy, a feature in a model, a report a human reads. Score them by that: does a reviewer find the clusters nameable and distinct? Does adding the cluster id as a feature move the target metric? A partition that is mediocre by silhouette but obviously useful to the team is a good partition. ## What to do when nothing supports a partition Accept the finding rather than shopping for a k that looks defensible on a chart. Practical exits: - **Drop the partition.** If the corpus is a continuum, nearest-neighbour retrieval serves most of the use cases people wanted clustering for, and it makes no false claim of discrete categories. - **Go hierarchical and let a human cut it.** A dendrogram lets a domain expert choose granularity where it is meaningful, which is more honest than an automatically selected k. - **Use density clustering and keep the noise.** Extracting a handful of genuinely dense pockets and admitting the rest is unstructured is often the true answer. - **Seed the clustering.** If you have a few labelled anchors, constrained or semi-supervised clustering can align groups with the distinctions you care about instead of the dominant variance in the space. - **Change the representation.** A general-purpose embedding model may not separate your domain's distinctions at all; a domain-adapted model, or embedding a normalized field rather than raw free text, can change the geometry more than any clustering hyperparameter will. And report the flat sweep itself. "There is no natural number of clusters here" is a finding, and stating it beats presenting a k chosen from noise as though it were discovered.

  • What silhouette value counts as good on text embeddings?
    There is no fixed threshold, and quoting one is a mistake. On high-dimensional text vectors the within- and between-cluster distances sit close together, so mean silhouettes around 0.1 can accompany perfectly usable clusters. Use it comparatively — across k, across algorithms, and against a randomly assigned baseline on the same data — rather than against an absolute bar borrowed from low-dimensional examples.
  • How exactly would you use a held-out labelling to judge the clustering?
    Sample a few hundred documents and have a person categorize them blind to the clusters. Then compute adjusted Rand index or normalized mutual information between the cluster ids and those labels — both correct for chance — and inspect per-cluster purity to see which clusters map cleanly and which mix categories. Low agreement is informative twice over: either the clusters are not meaningful, or the embedding does not encode the distinction the humans are making.
  • Re-clustering a bootstrap resample gives quite different assignments. What do you conclude?
    That the chosen k is not supported by the data. Genuine structure survives resampling — the same documents keep landing together — so high run-to-run disagreement on shared points means the partition is being imposed rather than found. I would sweep k for stability as well as for silhouette and prefer a k where agreement is high, or accept that no k is stable and abandon the partition.

saying these in an interview costs you the question

  • Picks the k with the highest silhouette even when the curve is flat
  • Quotes a universal 'good silhouette' threshold for any dataset
  • Treats internal indices as validation of meaning rather than geometry
  • Never checks whether clusters survive a different seed or resample
  • Assumes a flat curve means the code is broken rather than a real finding

context