How do you work out what a cluster of document embeddings actually represents?
answer
- group ids are not meaning
- read the members near the centre
- frequent inside, rare outside
- sample randomly, not just the core
- test the label on documents you did not read
basics
~20 sRead the documents nearest the cluster's centre, extract the terms that are common inside the cluster but rare outside it, and have a model draft a label from a random sample of members. Then check that label against members you did not look at.
solid answer
~50 sClustering gives you group ids, not meaning, so labelling is its own step with three complementary techniques. First, pull exemplars: the documents closest to the cluster centroid, which read as the prototypical cases. Second, compute distinctive terms by weighting each term's frequency inside the cluster against its frequency in the rest of the corpus, so you see what makes this cluster different rather than what is common everywhere. Third, hand a *random* sample of members to a model and ask for a short label plus a one-line description. Then verify: draw fresh members you did not use and check whether the label fits them. That last step catches the main failure, a cluster with two themes where the centroid neighbourhood shows only one — clustering 12,000 free-text statistics answers, a group labelled "confuses mean and median" may quietly also hold answers about skew, and only the held-out check reveals it.
code
python · 15 linesimport numpy as np
from sklearn.feature_extraction.text import TfidfVectorizer
docs = ["the mean is the same as the median",
"median and mean are interchangeable here",
"a p value is the chance the null is true",
"p value gives the probability the null hypothesis holds"]
labels = np.array([0, 0, 1, 1])
vec = TfidfVectorizer(stop_words="english").fit(docs)
terms = np.array(vec.get_feature_names_out())
mat = vec.transform(docs).toarray()
for c in (0, 1):
contrast = mat[labels == c].mean(axis=0) - mat[labels != c].mean(axis=0)
print(c, terms[np.argsort(contrast)[-3:]][::-1])go deeper
Know that the algorithm gives you numbers, not names, and be able to list the three moves: read the most central documents, pull the terms that distinguish the cluster, and draft a label from a sample.
Explain why raw term frequency returns stopwords and contrast weighting does not, and why a sample drawn only from the cluster core produces a label that misses part of the membership.
Show the verification discipline — held-out members, a recorded fit rate, cross-checking a generated label against distinctive terms — and name the multimodal-cluster failure this catches before it reaches a stakeholder.
Own what gets presented: which clusters are worth naming at all, how labels stay auditable with exemplars and terms attached, and how cluster identity is matched across re-runs so a shifting taxonomy is reviewed rather than quietly reissued.
## Why labelling is a separate problem A clustering algorithm returns integers. Turning cluster 7 into "students who treat a p-value as the probability the null hypothesis is true" is an interpretation step, and it is where most of the errors in an unsupervised analysis get introduced — because a plausible-sounding label is very easy to produce and very easy to believe. ## Technique 1: exemplars Rank the cluster's members by distance to the centroid and read the closest handful. Note that the centroid is an average vector and usually corresponds to no real document, so what you are reading are the members nearest that average — the medoid (the actual member with the smallest total distance to the others) is the closest thing to a canonical representative. Exemplars give you the prototype fast, but they are systematically unrepresentative: they show the cluster's core and hide its periphery. Always pair them with a sample drawn from the middle and outer bands of the distance distribution. ## Technique 2: distinctive terms Raw term frequency inside a cluster mostly returns stopwords and corpus-wide vocabulary. What you want is contrast: terms frequent *inside* the cluster and rare *outside* it. The standard move is a TF-IDF-style weighting where the cluster is treated as one document and the rest of the corpus as the background — the class-based TF-IDF used by topic-modelling tooling is exactly this idea. Subtracting the mean term weights of non-members from the mean term weights of members gives a usable ranking in a few lines. Distinctive terms are cheap, deterministic and auditable, which makes them a good cross-check on any generated label: if the model called the cluster "login problems" and the distinctive terms are about refunds, you have caught a bad label without reading anything. ## Technique 3: model-generated labels Sample members, present them, and ask for a short noun-phrase label plus a one-sentence description. Three details matter. Sample **randomly** rather than taking the nearest-centroid documents, or the label will describe the core and miss the rest. Ask for a description alongside the label, because a description that hedges ("various answers about statistics") is itself a signal that the cluster is not coherent. And label clusters with the other clusters' labels visible where you can, so the labels come out mutually distinct rather than four variations on the same phrase. ## Verifying the label This is the step people skip. Hold out members that were not shown during labelling, then check whether the label applies to them — a human can do this on 20 documents in a few minutes, or you can ask a model, one document at a time, whether it fits the label, and record the hit rate. A cluster where the label fits 90% of held-out members is a real finding you can present. One where it fits 55% is two clusters wearing one name, and the fix is to split it, raise k, or report it as a heterogeneous bucket. The same check catches the most damaging failure mode: a cluster that is multimodal. Consider clustering 12,000 free-text answers from an introductory statistics course, hoping to surface misconceptions the rubric missed. One dense group's nearest-centroid answers all conflate the mean with the median, so "mean/median confusion" is the obvious label — but half the held-out members turn out to be answers about skewed distributions that merely share the vocabulary. Presenting the first label to the teaching team would have sent them to fix the wrong thing. ## Practical notes Name the clusters that matter and leave the rest unnamed. In a long-tailed corpus a handful of clusters carry most of the mass and most of the actionable signal; forcing a crisp label onto a diffuse leftover group manufactures a category that does not exist. Keep the exemplars and distinctive terms alongside the label wherever the results are shown, so a reader can audit the interpretation instead of trusting it. And if you re-run the clustering on new data, do not assume cluster ids carry over — labels must be re-derived or explicitly matched, because the numbering is arbitrary from run to run.
- Why sample randomly instead of just using the ten documents nearest the centroid?Because centre-nearest members are the prototype, not the cluster. A cluster can hold a second theme at its periphery that never appears in the top ten, and a label drafted from the core will silently exclude it. Random sampling across the whole membership, or stratified sampling across distance bands, exposes the spread — and if the sample looks incoherent, that is a finding about the clustering rather than a labelling problem.
- How would you check that a generated label is actually right?Hold out members the labeller never saw, then test the label against them one document at a time and record the fit rate. High agreement means the label generalizes; around half means the cluster is heterogeneous and should be split or reported as mixed. Cross-checking the label against the cluster's distinctive terms is a cheap second signal that catches confidently wrong labels.
- Do cluster labels carry over when you re-run the pipeline next month?No — cluster ids are arbitrary and unstable across runs, seeds and data refreshes, so label 3 this month may be a different group entirely. If you need continuity, match clusters across runs explicitly, for example by centroid similarity or by overlap of shared documents, and re-derive labels rather than assuming they persist. Treat any change in the label set as something to review, not to hide.
saying these in an interview costs you the question
- Treats the centroid as if it were a real document
- Ranks terms by raw frequency and reports stopwords
- Labels from the five most typical members and stops there
- Never checks the label against unseen members of the cluster
- Forces a crisp name onto a diffuse leftover cluster