skip to content

Clustering & Visualization

Using embeddings without a query: k-means or DBSCAN to group documents, UMAP or t-SNE to project them down to two dimensions, and silhouette scores to judge whether the clusters mean anything. Interviewers ask how much you can safely conclude from a 2D picture of a 1536-dimensional space.

on this pageshow

questions

4

What can a UMAP or t-SNE plot of embeddings actually tell you?

level: middleimportance: must knowfreq 66%

answer

  1. local structure, not global geometry
  2. the empty space is not a measurement
  3. axes have no units at all
  4. perplexity and neighbour count reshape the picture
  5. re-run the seed before you believe it

basics

~20 s

UMAP and t-SNE preserve which points sit near each other locally, so tight visible groups usually reflect real neighbourhoods. The gaps between blobs, the relative blob sizes and the axes carry no reliable meaning and shift with hyperparameters and seed.

solid answer

~50 s

These are neighbour-embedding methods: they optimize a layout so that each point's near neighbours in the original space stay near in two dimensions. What survives is local structure, so a tight visible group is usually a real one. What does not survive is everything quantitative — the distance between two blobs, their relative sizes and densities, and the axes, which have no units and can be rotated or reflected freely. Both are stochastic and hyperparameter-sensitive: t-SNE's perplexity and UMAP's neighbour count and minimum distance change how much local versus global structure the layout tries to keep, and a different seed gives a different picture. Take a UMAP of a music catalogue's track descriptions: neighbourhoods of similar tracks are trustworthy, but two genre blobs sitting far apart says nothing about how different those genres are. Use the plot to communicate and generate hypotheses, and do the clustering and measurement in the original space.

go deeper

for a junior

Know the headline rule: nearby points on the plot are genuinely similar, but the distance between groups, the size of the blobs and the axes mean nothing. Never quote a gap on the chart as a number.

for a middle

Be ready to explain why — the objective only rewards preserving each point's neighbours — and to name the knobs that reshape the picture, such as t-SNE's perplexity and UMAP's neighbour count, plus the fact that both are stochastic.

for a senior

Demonstrate the working discipline: cluster and measure in the original space, use the projection only as a display coloured by independently computed labels, and check that any structure you report survives new seeds and settings.

for a principal

Own how these pictures are used in decisions. Set the norm that a projection is an illustration, never evidence, insist claims come with a measurement in the original space, and be ready to say out loud that a striking chart shows nothing.

## What these methods are doing t-SNE and UMAP are *neighbour embedding* methods. Both start by describing, for every point, which other points are its neighbours and how strongly — t-SNE converts distances into probabilities with a bandwidth set by the **perplexity** parameter, UMAP builds a weighted neighbour graph with `n_neighbors` edges per point. Then both search for a 2D (or 3D) layout in which that neighbour structure is reproduced as closely as possible, optimizing by gradient descent from a random or spectral initialization. Nothing in that objective asks the layout to preserve absolute distance, scale, or orientation. The optimizer is free to place a well-separated group anywhere on the canvas as long as its members stay together. That single fact explains almost every misreading of these plots. ## What survives the projection - **Local neighbourhoods.** If two documents are near each other in the plot, they were usually near each other in the embedding space. This is the property the objective actually rewards. - **The existence of well-separated groups.** If the corpus really contains disjoint dense regions, both methods will typically show them as separate blobs. - **A qualitative sense of continuity.** A long smear rather than discrete blobs is a genuine signal that the corpus is a continuum, not a set of clean categories. ## What does not survive - **Distance between clusters.** The empty space between two blobs is not a measure of how different they are. Two blobs at opposite ends of the canvas may be closer in the original space than two adjacent ones. t-SNE is worse at this than UMAP, but UMAP's global geometry is also not metric — treat it as suggestive at best. - **Cluster size and density.** t-SNE in particular equalizes apparent density: a tight cluster of 50 near-duplicates and a diffuse cluster of 5,000 documents can be drawn at similar sizes. Do not read "bigger blob" as "more documents" or "broader topic". - **Axes.** The horizontal and vertical directions have no units and no interpretation. The whole layout can be rotated, flipped or translated with no change in objective value, so "this group is to the left" is not a finding. - **Certainty that a visible group is real.** At small perplexity or small `n_neighbors`, both methods can shatter genuinely uniform data into visually convincing islands. Apparent clusters in noise are a well-known failure of these plots. ## Hyperparameters and stochasticity t-SNE's perplexity roughly sets the effective number of neighbours each point is balanced against; small values emphasize very local structure and fragment the picture, large values smooth it and merge groups. UMAP's `n_neighbors` plays the analogous role, while `min_dist` controls only how tightly points may be packed visually — lowering it makes blobs look crisper without changing what was learned. Both use random initialization and stochastic optimization, so re-running with a new seed changes the layout. The operational consequence: never report structure you have seen in exactly one run at one setting. Re-run across two or three seeds and a couple of neighbour settings. Structure that persists is worth investigating; structure that appears in one panel and vanishes in the next is an artefact of the projection. ## How to use the plot honestly The reliable workflow inverts the naive one. Do the analysis in the original space — cluster there, or compute nearest neighbours there — and use the projection purely as a **display surface**, colouring points by labels or metadata you computed independently. If the independently derived clusters land as coherent regions on the plot, that is mild corroboration; if they are shredded across the canvas, that is a prompt to look harder, not proof either way. A concrete case: a UMAP of a music catalogue's track descriptions. Zooming in, the local neighbourhoods are convincing — acoustic singer-songwriter tracks sit with acoustic singer-songwriter tracks, and pulling up the ten nearest points to any track gives sensible results you can verify by reading them. But the layout also shows two large blobs at opposite corners with a wide void between them, and stakeholders will inevitably ask what the void means. It means nothing. It is the optimizer having found a low-cost arrangement, and a different seed may place those blobs side by side. The verifiable claims from that plot are all local; the eye-catching global ones are not claims at all. ## Alternatives when you need evidence rather than a picture If the question is "are these two groups distinct?", answer it with numbers in the original space: nearest-neighbour purity, the fraction of each group's k nearest neighbours that share its label, or an internal cluster index computed on the full-dimensional vectors. If the question is "what is near this item?", print the ranked neighbour list — it is more informative than the picture and cannot be misread. Save the projection for the slide, and label it as an illustration.

  • A stakeholder points at two far-apart blobs and asks how different those groups are. What do you say?
    That the plot cannot answer it. Between-cluster distance is not preserved by either method, and the layout can be rotated or rearranged without changing its objective. I would answer with a measurement in the original space instead — for example the mean cosine similarity between the two groups' centroids, or how often members of one group appear among the other's nearest neighbours — and show the plot only as an illustration.
  • How would you tell an artefact cluster from a real one?
    Perturb the things that should not matter. Re-run with several random seeds and two or three neighbour or perplexity settings; a real dense region keeps its membership, an artefact dissolves or reshuffles. Then verify in the original space: take the candidate group's members and check whether they are genuinely each other's nearest neighbours there, and read a sample to see if they share a theme.
  • Does colouring the plot by cluster labels prove the clustering is good?
    No, and it is a common trap, especially if the projection ran on the same vectors. Clean-looking colours can come from the projection's own tendency to separate groups, and messy colours can hide a perfectly serviceable partition. Judge the clustering with quantitative validation on the full-dimensional data or against a held-out human labelling; the coloured plot is communication, not evidence.

A subway map is drawn to show which stops connect to which, not how far apart they really are — you can trust the adjacency and not the geography. A t-SNE or UMAP plot is the same bargain: neighbourhoods are meaningful, the map distances are not.

saying these in an interview costs you the question

  • Reads the gap between two blobs as a similarity measurement
  • Interprets the axes as if they were latent factors
  • Judges cluster quality by how clean the picture looks
  • Treats one run at default settings as the definitive layout
  • Assumes bigger blobs contain more documents

context

open as a page

How do you choose between k-means and DBSCAN for clustering document embeddings?

level: middleimportance: must knowfreq 62%

basics

~20 s

k-means forces every document into one of k clusters you fix in advance, so it fits a corpus you want fully partitioned. DBSCAN groups by density, discovers how many clusters exist, and leaves sparse points unlabelled as noise.

open as a page

How do you work out what a cluster of document embeddings actually represents?

level: juniorimportance: should knowfreq 38%

basics

~20 s

Read the documents nearest the cluster's centre, extract the terms that are common inside the cluster but rare outside it, and have a model draft a label from a random sample of members. Then check that label against members you did not look at.

open as a page

Silhouette scores stay flat from k=2 to k=30 on document embeddings — what does that mean?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A flat curve means no value of k gives a materially better partition than any other: the corpus behaves more like a continuum than a set of separated groups. Do not take the argmax of noise — validate against a held-out labelling or a downstream task instead.

open as a page