On a t-SNE map, why is the gap between two clusters not a real distance?
answer
- local neighbourhoods, not global geometry
- the divergence is weighted asymmetrically
- heavy-tailed map kernel spreads non-neighbours
- per-point bandwidth normalises density away
- rerun with a new seed, gap moves
basics
~20 st-SNE only preserves each point's near neighbours. It minimises a divergence that punishes tearing neighbours apart but barely punishes moving distant points, so between-cluster gaps and island sizes are artefacts of the layout, not measured distances.
solid answer
~50 st-SNE turns data-space distances into neighbour probabilities and then moves 2-D points to minimise `KL(P || Q)`. That cost is weighted by the data-space similarity, so separating two genuine neighbours is expensive while stacking two far-apart points is almost free — the optimiser fights for local neighbourhoods and is nearly indifferent to where whole clusters sit relative to each other. The heavy-tailed kernel used in the map actively flings non-neighbours apart, which is why islands look so cleanly separated. Cluster area is just as unreliable: the per-point kernel width is tuned to a fixed perplexity, which normalises density away, so a tight group and a diffuse group can occupy similar area. In a single-cell UMAP or t-SNE view, the wide band between two cell-type islands and their relative sizes are layout artefacts — read them qualitatively and go back to the original features for any number you plan to report.
go deeper
Recall the one-line rule: a t-SNE or UMAP map is trustworthy about who is near whom and untrustworthy about how far apart and how big. Never quote a distance or an island size off the picture.
Be ready to say why: the divergence is weighted by data-space similarity, so tearing neighbours apart is costly and moving strangers is cheap, and the per-point bandwidth normalises density away.
Show how you police it in practice — fixed seeds, a principal-component initialisation, several settings compared, and every claim taken from the map back to the original features before it is reported.
Own the reporting standard. Decide when a projection may appear in a deck at all, what caption must sit under it, and how the team distinguishes an exploratory instrument from evidence.
## What t-SNE is actually optimising t-SNE (t-distributed stochastic neighbour embedding) takes a table of high-dimensional rows and produces one 2-D point per row, in three steps. 1. For each point `i`, distances to every other point are turned into a probability distribution `p(j|i)` with a Gaussian kernel centred on `i`. The kernel width is chosen separately for each point, by binary search, so that the resulting neighbour distribution has the perplexity the user asked for — loosely, so every point ends up with the same effective number of neighbours. These conditionals are symmetrised into a joint distribution `P` over pairs. 2. In the 2-D map a second distribution `Q` is defined over pairs using a Student-t kernel with one degree of freedom: `q_ij` is proportional to `1 / (1 + d_ij^2)`, where `d_ij` is the map distance. That kernel has heavy tails. 3. The map points are moved by gradient descent to minimise the Kullback-Leibler divergence `KL(P || Q) = sum p_ij * log(p_ij / q_ij)`. Every property people misread follows from those three choices. ### The objective is asymmetric Each term in the sum is weighted by `p_ij`, the similarity **in the data**. If two rows are genuine neighbours (`p_ij` large) and the map puts them far apart (`q_ij` small), the term blows up — very expensive. If two rows are far apart in the data (`p_ij` near zero) and the map places them close together, the weight kills the term — nearly free. So the optimiser works hard to keep neighbourhoods intact and is close to indifferent about the arrangement of one cluster relative to another. Local structure is what you are buying; global geometry is largely unconstrained. ### The heavy tail pushes clusters apart The Student-t kernel exists to fix the crowding problem: a 2-D disc has vastly less room than a 50-dimensional ball, so moderate distances cannot all be honoured. Heavy tails let a moderate data-space distance be drawn as a large map distance at low cost. The visible side effect is clean islands with white space between them. The white space means only *these rows were not each other's neighbours*. Its width is not a measurement, and two gaps of different widths do not rank two pairs of clusters. ### Density is equalised Because each point's kernel width is tuned to hit a fixed perplexity, a point in a dense region gets a narrow kernel and a point in a sparse region a wide one — relative density is normalised away. A compact group and a diffuse group of the same membership can end up as islands of similar area, and a small group can look large. So island area reports neither the number of rows nor the spread of the group. ### It is not deterministic The objective is non-convex and the map starts from a random initialisation, with an early phase that exaggerates `P` to let clusters separate. Two runs at identical settings settle in different local minima: islands rotate, swap sides, occasionally split or merge. What tends to survive across runs is *which points sit together*, not the coordinates, the orientation or the gaps. Fixing the random seed makes one run reproducible; initialising the map from the first two principal components rather than at random makes the coarse arrangement much more stable from run to run. ### The concrete failure In a single-cell UMAP, two cell-type islands sit far apart with a wide empty band, and one island looks three times the area of the other. The tempting report is *these two populations are very different, and the first one dominates*. Neither claim is supported. The band says the two groups were not each other's neighbours at the chosen neighbourhood scale — nothing about how different they are. The areas reflect density normalisation, not counts, and the count is a number you can simply compute. Any statement about how far apart the populations are belongs in the original expression space, using a statistic you can defend. ### What you may legitimately read - Which rows land together locally. - Whether a label you already hold (class, batch, cell type) paints coherently onto the map — good evidence of separability, and one of the fastest ways to spot a batch effect. - Isolated pockets worth going back and inspecting in the original features. Everything quantitative goes back to the original space. ### Does UMAP fix it? UMAP builds a weighted nearest-neighbour graph and optimises a cross-entropy layout, and it does usually retain more of the coarse arrangement than t-SNE, especially with a larger neighbourhood setting. But it is still a neighbourhood-driven layout produced by a stochastic optimiser, and between-island distance is still not a metric. For both methods the honest register is qualitative: the map is an instrument for generating hypotheses, not for measuring them.
- Does the same warning apply to a UMAP plot, or does UMAP preserve global structure?UMAP typically retains more of the coarse arrangement than t-SNE, and a larger neighbourhood setting retains more still. But it is the same species of object: a stochastic layout driven by a nearest-neighbour graph. Between-island distance is not a metric, so report it qualitatively and measure anything you intend to publish in the original feature space.
- Two runs with the same settings gave visibly different layouts — is one of them wrong?Neither. The objective is non-convex and the initialisation is random, so runs land in different local minima and islands rotate or swap position. Membership of neighbourhoods is the stable part. Fix the seed for reproducibility, and initialise from the first two principal components if you want the coarse arrangement to be comparable across runs.
- What can you legitimately conclude from a t-SNE map, then?That certain rows are consistently near neighbours; that an existing label paints coherently or does not, which is real evidence about separability; and that specific pockets deserve a closer look. Treat every one of those as a hypothesis and confirm it in the original features before it reaches a slide.
It is like a seating chart built only from who wants to sit next to whom. Every friendship is honoured, but how far apart two tables end up says nothing about the guests.
saying these in an interview costs you the question
- Reads the gap between two islands as a similarity score
- Compares island areas as if they were cluster sizes
- Says an empty band proves the groups are unrelated
- Treats one run's layout as the definitive picture
- Reports numbers computed from the 2-D coordinates as facts about the data