skip to content

Embedding Space Geometry

A similarity number only means something given the shape of the space it came from. Embedding spaces are anisotropic, cosine scores are not calibrated across models, and some vectors match almost everything.

on this pageshow

questions

5

Why do unrelated texts still score 0.8 cosine in an embedding space?

level: middleimportance: must knowfreq 55%

answer

  1. scores are relative, not absolute
  2. the whole space points one way
  3. narrow cone, not a full sphere
  4. measure random-pair baseline first
  5. anisotropy inflates every cosine

basics

~20 s

Learned embedding spaces are anisotropic: vectors crowd into a narrow cone instead of spreading over the sphere, so almost every pair shares a large common direction. That inflates all cosine scores, making 0.8 a baseline rather than a match.

solid answer

~50 s

Cosine similarity measures the angle between two vectors, and in a trained text-embedding space the angles are not spread evenly. The vectors occupy a narrow cone around one or a few dominant directions — a property called **anisotropy**. Because every embedding carries a large shared component, even completely unrelated texts land at a high cosine; a noise floor of 0.7–0.85 is typical for many encoders. The absolute number therefore carries almost no information, and only the *separation* between related and unrelated pairs does. The practical move is to measure your own baseline: sample a few thousand random pairs from your corpus, compute the mean and 95th percentile of their cosine, and treat that band as zero. Any cutoff that is not comfortably above it will match everything. Mean-centering or whitening the space widens the spread if you need better-behaved raw scores.

code

python · 12 lines
python
import numpy as np

rng = np.random.default_rng(0)
# a synthetic anisotropic space: every vector carries one strong shared direction
common = rng.normal(size=768)
X = 0.8 * common + 0.2 * rng.normal(size=(2000, 768))
X /= np.linalg.norm(X, axis=1, keepdims=True)

i, j = rng.integers(0, len(X), size=(2, 20000))
keep = i != j
baseline = (X[i[keep]] * X[j[keep]]).sum(axis=1)
print(f"random-pair cosine: mean={baseline.mean():.3f} p95={np.percentile(baseline, 95):.3f}")

go deeper

for a junior

Know that a cosine similarity is a relative score, not a percentage, and that unrelated texts can still score high. Say plainly that you would compare scores within one model rather than trusting a fixed number.

for a middle

Be ready to name anisotropy, explain that vectors occupy a narrow cone with a shared dominant direction, and describe how to measure the random-pair baseline before choosing any cutoff.

for a senior

An interviewer expects you to have diagnosed this in production: showing that a threshold sat below the noise floor, widening separation with centering or whitening, and moving decision logic from absolute scores to ranks and margins.

for a principal

Own the consequence for system design — that score scales are properties of a model-and-corpus pair, so thresholds are versioned artifacts, cross-model score comparisons must be forbidden by convention, and downstream teams should consume calibrated outcomes rather than raw similarity.

## What anisotropy means A vector space is *isotropic* when its points spread evenly in all directions: pick two at random and their angle averages ninety degrees, so their cosine is about zero. That is the mental model most people bring to cosine similarity — 0 means unrelated, 1 means identical, and the values in between form an interpretable scale. Learned text-embedding spaces do not behave that way. They are **anisotropic**: the vectors crowd into a narrow cone, all pointing broadly the same way. A handful of directions dominate the variance, and every embedding carries a large component along them. Two texts that share nothing semantically still share that common direction, so their cosine is high before any real similarity is measured. ## Why trained encoders end up like this The cone is a by-product of how the encoders are trained, not a defect in your pipeline. Language-model training pushes representations of frequent and rare tokens into systematically different regions, and the shared statistical structure of natural language — the same function words, the same discourse patterns, the same document conventions — leaves a common signature in every vector. Models trained with a contrastive retrieval objective are noticeably flatter than raw language-model hidden states, because the objective explicitly separates positives from negatives, but flatter is not isotropic. Effectively every production embedding space you will meet has an inflated baseline. ## What it does to your numbers Three consequences matter in interviews and in production. First, **absolute thresholds are meaningless in isolation**. Consider a duplicate-bug-report detector: two reports about completely unrelated subsystems come back at 0.82 cosine. The intuitive "0.9 means duplicate" rule inherited from a tutorial is not just badly tuned, it is measuring against the wrong zero. Whether 0.82 is high depends entirely on where this model's random-pair distribution sits. Second, **all the signal lives in a thin band**. If unrelated pairs average 0.78 and true duplicates average 0.91, your entire dynamic range is thirteen hundredths. Small score differences that look like rounding noise are actually the whole decision, which is why score-based logic in these systems is so fragile and why rank-based logic is usually sturdier. Third, **scores do not transfer across models**. Every model has its own cone, so its own baseline. A cutoff tuned for one encoder, copied to another, silently changes the precision/recall operating point. ## Diagnosing it: the random-pair baseline The measurement is cheap and should be a standard step whenever you adopt an embedding model. Sample a few thousand random pairs from your own corpus — random pairs are overwhelmingly unrelated — and compute the distribution of their similarity. The mean tells you where zero really is; the 95th percentile tells you the score a non-match can reach by luck. Then sample known-related pairs and compare the two distributions. If they overlap heavily, no threshold will separate them and the problem is the representation, not the cutoff. ## What helps **Centering.** Subtract the corpus mean vector from every embedding before comparing. This removes the shared direction and immediately widens the score spread. **Whitening** goes further, rescaling the principal directions so variance is more evenly distributed; both are post-processing steps that leave the model untouched. Both must be applied identically at index time and query time, and the mean must be recomputed if the corpus shifts substantially. **Think in ranks and margins.** Instead of "score above T", use "top-1 and at least M above top-2", or percentile-of-baseline rather than a raw cosine. These survive a model swap far better than a bare number. **Normalize your expectations, not just your vectors.** Report calibrated quantities to stakeholders — precision at a chosen operating point — never raw cosine, which invites people to read it as a percentage. ## What anisotropy is *not* It is not fixed by L2-normalizing the vectors: normalization removes magnitude, and the cone is about direction, so the baseline stays exactly where it was. It is not fixed by adding dimensions. It is not evidence that the model is broken — a model with a 0.8 baseline can still rank beautifully. And it does not mean cosine is the wrong metric; it means the metric's output is a relative quantity in a space you have to characterize before you can read it.

  • Does L2-normalizing the embeddings reduce the inflated baseline?
    No. Normalization sets every vector's length to one, which is about magnitude; anisotropy is about direction. After normalizing, all the vectors still sit in the same narrow cone, so the mean cosine between random pairs is unchanged. What does help is centering — subtracting the corpus mean vector before comparison — because that removes the shared direction itself. Normalization and centering solve different problems and are often applied together.
  • How would you report similarity to a product stakeholder who wants a percentage?
    Never hand over raw cosine, because it reads as a percentage and is not one. Convert to a calibrated quantity instead: on a labelled sample, measure precision and recall at the operating point you actually run, and report those. If you need a per-pair number, fit a mapping from raw score to estimated match probability on held-out labelled pairs and report the probability, noting it is valid only for this model and corpus.
  • Two models both give 0.9 on your test pair. Which is better for retrieval?
    You cannot tell from that number. Compare each model's separation instead: measure the random-pair baseline and the related-pair distribution for both, then compare the gap, or simply compare ranking quality on labelled queries. A model with a 0.6 baseline and 0.9 on matches is discriminating far more than one with a 0.85 baseline and the same 0.9 on matches.

Everyone in the photo is facing roughly the same direction, so the angle between any two people is small regardless of whether they came together. You learn who arrived together only by comparing small differences in where they face.

saying these in an interview costs you the question

  • Treating cosine 0.8 as roughly 80 percent similar
  • Copying a similarity threshold between different embedding models
  • Believing normalization removes the inflated similarity baseline
  • Concluding the model is broken because unrelated pairs score high
  • Never measuring the random-pair baseline for the corpus

context

open as a page

Why do embedding models like E5 and BGE need different prefixes for queries and passages?

level: middleimportance: should knowfreq 48%

basics

~20 s

They were trained asymmetrically: short questions and long documents are mapped into the shared space by different roles, and the prefix tells the model which side it is encoding. Encode both sides identically and retrieval quality degrades silently.

open as a page

What breaks when only half a vector index's embeddings are L2-normalized?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Dot-product ranking scales with vector length, so unnormalized entries score by magnitude as well as direction. A half-normalized collection splits into two score scales: one group crowds every result list while the other becomes effectively unreachable.

open as a page

How do you choose an embedding similarity threshold for auto-closing incidents?

level: principalimportance: should knowfreq 42%

basics

~20 s

Derive it from labelled pairs on that exact model and corpus, choose the operating point from the cost of a wrong auto-close, and prefer a margin over the runner-up to a bare number. Re-validate on every model or corpus change.

open as a page

In vector search, why does one chunk appear in the top-10 for most queries?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

That chunk is a hub. In high-dimensional spaces a few vectors sit near the centre of the distribution and land in many nearest-neighbour lists regardless of the query. Generic boilerplate is the textbook hub.

open as a page