skip to content

What breaks when only half a vector index's embeddings are L2-normalized?

level: seniorimportance: should knowfreq 38%

answer

  1. magnitude leaks into the ranking
  2. one invariant, half enforced
  3. check the norm histogram
  4. two populations, two score scales
  5. queries cannot compensate for stored lengths

basics

~20 s

Dot-product ranking scales with vector length, so unnormalized entries score by magnitude as well as direction. A half-normalized collection splits into two score scales: one group crowds every result list while the other becomes effectively unreachable.

solid answer

~50 s

An inner-product index ranks by length times length times the cosine of the angle. When every stored vector has unit length, that reduces to pure angle and the index behaves exactly like cosine. Break the invariant for part of the corpus and the length term comes back — but only for those rows. If the unnormalized vectors are longer than unit length, they win comparisons on magnitude alone and appear near the top for unrelated queries; if shorter, they are pushed off every result list and become invisible. The bias is not random, because embedding norms vary systematically with things like text length and lexical frequency, so a whole class of documents is affected together. Detect it by plotting the distribution of stored norms and looking for two modes. Fix it by normalizing in one shared write path and re-embedding or re-upserting the affected rows — or by configuring the index for cosine so it normalizes for you.

code

python · 11 lines
python
import numpy as np

rng = np.random.default_rng(1)
normalized = rng.normal(size=(500, 384))
normalized /= np.linalg.norm(normalized, axis=1, keepdims=True)
raw = rng.normal(size=(500, 384)) * rng.uniform(3, 12, size=(500, 1))
stored = np.vstack([normalized, raw])

norms = np.linalg.norm(stored, axis=1)
print(f"min={norms.min():.2f} median={np.median(norms):.2f} max={norms.max():.2f}")
print(f"within 1% of unit length: {(abs(norms - 1) < 0.01).mean():.0%}")

go deeper

for a junior

Know that inner-product ranking is affected by vector length, so many pipelines normalize embeddings to unit length before storing them, and that this must be done consistently.

for a middle

Explain that with unit-length vectors the length terms cancel and inner product ranks by angle alone, and describe what happens to a document whose stored vector is much longer or shorter than the rest.

for a senior

Demonstrate the diagnosis: a norm histogram over the stored vectors, retrieval-frequency counting over a query sample, and a ranking diff against freshly normalized re-embeddings to pin the offending ingestion batch.

for a principal

Own the invariant as an architectural decision — one enforced write path with a boundary assertion, or an index metric that removes the invariant entirely — and require ingestion metadata that makes any future violation attributable.

## The invariant an inner-product index depends on Many vector stores let you choose how vectors are compared, and the inner-product option is popular because it is the cheapest to compute. Inner product accounts for both the angle between two vectors and their lengths. The usual reason this is acceptable is an invariant maintained *outside* the index: every vector written to it has unit length. Under that invariant the length terms vanish and inner product ranks identically to angle-based similarity. The invariant is enforced by your code, not by the store. Nothing rejects a vector of length 7.3. That is the setup for the bug. ## What a mixed collection actually does Suppose an initial backfill normalized at write time, and a later ingestion path — a new service, a migrated job, a different model wrapper — did not. Now the collection holds two populations. For a query vector of fixed length, the score of each candidate is proportional to that candidate's length multiplied by its cosine with the query. The unnormalized rows are typically much longer than unit length, so their scores are scaled up by a constant-ish factor that has nothing to do with relevance. A weakly relevant long-norm document beats a strongly relevant unit-norm one. In practice you see one cohort dominating the top-k across unrelated queries, exactly like a hub, except the cause here is arithmetic rather than the shape of the space. The mirror case is worse to diagnose. If the unnormalized vectors happen to be *shorter* than unit length — some encoders and some pooling choices produce small-norm outputs — those documents are suppressed. They are in the index, they are perfectly relevant, and they never surface. Recall drops for a subset of the corpus, and no error, latency change, or score anomaly reveals it. Crucially the affected set is not random. Embedding norms correlate with properties of the text such as length and how frequent its vocabulary is, so an entire category of content moves together: all the long policy documents, all the short ones, everything ingested after a given date. ## Detecting it The fastest check is to read the stored vectors back and plot the distribution of their norms. A healthy normalized collection shows a spike at 1.0 with a floating-point-width spread. A mixed one shows two modes, or a long tail. Report the share of vectors within, say, one percent of unit length; anything less than 100% is the bug. A second check works from the outside: sample a large set of realistic queries, run retrieval, and count how often each document identifier appears in the top-k. A heavy-tailed retrieval-frequency distribution — a handful of documents in a large fraction of result lists — means something is favouring those rows independently of the query, and norm inflation is the first hypothesis to test. A third check is a ranking diff: re-embed a sample of documents, normalize them, and compare the top-k those queries return against what the live index returns. Large disagreement concentrated on a specific ingestion batch identifies the culprit. ## Fixing and preventing it The repair is to make the whole collection consistent. Either re-upsert the affected rows with normalized vectors, or rebuild the collection if you cannot cleanly identify them. There is no query-side patch: you cannot compensate for stored magnitudes by scaling the query, because a query multiplier applies uniformly to all candidates and leaves the relative bias untouched. Prevention is structural. Put normalization in exactly one write path and make it impossible to bypass — a single ingestion function that every producer must call. Assert the invariant at the boundary: check the norm of each vector before upsert and fail loudly rather than silently accepting it. If the store offers an explicit cosine metric, choosing it removes the invariant entirely by normalizing internally, which is often the right call for a team that has more than one ingestion path. Be alert to the failure mode this bug hides behind: some providers return unit-length vectors by default, so a codebase that never normalized anything can work perfectly for years and then break the day someone adds a second model that does not. The code did not change; the assumption did. ## What the interviewer is testing They want to hear that you know inner product is only equivalent to angle-based comparison *under an invariant your code maintains*, that a partially maintained invariant is more dangerous than an abandoned one because it corrupts only part of the ranking, and that you would diagnose it with a norm histogram and retrieval-frequency counting rather than by staring at similarity scores — which look entirely normal throughout.

  • Could you scale the query vector to compensate for the unnormalized documents?
    No. Multiplying the query by any constant scales every candidate's score by the same factor, so the ordering is unchanged — and ordering is what retrieval returns. The bias comes from per-document lengths, which differ row by row, so only a per-document correction fixes it. That correction is exactly normalization, which means re-writing the vectors.
  • How would you find which ingestion batch introduced the unnormalized rows?
    Join the norm outlier set against ingestion metadata — source, model version, timestamp — which is one reason to store those fields alongside every vector. Typically the outliers cluster tightly on one producer or one date range. If metadata is missing, re-embed a sample with each candidate pipeline and match the resulting norms against the stored ones.
  • If the store offers an explicit cosine metric, is normalizing at write time still worth it?
    Usually yes, for consistency rather than correctness. A cosine metric normalizes internally, so the bug cannot occur through that index — but the same vectors are often reused elsewhere, in clustering, deduplication, or an offline job that does raw dot products. Normalizing once at write time gives every consumer the same guarantee instead of relying on each one to normalize correctly.

saying these in an interview costs you the question

  • Assuming the vector store rejects non-unit-length vectors
  • Thinking dot product and cosine always rank identically
  • Trying to fix stored-magnitude bias by rescaling the query
  • Diagnosing from similarity scores, which look completely normal
  • Normalizing in several call sites instead of one write path

context