skip to content

Why must IVF and PQ indexes be trained, and what breaks when the data drifts?

level: seniorimportance: should knowfreq 40%

answer

  1. the index is fitted, not merely built
  2. centroids come from a sample
  3. yesterday's sample, today's data
  4. recall decays while latency looks fine
  5. count training vectors per centroid

basics

~20 s

IVF centroids and PQ codebooks are fitted from a sample of the corpus before ingestion, then frozen. When later data drifts away from that sample, cells become unbalanced and codes reconstruct poorly, so recall degrades quietly while queries keep succeeding.

solid answer

~50 s

Both halves of this index family are *learned* structures. The coarse quantizer's centroids and each PQ subspace codebook come from a k-means pass over a sample of the corpus, and after that they are constants: every subsequent insert is assigned to an existing cell and encoded against existing codebooks. Two things go wrong. First, an unrepresentative or too-small sample produces bad centroids from day one — a rough rule is tens of training vectors per centroid, and FAISS warns when you supply fewer than roughly 39 per centroid. Second, distribution drift: new content lands in regions the sample never covered, so those vectors pile into a handful of cells (imbalance, tail latency) and their residuals exceed what the codebooks model well (reconstruction error, recall loss). The failure is silent — no errors, stable latency, slowly worsening results — so you need a standing recall check against exact search on held-out queries, plus a periodic retrain-and-rebuild.

go deeper

for a junior

Know that these indexes have a training step before vectors are added, and that the centroids and codebooks it produces come from a sample of the data.

for a middle

Be ready to explain the train-add-search sequence, why the sample must be random and large enough per centroid, and that parameters are frozen once training completes.

for a senior

Show the production instinct: drift degrades recall silently, so you keep a ground-truth recall check running, watch list imbalance, and treat retrain-and-rebuild as scheduled maintenance rather than an incident response.

for a principal

Own the lifecycle cost — rebuild compute, blue/green swap procedure, and the fact that an embedding-model upgrade is a full re-embed of the corpus, which should be budgeted before the model change is approved.

## Learned indexes, not just built ones A plain exhaustive index has no parameters: you append vectors and you are done. IVF and PQ are different in kind — they contain *fitted model parameters*, and every guarantee they offer is conditional on those parameters matching the data. - **IVF** needs `nlist` centroids, obtained by k-means over a sample. They define the Voronoi cells that route both inserts and queries. - **PQ** needs one codebook of 256 centroids per subspace, also from k-means, fitted on the corresponding slices of the same sample (or on residuals, if residual encoding is used). The workflow is therefore three phases, not two: **train**, then **add**, then **search**. Skipping or short-changing the first is a leading cause of "my ANN index has bad recall and I don't know why". ## How much training data The number that matters is vectors *per centroid*, not vectors in total. Fitting 4096 centroids from 10,000 samples gives roughly two points per cluster — the centroids are noise. Practical guidance: - Tens of vectors per centroid as a floor; hundreds is comfortable. FAISS emits a warning below roughly 39 points per centroid and a stronger one below about 10. - `nlist` itself is usually chosen around the square root of the corpus size, which conveniently keeps the training requirement modest even for huge corpora: a few million samples suffices for a billion-vector index. - The sample must be **random across the whole corpus**, not the first N rows. Sorting by ingestion date, by tenant, or by source is the standard way to poison it — you fit centroids to one region and then serve the rest. At satellite-imagery scale — say 1.2 billion patch embeddings on one machine with 4096 coarse cells and 64-byte PQ codes — training might use a few million randomly sampled patches, take minutes, and then govern the recall of every one of those 1.2 billion vectors for as long as the index lives. ## What drift actually does Drift means the distribution of new vectors differs from the training sample. Causes are mundane: the product expands into a new language or region, a new document type is ingested, seasonal content arrives, or — the big one — **the embedding model is upgraded**, which relocates the entire space. Two distinct degradations follow. 1. **Partition imbalance.** New vectors that fall outside the sampled regions are still assigned to their nearest existing centroid, however far away. They pile into a few cells. Those cells grow fat, so a query landing there scans far more than average (p99 latency), and the boundary between the crowded region and everything else is drawn in the wrong place, so edge-of-cell misses multiply. 2. **Reconstruction error.** PQ codebooks model the variance the sample showed them. Vectors from an unseen region get encoded to distant centroids, so their reconstructed distances are badly wrong, and they are ranked essentially arbitrarily against well-modelled vectors — a subtle fairness problem too, because newer or minority content is systematically retrieved worse. Critically, an embedding-model change is not drift you can absorb by retraining the index — it invalidates every stored vector, so the whole corpus must be re-embedded and the index rebuilt from scratch. ## Why it is hard to notice Nothing fails. Every query returns k results with plausible-looking scores; latency is unchanged or only mildly worse; no exception is thrown. The degradation shows up as slightly worse downstream answers, which teams typically attribute to the model or the prompt. Detection has to be deliberate: - **Standing recall check.** Keep a held-out query set with exact ground truth computed by brute force, and run recall@k on a schedule. This is the only direct signal, and it costs almost nothing at a few thousand queries. - **Refresh the ground truth periodically** against the current corpus — an old ground-truth set measures an old dataset. - **List-length distribution.** Track the ratio of the largest inverted list to the median. A rising ratio is the earliest structural warning. - **Fraction of new vectors landing in the top few cells.** A direct drift proxy. - **Sampled reconstruction error.** For a sample of recent vectors, compare the true distance to the PQ-estimated distance; a widening gap says the codebooks are stale. ## Remediation Retrain on a fresh random sample of the *current* corpus and rebuild — for IVF that means re-assigning every vector to the new cells, and for PQ re-encoding every vector, so it is a full offline pass. Standard practice is a blue/green rebuild: construct the new index alongside the live one, validate recall against ground truth, then swap. Cadence is empirical — driven by the recall check rather than the calendar — but any system with a live ingest pipeline should assume periodic rebuilds are part of its operating cost, and should budget the compute for them up front rather than discovering the need after quality complaints arrive.

  • How would you detect this degradation in production?
    Keep a held-out query set with exact brute-force ground truth and run recall@k on a schedule — that is the only direct signal, and it is cheap at a few thousand queries. Supplement it with structural proxies: the ratio of the largest inverted list to the median, the share of newly ingested vectors landing in a few cells, and the gap between true and quantized distances on a recent sample.
  • Does retraining the index fix a change of embedding model?
    No. A new embedding model produces a different vector space, so every stored vector is meaningless against new queries. The whole corpus must be re-embedded and the index rebuilt from those new vectors; retraining centroids over stale embeddings fixes nothing. Plan model upgrades as full re-index events, typically with a blue/green swap after validating recall.
  • What is the most common way teams poison the training sample?
    Taking the first N vectors instead of a random sample. Corpora are almost always ordered by ingestion time, tenant, or source, so the first slice represents one region of the space; centroids and codebooks fit that region and every later vector is served by a partition that was never designed for it. Sample uniformly across the whole corpus.

saying these in an interview costs you the question

  • Assumes IVF and PQ need no training step at all
  • Trains on the first N rows instead of a random sample
  • Thinks retraining fixes an embedding-model upgrade
  • Expects drift to surface as errors or slow queries
  • Sizes training data in total vectors, not per centroid

context