skip to content

Training and Recall Tuning

IVF and PQ indexes must be trained on representative data before you add vectors, and recall is then bought back with search-time parameters. This is where most FAISS mistakes actually happen.

on this pageshow

questions

5

Why must a FAISS IVF index be trained before add(), and on what data?

level: middleimportance: must knowfreq 76%

answer

  1. one-time fitting step, not a container
  2. k-means centroids and PQ codebooks
  3. is_trained flips false to true
  4. roughly 39 points per centroid
  5. sample from the real corpus

basics

~20 s

IVF and PQ indexes learn structure from data — k-means centroids for IVF, codebooks for PQ. Until train() runs, is_trained is False and add() raises. Train on a representative sample of the vectors you will actually search.

solid answer

~50 s

An `IndexIVFFlat` or `IndexIVFPQ` is not a plain container: it partitions the space into `nlist` cells, and those cell centroids come from running k-means over a training set. PQ additionally learns one codebook per sub-vector. Both happen inside `train()`, which flips `is_trained` from False to True; calling `add()` first raises a RuntimeError. The training data must come from the same distribution as your corpus — a random subsample of the real embeddings, not synthetic or out-of-domain vectors, because unrepresentative centroids give unbalanced cells and queries then land in cells that do not hold their true neighbours. FAISS wants roughly 39 training points per centroid and will warn below that; it also subsamples above 256 per centroid, so feeding it tens of millions of vectors buys nothing. `IndexFlatL2`, `IndexFlatIP` and `IndexHNSWFlat` need no training at all.

code

python · 14 lines
python
import faiss
import numpy as np

d, nlist = 128, 1024
xt = np.random.random((100_000, d)).astype('float32')   # representative sample
xb = np.random.random((1_000_000, d)).astype('float32')  # corpus

quantizer = faiss.IndexFlatL2(d)
index = faiss.IndexIVFFlat(quantizer, d, nlist)
print(index.is_trained)   # False
index.train(xt)           # k-means -> nlist centroids
print(index.is_trained)   # True
index.add(xb)
print(index.ntotal)       # 1000000

go deeper

for a junior

Remember the order: construct, train, add, search. Be able to say that IVF and PQ indexes must be trained and that flat indexes need no training at all.

for a middle

Explain what train() actually computes — k-means cell centroids for IVF and per-sub-vector codebooks for PQ — and why add() raises before it. Know the rough sizing rule of tens to a couple of hundred training points per centroid.

for a senior

Show that you can diagnose an unrepresentative training sample: unbalanced cells, recall that stays flat no matter how you raise nprobe, and skew introduced by sampling one slice of the corpus. Treat the trained index as an immutable build artifact.

for a principal

Own the lifecycle: when a corpus or model change forces a rebuild, how the new index is validated against the old one before a swap, and how index build cost and rebuild cadence factor into the cost of changing embedding models at all.

## What "training" means in FAISS FAISS indexes fall into two groups. Some are pure containers: `IndexFlatL2` and `IndexFlatIP` store raw float32 vectors and compare a query against every one of them, and `IndexHNSWFlat` builds its graph incrementally as vectors arrive. These report `is_trained == True` from the moment they are constructed, and you can call `add()` immediately. The other group is *data-dependent*. An inverted-file index (`IndexIVFFlat`, `IndexIVFPQ`) does not search all vectors; it first partitions the vector space into `nlist` cells and stores each vector in the cell whose centroid is nearest. Those centroids are learned by running k-means over a training set. Product quantization goes further: it splits each vector into `m` sub-vectors and learns a codebook of 2^nbits centroids for each slice, so that a vector can later be stored as `m` small integer codes instead of floats. Neither the cell centroids nor the codebooks can be invented — they are fitted parameters, and `train()` is the fitting step. ## What happens if you skip it `index.is_trained` is False until `train()` returns. Calling `add()` on an untrained IVF index raises a RuntimeError from FAISS's internal assertion — it does not silently train itself, and it does not silently degrade. This is the single most common first-hour FAISS error, and the fix is always the same: train, then add. ``` index = faiss.IndexIVFFlat(faiss.IndexFlatL2(d), d, nlist) index.is_trained # False index.train(xt) index.add(xb) ``` ## How much training data FAISS's k-means implementation has two internal bounds that drive the practical answer. It expects at least about 39 points per centroid and prints a warning like *"WARNING clustering N points to K centroids: please provide at least ... training points"* when you give it fewer. Above about 256 points per centroid it randomly subsamples, so a training set larger than `256 * nlist` gives you no better centroids and only costs time. Product quantization has a parallel requirement: each sub-quantizer learns 2^nbits centroids, so with the default nbits=8 you want at least a few tens of thousands of training vectors before the codebooks mean anything. That gives a simple sizing rule. Pick `nlist` first (commonly on the order of sqrt(N) for N vectors), then take a random sample of roughly 40–256 times `nlist` vectors to train on. For a million vectors with nlist=1024, something like 100k–260k training vectors is plenty; training on the full million is wasted work. ## Representativeness beats volume The centroids define where every future vector and every future query will be routed. If the training sample is drawn from a different distribution than the corpus — one language when production is multilingual, one product category when the catalogue is broad, or synthetic random vectors during a prototype — the cells are unbalanced. Some cells hold a huge share of the corpus and others are nearly empty, so scanning `nprobe` cells scans far more or far less than the intended fraction, and queries frequently miss neighbours that sit just over a badly-placed boundary. The symptom is recall that is poor at every `nprobe` value, which is the tell that the problem is the training set rather than the search parameter. A random subsample of the real corpus is the right default. If your data has known strata, sampling proportionally across them protects against a skewed draw. ## Training is a build-time artifact Because training is expensive and deterministic given its input, the normal production pattern is: train once, add the corpus, `faiss.write_index(...)`, and ship the file. Serving processes only read it. You cannot "retrain in place" — changing the centroids invalidates every code and every cell assignment already stored, so a retrain means building a new index from scratch and swapping it in. Plan for that when the embedding model changes or when the corpus drifts substantially, and treat the trained index as an immutable build artifact tied to a specific model version. ## Failure modes to recognise - `add()` raising because `is_trained` is False — train first. - A clustering warning at build time — your training sample is too small for `nlist`; either sample more or lower `nlist`. - Uniformly poor recall across all `nprobe` values — suspect unrepresentative training data or an `nlist` far too large for the corpus size. - Someone "fixing" drift by training on new data and adding to the existing index — that is a rebuild, not an incremental update.

  • Your corpus doubles over six months. Do you retrain?
    Not automatically. Adding vectors to an already-trained index is fine as long as they come from the same distribution — the centroids stay usable. Retrain when the distribution actually shifts (new languages, new domains, a new embedding model) or when cells become badly unbalanced, which shows up as recall that no longer responds to nprobe. Retraining means building a fresh index and swapping it, since existing cell assignments and PQ codes are invalidated.
  • Does giving train() ten million vectors instead of two hundred thousand improve the centroids?
    Almost never. FAISS's clustering subsamples the training set to roughly 256 points per centroid, so beyond about 256 * nlist vectors you are paying for a random subsample you could have taken yourself. The extra data only helps if your smaller sample was skewed; fix that by sampling better, not by sampling more.
  • Which common FAISS index types report is_trained as True on construction?
    IndexFlatL2 and IndexFlatIP, which store raw vectors and search exhaustively, and IndexHNSWFlat, which builds its graph as vectors are added. Anything with a coarse quantizer (IVF) or a learned codebook (PQ, scalar quantization, OPQ) must be trained. A quick guard in build code is to assert index.is_trained before calling add().

Training an IVF index is like laying out the aisles of a warehouse before any stock arrives: the aisle boundaries are chosen from a sample of what you expect to store, and if that sample was wrong, everything ends up crammed into two aisles.

saying these in an interview costs you the question

  • Thinking add() trains the index implicitly on the first batch
  • Training on random or synthetic vectors instead of real embeddings
  • Believing more training data always yields better centroids
  • Assuming a trained index can be retrained in place
  • Calling train() again after every batch of new vectors

context

open as a page

FAISS IVF recall is too low in production — how do you tune nprobe and prove it worked?

level: seniorimportance: must knowfreq 68%

basics

~20 s

Raise nprobe — the number of cells each query scans. It is a search-time knob needing no rebuild. Sweep it while measuring recall@k against exact brute-force results, then pick the smallest value that hits your recall target within the latency budget.

open as a page

How do you persist a FAISS index with write_index and read_index, and what is not saved?

level: juniorimportance: should knowfreq 55%

basics

~20 s

faiss.write_index(index, path) serialises the whole index — trained centroids, codebooks, codes and ids — and faiss.read_index(path) restores it ready to search, with no retraining. What it never stores is your documents or metadata; FAISS keeps only vectors and 64-bit ids.

open as a page

In FAISS IndexIVFPQ, what do the m and nbits parameters control and cost?

level: middleimportance: should knowfreq 52%

basics

~20 s

m is how many sub-vectors each vector is split into and nbits the bits per sub-code, so a stored vector shrinks to about m*nbits/8 bytes. Raising either improves accuracy but costs memory, distance-computation time and training data.

open as a page

What does OPQ preprocessing buy a FAISS PQ index, and what does it cost?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

OPQ learns a rotation applied before product quantization so that variance is spread evenly across sub-vectors instead of concentrating in a few. That typically lifts recall at the same code size, paying longer training and an extra matrix multiply on every vector and query.

open as a page