skip to content

FAISS IVF recall is too low in production — how do you tune nprobe and prove it worked?

level: seniorimportance: must knowfreq 68%

answer

  1. search-time knob, no rebuild needed
  2. how many cells each query scans
  3. cost roughly linear, recall saturating
  4. ground truth from an exact flat index
  5. a plateau means a different bottleneck

basics

~20 s

Raise nprobe — the number of cells each query scans. It is a search-time knob needing no rebuild. Sweep it while measuring recall@k against exact brute-force results, then pick the smallest value that hits your recall target within the latency budget.

solid answer

~50 s

`nprobe` is the only knob that changes the recall/latency balance of an IVF index without touching the data: it decides how many of the `nlist` cells a query scans, so cost is roughly linear in it and recall rises with diminishing returns. To tune it honestly you need ground truth, so build an `IndexFlatL2` over the same vectors, run a held-out query set through it, and treat those neighbours as correct. Then sweep `index.nprobe` over, say, 1, 4, 8, 16, 32, 64, record recall@10 and p95 latency at each point, and pick the smallest value that clears your target. Set it per search with `faiss.SearchParametersIVF` when different callers need different tradeoffs. If recall plateaus well below target at every `nprobe`, the ceiling is elsewhere — an oversized `nlist`, an unrepresentative training sample, or PQ codes too short to preserve distances.

code

python · 20 lines
python
import faiss
import numpy as np

d = 64
xb = np.random.random((100_000, d)).astype('float32')
xq = np.random.random((1_000, d)).astype('float32')

exact = faiss.IndexFlatL2(d)
exact.add(xb)
_, gt = exact.search(xq, 10)          # ground truth

index = faiss.IndexIVFFlat(faiss.IndexFlatL2(d), d, 256)
index.train(xb)
index.add(xb)

for nprobe in (1, 4, 16, 64, 256):
    index.nprobe = nprobe
    _, I = index.search(xq, 10)
    recall10 = (I == gt[:, :1]).any(axis=1).mean()
    print(nprobe, round(float(recall10), 3))

go deeper

for a junior

Know that nprobe controls how many clusters a query looks at, that raising it improves recall and costs latency, and that it can be changed at any time without rebuilding.

for a middle

Explain the cost model — roughly ntotal * nprobe / nlist vectors scanned — and describe how to compute recall@k against an exact flat index rather than trusting the approximate index's own scores.

for a senior

Demonstrate the whole loop: held-out queries, exact ground truth, a recall-versus-latency sweep, and the judgment to recognise a plateau as a quantization or nlist problem rather than pushing nprobe higher. Re-measure on every rebuild.

for a principal

Own the target itself: what recall the product actually needs, how much latency and hardware that costs, and whether a refine stage, a bigger code or a different index type is the cheaper way to buy the last few points. Make recall a build-gate metric, not a one-off tuning session.

## What nprobe actually does An IVF index divides the corpus into `nlist` cells. At query time the coarse quantizer ranks the cells by distance from the query to their centroids and scans the closest `nprobe` of them, ignoring the rest. With `nlist=1024` and `nprobe=8` you touch roughly 8/1024 of the corpus, which is where the speedup comes from — and also where the recall loss comes from, because a true neighbour sitting in the ninth-closest cell is simply never seen. That makes the cost model easy: work per query is approximately `ntotal * nprobe / nlist` distance computations plus the cost of ranking centroids, so latency grows close to linearly with `nprobe`. Recall grows fast at first and then flattens, because the first few cells contain most of the true neighbours. At `nprobe == nlist` on an `IndexIVFFlat` every cell is scanned and results match exact search exactly — at brute-force cost, which is the point of not doing it. ## Why it is the knob you reach for first `nprobe` is a search-time parameter. It lives on the index object, not in the stored data, so changing it requires no retraining, no re-adding and no redeploy of the index file. `nlist`, the PQ code size and the choice of index type are all build-time and cost you a rebuild. When recall is missed in production, always establish whether `nprobe` can close the gap before you consider rebuilding. ``` index.nprobe = 32 # global for this index object index.search(xq, 10, params=faiss.SearchParametersIVF(nprobe=64)) # per-query ``` The per-search form matters operationally: a latency-sensitive autocomplete path and an offline batch job can share one index and still run at different recall points. For indexes wrapped in a pre-transform, assign through `faiss.extract_index_ivf(index).nprobe` or `faiss.ParameterSpace().set_index_parameter(index, "nprobe", 32)`, since the outer object has no `nprobe` attribute of its own. ## Measuring recall honestly You cannot tune what you do not measure, and the index cannot grade itself — the distances it returns look perfectly reasonable even when it missed the true nearest neighbour entirely. The ground truth must come from exact search: 1. Hold out a query set that looks like production traffic (real queries if you have them; a random sample of corpus vectors otherwise). 2. Build an `IndexFlatL2` (or `IndexFlatIP` if you use inner product) over the same vectors and search it for k neighbours. Those are the correct answers. 3. For each candidate `nprobe`, search the IVF index and compute recall@k — the fraction of the exact top-k that appears in the approximate top-k. Track the 1-recall@k variant (was the true nearest neighbour returned at all?) when your application only shows one result. 4. Record latency at the same time, at the batch size and concurrency you actually serve. One subtlety: on a corpus of a few million vectors, exact search over a query set of a thousand takes seconds, so this is cheap to run in CI. Do it on every index rebuild and treat a recall regression as a build failure. ## Reading the curve The recall-versus-nprobe curve tells you what to do next: - **Rises steeply, reaches target at a modest nprobe.** Ship that value with a little headroom. This is the healthy case. - **Rises but only reaches target at nprobe close to nlist.** Your cells are too fine or too unbalanced for the corpus; a smaller `nlist`, or better training data, will let you get the same recall at a fraction of the scan. - **Flattens well below target.** `nprobe` is not your bottleneck. On an IVFPQ index, the compressed codes themselves cap achievable recall no matter how many cells you scan — the distances are approximations. Fix that with larger PQ codes, or by re-ranking the shortlist against full-precision vectors (`faiss.IndexRefineFlat` with a `k_factor`), not by raising `nprobe`. - **Recall is poor even at nprobe == nlist on an IVFFlat.** That is impossible from the index's perspective, so your measurement is wrong: mismatched metric between the ground truth and the index, unnormalised vectors under inner product, or an id mapping mixed up. ## The HNSW analogue If the index is an `IndexHNSWFlat` rather than IVF, the equivalent search-time knob is `index.hnsw.efSearch` — how wide the graph traversal keeps its candidate list. It behaves the same way: raise it for recall, pay latency, no rebuild required. `efConstruction` is its build-time counterpart and cannot be changed after the fact. Same tuning loop, different attribute. ## Production practice Pin the chosen `nprobe` in configuration, not in scattered call sites, and record alongside it the recall it was measured to deliver and the corpus version it was measured on. Re-measure after every rebuild, because a new training run moves the centroids and a value that gave 0.95 recall last month may give 0.90 this month. Alerting on a serving-side proxy — for example, how often the top result changes when a shadow query runs at a much higher `nprobe` — catches drift between rebuilds.

  • Recall plateaus at 0.82 no matter how high you push nprobe on an IVFPQ index. What now?
    The ceiling is quantization, not cell selection: PQ stores compressed codes, so the distances are approximate even when the right cell is scanned. Options are a larger code (more sub-vectors or more bits), or keeping the compressed index for candidate generation and re-ranking the shortlist against full-precision vectors with IndexRefineFlat and a k_factor above 1. Raising nprobe past that plateau only buys latency.
  • How would you choose nprobe when different callers have different latency budgets?
    Do not fork the index. Measure the whole recall/latency curve once, then pass a per-search faiss.SearchParametersIVF(nprobe=...) so an interactive path can run lean while a batch or high-stakes path runs wide. Pin both values in configuration next to the recall each was measured to deliver, and re-measure after every index rebuild.
  • What is the equivalent knob on an IndexHNSWFlat, and what is its build-time counterpart?
    index.hnsw.efSearch controls how wide the candidate list stays during traversal — raise it for recall, at roughly proportional latency cost, with no rebuild. Its build-time counterpart is efConstruction, which affects graph quality and is fixed once vectors are added. The tuning method is identical: sweep efSearch against exact ground truth and take the smallest value that meets target.
  • Your recall measurement reports near zero at every nprobe. What do you check before touching the index?
    The measurement. Confirm the ground truth index uses the same metric as the IVF index (L2 versus inner product), that vectors are normalised if you rely on cosine via inner product, that both searched the identical vector set, and that ids line up if IndexIDMap is involved. A genuine index problem degrades recall gradually; near-zero recall is almost always a harness bug.

saying these in an interview costs you the question

  • Judging recall from the distances the approximate index itself returns
  • Assuming higher nprobe always fixes recall, including on PQ indexes
  • Believing nprobe changes require retraining or rebuilding
  • Tuning nprobe on synthetic queries unlike production traffic
  • Setting nprobe once and never re-measuring after a rebuild

context