skip to content

How do nprobes and refine_factor change recall on a LanceDB vector search?

level: seniorimportance: should knowfreq 52%

answer

  1. Two knobs, two distinct failure modes
  2. One widens the search, one rescores it
  3. Missed entirely versus merely misordered
  4. Over-fetch, then read the real vectors
  5. Refinement costs random reads, not compute

basics

~20 s

nprobes sets how many index partitions a query searches, so raising it widens coverage at linear latency cost. refine_factor over-fetches candidates and rescores them against the full-precision vectors, fixing ordering errors caused by compression at the price of extra reads.

solid answer

~50 s

Both are per-query knobs on the search builder, so you can tune them without rebuilding the index. `.nprobes(n)` controls how many IVF partitions the query visits; the default is small (20), and raising it is the first lever when recall is poor, because a missed neighbour usually lives in a partition you never opened. Cost is close to linear in n. `.refine_factor(k)` attacks a different failure: with an IVF_PQ index the candidate distances come from compressed codes, so even a candidate you did find can be ranked wrongly. Setting `refine_factor(k)` retrieves `limit × k` candidates, reads their full-precision vectors, recomputes exact distances and returns the true top `limit`. That costs extra random reads — expensive on object storage, cheap on local NVMe. The diagnostic is: if the right document is not in the candidate set at all, raise nprobes; if it is present but ranked badly, add refinement.

code

python · 8 lines
python
rows = (
    tbl.search(query_vec)
    .metric("cosine")
    .nprobes(60)         # visit 60 IVF partitions instead of the default 20
    .refine_factor(10)   # fetch 10x limit, rescore on full-precision vectors
    .limit(10)
    .to_list()
)

go deeper

for a junior

Know that both are chained onto the search query and that neither needs an index rebuild. nprobes searches more of the index; refine_factor double-checks the top candidates more carefully.

for a middle

Explain the mechanics: nprobes is how many IVF partitions are visited, refine_factor over-fetches limit × k candidates and rescores them against full-precision vectors, and latency grows roughly linearly with nprobes.

for a senior

Diagnose from symptoms rather than guessing — missing versus misranked — and show a measured sweep against a flat-scan ground truth with p99 latency recorded alongside recall. Set different values per query path.

for a principal

Own recall as a product-level target with an explicit latency and cost budget, especially on object storage where refinement's random reads are billed requests. Decide who is allowed to move these knobs and how the target is monitored over time.

## Two different failure modes An IVF_PQ search can go wrong in two independent ways, and each knob addresses exactly one of them. Conflating them is the most common tuning mistake, and it wastes latency budget on the wrong axis. **Miss**: the true neighbour sits in a partition the query never probed, so it is not in the candidate set at all. No amount of rescoring recovers it — the candidate list simply does not contain it. **Misrank**: the true neighbour is in the candidate set, but its distance was computed from a compressed code rather than the real vector, so it lands below worse candidates and falls outside the returned top-k. `nprobes` fixes misses. `refine_factor` fixes misranks. ## nprobes The IVF stage assigns the query to the nearest partition centroids and searches only those partitions. `.nprobes(n)` on the query builder sets how many. LanceDB defaults to a small value — 20 — chosen for latency, not for recall. Raising n adds partitions in order of centroid distance, so early increases buy a lot of recall and later ones buy progressively less: the recall curve rises steeply, then flattens. Latency, by contrast, grows roughly linearly, because each additional partition is more vectors to score. That asymmetry is what makes tuning tractable — there is usually a clear knee where recall has plateaued and further probes are pure cost. How many you need depends on how the data clusters. Well-separated embeddings need few probes; embeddings that are nearly uniform, or a query distribution that sits on cluster boundaries, need many. This is why the correct value is measured on your own data, not inherited from a blog post. Note the interaction with build parameters: probing 20 of 256 partitions covers a very different fraction of the corpus than probing 20 of 4096. Changing the partition count and leaving the probe count alone silently changes recall. ## refine_factor Product quantization stores an approximation of each vector. Distances computed from those codes are close but not exact, and near-ties get ordered wrongly. `.refine_factor(k)` addresses that directly: the engine fetches `limit × k` candidates using the fast approximate distances, then reads the actual stored vectors for those candidates, recomputes exact distances, re-sorts, and returns the top `limit`. The effect is that PQ is demoted to a candidate generator and the final ordering is exact with respect to the candidates it produced. In practice a modest factor — 5 to 10 — recovers most of the ranking loss. The cost is I/O, not compute. Refinement reads full-precision vectors for `limit × k` rows, and those reads are scattered. On local NVMe this is nearly free. On object storage each read is a request with real latency and a real price, so a high refine factor on a remote table can dominate query time — the exact deployment LanceDB is often chosen for. Budget it accordingly. Refinement cannot repair a miss. If the neighbour was never in the `limit × k` candidates, rescoring will not conjure it. ## A tuning procedure 1. Build ground truth: sample a few hundred real query vectors and record the exact top-k, from a flat scan or a small unindexed copy of the data. 2. Measure baseline recall@k with defaults, and record p50/p99 latency alongside it. Recall without latency is meaningless. 3. Sweep `nprobes` alone — 20, 40, 80, 160 — and find the knee where recall stops improving. 4. If recall is still short of target at that knee, the residual loss is ranking loss, so add `refine_factor` at 5, then 10, and watch p99 rather than the mean; refinement's random reads show up in the tail first. 5. Fix the values per query path. A user-facing autocomplete and an offline batch enrichment over the same table should not share settings — the batch job can afford far more of both. ## Diagnosing in production When someone reports bad results, do not start turning knobs. Take the failing query, run it against a flat scan, and check whether the expected document is in the exact top-k at all. If it is not, the problem is upstream — the embedding model or the chunking — and no index setting will help. If it is, check whether it appears anywhere in the approximate candidate set at a large `nprobes`: present means raise refinement, absent means raise probes. That three-way split turns a vague quality complaint into a specific, fixable cause.

  • A query returns the right document but ranked eighth instead of first. Which knob do you reach for?
    refine_factor. The document was retrieved, so the candidate generation worked and probing wider will not change the ordering. The misranking comes from distances computed on compressed codes; refinement re-reads the full-precision vectors for the over-fetched candidates and re-sorts them exactly, which is precisely this failure. Raising nprobes here would add latency and fix nothing.
  • Why is a high refine_factor riskier on an object-storage-backed table than on local disk?
    Refinement reads full-precision vectors for limit × refine_factor rows, and those reads are scattered rather than sequential. On NVMe a scattered read is microseconds; against object storage it is a network request with tens of milliseconds of latency and a per-request charge. The tail latency degrades before the mean does, so watch p99 and consider a smaller factor or caching hot vectors locally.

saying these in an interview costs you the question

  • Treating nprobes and refine_factor as interchangeable recall knobs
  • Expecting refinement to recover a neighbour never retrieved
  • Tuning recall without recording latency alongside it
  • Changing num_partitions and leaving nprobes untouched
  • Using one setting for both interactive and batch query paths

context