How do oversampling and rescore in Qdrant's QuantizationSearchParams recover recall?
answer
- two switches, one on the request
- cast a wider net, then check properly
- the multiplier applies to limit
- cheap pass wide, expensive pass narrow
- binary needs a much wider net than int8
basics
~20 sOversampling makes Qdrant fetch more candidates than requested using cheap quantized distances; rescoring then re-ranks those candidates with the original full-precision vectors and returns the true top results. Together they trade a little extra work for most of the accuracy quantization gave away.
solid answer
~40 sQuantized distances are approximate, so the top-k they produce is roughly right but locally misordered — and sometimes a true neighbour lands just outside k. `models.QuantizationSearchParams(oversampling=3.0, rescore=True)`, passed inside `models.SearchParams`, fixes both halves. `oversampling=3.0` with `limit=10` makes Qdrant retrieve about 30 candidates using the fast quantized representation; `rescore=True` then reads the original vectors for those candidates, computes exact distances, and returns the best 10. Cost is asymmetric in a useful way: the wide part of the search runs on the small in-memory quantized copy, and only a few dozen full-precision reads are needed — which is why originals can sit on disk with `on_disk=True`. Rescoring is on by default; `rescore=False` skips it for maximum speed, and `ignore=True` bypasses quantization for a query entirely. Tune oversampling by sweeping it against recall measured with `exact=True`.
code
python · 15 linesfrom qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
hits = client.query_points(
collection_name="docs",
query=[0.02] * 1536,
limit=10,
search_params=models.SearchParams(
hnsw_ef=128,
quantization=models.QuantizationSearchParams(
oversampling=3.0,
rescore=True,
),
),
).pointsgo deeper
Know that quantized distances are approximate, and that Qdrant can fetch extra candidates and then re-check them against the original vectors before returning results.
Explain the two-phase mechanism: oversampling widens the candidate set during the cheap quantized pass, rescore re-ranks that set with full-precision vectors, and both are per-request settings.
Show the tuning loop and the storage interaction — sweep oversampling against exact-search ground truth, and recognise that rescoring turns each query into a bounded set of random reads whose cost depends on the storage class.
Own the end-to-end budget: this pattern is what allows vector RAM to grow sub-linearly with the corpus, so decide the recall floor, the storage class that makes rescoring affordable, and when a query path should bypass quantization entirely.
## The problem being solved When a Qdrant collection is quantized, the HNSW traversal computes distances on the compressed representation. Compression means the computed distance is an estimate of the true one. Two failure modes follow. First, **misordering**: the k results returned are mostly the right points, but their relative ranks are wrong, which matters when downstream logic cares about the top result or feeds scores into a reranker. Second, **misses**: a genuine nearest neighbour is estimated slightly too far away and falls out of the top k entirely. No amount of re-sorting the returned k can bring it back. Oversampling addresses the misses; rescoring addresses the misordering. They are designed to be used together. ## The mechanism Both live in `models.QuantizationSearchParams`, nested inside `models.SearchParams` on the request: `client.query_points(collection_name="docs", query=vector, limit=10, search_params=models.SearchParams(quantization=models.QuantizationSearchParams(oversampling=3.0, rescore=True)))` `oversampling` is a multiplier on `limit`. With `limit=10` and `oversampling=3.0`, Qdrant collects roughly 30 candidates during the quantized traversal. The candidate net is wider, so a neighbour whose estimated distance was slightly inflated is still caught. `rescore` then loads the **original** vectors for exactly those candidates, computes true distances, sorts, and truncates to `limit`. The final ordering is therefore exact with respect to the candidate set — the only remaining error is a true neighbour that never made it into the candidates at all, which is precisely what a larger oversampling factor reduces. ## Why the cost is acceptable The expensive-looking part — reading full-precision vectors — happens for only a few dozen points, after the graph walk is done. The graph walk itself, which touches orders of magnitude more vectors, runs entirely on the compressed copy. That asymmetry is the architectural point of Qdrant's quantization design. It is what makes the canonical configuration work: `models.VectorParams(size=1536, distance=models.Distance.COSINE, on_disk=True)` for the originals, plus a quantization config with `always_ram=True`. Traversal never touches disk; rescoring touches it for a bounded, small number of random reads. On NVMe that is a few milliseconds; on network storage it is not, which is why storage class matters for this pattern. ## Tuning Oversampling is empirical and interacts with the quantization mode. Scalar quantization's error is small enough that a factor near 1.5-2.0 often suffices. Binary quantization has much coarser distances and typically needs substantially more — factors of 3 or higher are common, and it is effectively unusable with oversampling at 1.0. Product quantization sits between them and depends on the compression ratio chosen. The method mirrors any other recall tuning: sample real query vectors, get ground truth with `models.SearchParams(exact=True)`, then sweep the oversampling factor and record recall@k against p95 latency. Pick the smallest factor that clears your recall floor. Re-run when the embedding model or the quantization mode changes. ## The other switches `rescore=False` skips the full-precision pass entirely. It is the right choice only when the compressed distances are good enough on their own — plausible with mild scalar quantization on a latency-critical path, dangerous with binary. Turning it off is also how you would isolate, during debugging, how much of your recall is coming from rescoring. `ignore=True` tells the query to bypass the quantized representation and use original vectors throughout. That gives the accuracy of an unquantized collection at the cost of losing the memory-locality benefit for that query — useful for a small number of high-value queries, or for comparing quantized against unquantized behaviour on the same index without rebuilding anything. ## Common mistakes The most common is enabling quantization and never touching these parameters, then reporting that "quantization destroyed our recall". Half the recall in a well-tuned quantized setup comes from rescoring. The second is pairing binary quantization with default oversampling and concluding binary quantization does not work — the candidate net was simply too narrow for how coarse the distances are. The third is putting the original vectors on slow shared storage. Rescoring turns every query into a handful of random reads; on high-latency storage those reads dominate the request and the whole design inverts from a win into a bottleneck.
- Why does the rescoring pass not blow up latency even when the original vectors are on disk?Because it runs only on the candidate set — with `limit=10` and `oversampling=3.0` that is about thirty vectors, not the millions the graph walk touched. The wide phase runs entirely on the in-memory compressed copy; disk sees a bounded number of random reads at the end. On NVMe that costs single-digit milliseconds. On high-latency network storage the same design inverts and rescoring becomes the bottleneck, so storage class is part of the decision.
- When would you deliberately set rescore=False?When the compressed distances are already accurate enough and latency is the binding constraint — mild scalar quantization on a hot path is the plausible case, especially if a downstream reranker will reorder the results anyway. It is a poor choice under binary or high-ratio product quantization, where the raw ordering is coarse. It is also useful diagnostically: toggling it shows exactly how much recall the rescoring pass is contributing.
- Your recall dropped after enabling binary quantization with default settings. What do you change first?Raise `oversampling` and confirm `rescore` is on. Binary distances are coarse, so a narrow candidate net simply does not contain the true neighbours and no re-ranking can recover them; factors of 3 or more are common. Measure with a held-out query set against `exact=True` ground truth while sweeping the factor. If recall stays short even at wide oversampling, the embeddings are likely too low-dimensional for binary quantization and scalar is the correct fallback.
Sift a large scoop of gravel through a coarse screen to grab plenty of candidate stones, then weigh only those few on a precision scale. The coarse pass is fast and generous; the accurate pass is slow but tiny.
saying these in an interview costs you the question
- Thinks rescoring re-runs the whole search at full precision
- Enables quantization but never sets oversampling
- Believes raising hnsw_ef compensates for quantization error
- Assumes binary quantization needs the same oversampling as int8
- Puts original vectors on slow shared storage and expects cheap rescoring