skip to content

What does hnsw_ef in Qdrant's SearchParams do, and how do you tune it?

level: middleimportance: should knowfreq 58%

answer

  1. set on the request, not the collection
  2. beam width during traversal
  3. recall and latency move together
  4. must not be smaller than limit
  5. diminishing returns reveal a weak graph

basics

~20 s

hnsw_ef is the per-request beam width: how many candidates Qdrant keeps while walking the HNSW graph. Higher values raise recall and latency roughly together, and because it is set per query, different endpoints can use different values.

solid answer

~50 s

`hnsw_ef` lives in `models.SearchParams` and is passed on each `query_points` call, so it is the one recall knob you can change without touching the collection or rebuilding anything. Mechanically it bounds the size of the candidate list held during traversal. A larger beam explores more of the graph before converging, so it finds more of the true nearest neighbours — and costs proportionally more distance computations and memory traffic. Latency grows roughly with the beam size, though sub-linearly at the low end. It must also be at least as large as `limit`; asking for 100 results with a beam of 32 cannot return 100 good ones. Tuning is empirical: build a held-out query set, get ground truth with `models.SearchParams(exact=True)`, then sweep `hnsw_ef` and pick the smallest value that hits your recall target at acceptable p95 latency. Because it is per-request, a batch re-index job can use a large value while an interactive endpoint uses a small one.

code

python · 9 lines
python
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")
hits = client.query_points(
    collection_name="docs",
    query=[0.02] * 768,
    limit=10,
    search_params=models.SearchParams(hnsw_ef=256),
).points

go deeper

for a junior

Know that hnsw_ef is passed per query inside SearchParams and that larger values mean more accurate but slower search. Mention that it should not be smaller than the number of results you ask for.

for a middle

Explain it as the traversal beam width and describe the concave recall curve — big early gains, then diminishing returns while latency keeps rising roughly linearly.

for a senior

Demonstrate the tuning loop: sample real queries, get ground truth with exact=True, sweep the value, and pick the knee against a p95 latency budget. Note that different endpoints can use different values on one collection.

for a principal

Frame it as the per-request quality dial that lets one index serve several SLOs, and set the policy for when recall regressions trigger a rebuild of the index rather than another bump in search effort.

## What the parameter is Qdrant's search-time effort knob is `hnsw_ef`, a field of `models.SearchParams` supplied per request: `client.query_points(collection_name="docs", query=vector, limit=10, search_params=models.SearchParams(hnsw_ef=256))` It is the size of the dynamic candidate list maintained while traversing the graph's bottom layer. The traversal keeps up to `hnsw_ef` best-so-far candidates, expands them, and stops when no candidate in the list can improve on the current results. A larger list means the walk explores more branches before deciding it has converged. ## Why it is the knob you reach for first Unlike the build-time settings, `hnsw_ef` changes nothing about stored data. There is no re-indexing, no config change, no downtime, and no permanent memory cost. That makes it uniquely safe to experiment with in production: you can A/B two values on live traffic and roll back by editing a request. It is also **per request**, which is the operationally interesting part. A single collection can serve an interactive search box at a small beam for tight p99 latency, an offline evaluation job at a very large beam for near-exact results, and a recommendation batch somewhere in between — all against the same index. Systems that bake search effort into the index cannot do this. ## The shape of the curve Recall against `hnsw_ef` is steeply concave. Going from a small beam to a moderate one usually buys a large recall jump; past a point, each doubling buys a fraction of a percent while latency keeps climbing roughly linearly. The practical method is to find the knee: 1. Sample a few thousand real query vectors from traffic. 2. Compute ground-truth neighbours once with `models.SearchParams(exact=True)`, which bypasses the graph and scans exhaustively. 3. For each candidate `hnsw_ef`, run the same queries, compute recall@k against ground truth, and record p50/p95 latency. 4. Choose the smallest value meeting the recall target with latency headroom. Re-run this when the embedding model changes. Recall is a property of the data distribution, not just the parameter — a new model with different dimensionality or clustering can move the knee substantially. ## Interactions worth knowing **With limit.** The beam must be at least as wide as the number of results requested, otherwise the walk cannot even hold `limit` candidates. In practice you want `hnsw_ef` comfortably above `limit`, not equal to it — deep pagination with a large `limit` implicitly demands a larger beam and thus more work per request. **With the build parameters.** `hnsw_ef` can only exploit the graph it was given. If `m` and `ef_construct` produced a poorly connected graph, raising the search beam shows diminishing returns early: you pay full latency and never reach the recall target. A recall curve that flattens well below your target at large beams is a signal to fix the build config, not to keep raising the beam. **With quantization.** When the collection is quantized, the traversal computes approximate distances, so raising `hnsw_ef` alone will not recover accuracy lost to compression; that is what oversampling and rescoring in `models.QuantizationSearchParams` are for. **With filters.** Filtered queries change the effective graph connectivity, so a beam width validated on unfiltered traffic can under-deliver on heavily filtered traffic. Validate the two separately if your workload mixes them. **With exact=True.** Setting `exact=True` in `SearchParams` skips approximate search entirely and does a full scan. It is a ground-truth and debugging tool, not a production setting for large collections — cost grows linearly with collection size. ## Defaults and hygiene If you omit `hnsw_ef`, Qdrant applies a server-side default. Relying on that default is fine for a prototype but weak for a production service: nobody has verified it meets your recall target on your data, and it hides a real quality dial from whoever tunes latency later. Set it explicitly, record the value alongside the recall number it was validated at, and treat it as part of the service's configuration.

  • You raise hnsw_ef tenfold and recall barely moves. What does that tell you?
    That the graph itself is the limit, not the search effort. A poorly connected index — too small an `m` or `ef_construct` for the data's dimensionality and clustering — flattens the recall curve early, so extra beam width buys latency and nothing else. The fix is on the build side: raise `ef_construct`, and `m` if needed, then re-measure. If the collection is quantized, the ceiling may instead be distance-approximation error, which oversampling and rescoring address.
  • How would you measure recall in production without a labelled dataset?
    Use exact search as the oracle. Sample real query vectors from live traffic, run each with `models.SearchParams(exact=True)` to get true top-k, then compare against the approximate results at your candidate settings. It is expensive per query, which is why you sample offline rather than run it inline. Re-run the sweep whenever the embedding model or the data distribution changes, since the recall curve is a property of the data.

saying these in an interview costs you the question

  • Thinks hnsw_ef is set on the collection and requires reindexing
  • Sets hnsw_ef below limit and expects full results
  • Believes a bigger beam fixes recall lost to quantization
  • Assumes the default value is validated for their data
  • Uses exact=True in production to be safe

context