skip to content

HNSW Config and Quantization

Qdrant exposes HNSW build and search parameters plus scalar, product, and binary quantization directly. Tuning these is how you trade memory and latency against recall on a real cluster.

on this pageshow

questions

6

In Qdrant, what do hnsw_config's m and ef_construct control, and what do they cost?

level: middleimportance: must knowfreq 72%

answer

  1. two build-time knobs, one search-time
  2. links per node versus beam width
  3. defaults are 16 and 100
  4. one costs RAM forever, one costs CPU once
  5. changing them triggers background reindexing

basics

~20 s

m is how many graph links Qdrant keeps per vector (default 16); ef_construct is how wide the neighbour search is while building (default 100). Raising m costs RAM permanently, raising ef_construct costs indexing time only.

solid answer

~50 s

Both are **build-time** parameters of `models.HnswConfigDiff`, passed to `create_collection` or later to `update_collection`. `m` is the number of bidirectional links kept per vector on the upper graph layers (default 16). It is the main driver of index quality and of *permanent* memory: every stored vector carries roughly `m` link slots, so doubling `m` roughly doubles the graph's RAM footprint and also slows queries slightly because each hop examines more neighbours. `ef_construct` (default 100) is how many candidates the builder keeps in its beam while choosing those links. Higher values produce a better-connected graph and higher recall, but cost only indexing CPU time — not steady-state memory. So the usual advice is: raise `ef_construct` freely if build time allows, raise `m` only when high-dimensional or clustered data still shows poor recall, and remember that changing either forces Qdrant's optimizer to rebuild the index for existing segments.

code

python · 8 lines
python
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")
client.create_collection(
    collection_name="docs",
    vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE),
    hnsw_config=models.HnswConfigDiff(m=32, ef_construct=256),
)

go deeper

for a junior

Know that Qdrant's HNSW index has tunable settings, that m and ef_construct are set when the collection is created, and that the defaults are 16 and 100. Say plainly that higher values mean better recall for more cost.

for a middle

Explain the mechanics: m is links per vector and drives permanent memory, ef_construct is the build-time beam width and drives only indexing CPU. Be explicit that neither is a per-query parameter.

for a senior

Show judgment about changing them on a live collection — background re-indexing load, mixed segment behaviour during the rebuild, and validating recall against exact ground truth on a held-out query set before you commit.

for a principal

Own the trade at fleet scale: RAM per point times collection size is a hard budget line, so argue m from measured recall targets and cost, not intuition, and decide which collections deserve denser graphs at all.

## Where these parameters live Qdrant builds an HNSW graph per segment. Its build-time configuration is `models.HnswConfigDiff`, which you can pass when the collection is created: `client.create_collection(collection_name="docs", vectors_config=models.VectorParams(size=768, distance=models.Distance.COSINE), hnsw_config=models.HnswConfigDiff(m=32, ef_construct=256))` or apply later with `client.update_collection(collection_name="docs", hnsw_config=models.HnswConfigDiff(...))`. The same struct can also be attached per named vector inside `VectorParams`, which matters when one collection holds vectors of very different dimensionality. ## m — links per vector `m` (default 16) is the maximum number of bidirectional edges each vector keeps on the graph layers above the bottom one; the bottom layer typically gets about `2 * m`. Three consequences follow. **Memory.** Links are stored alongside the vectors, so the graph overhead grows linearly with `m`. For a collection where the raw vectors themselves are large, the graph is a modest addition; for short vectors or heavily quantized collections it can become a significant share of RAM. This cost is permanent — it is paid for every point, forever, not just during the build. **Recall.** A denser graph has more escape routes out of local minima, so a traversal is less likely to get stuck far from the true nearest neighbours. Data that is highly clustered, or embeddings of 1000+ dimensions, benefit most from `m` above the default; well-spread 384-dimensional embeddings often do not improve measurably past 16. **Latency.** Every hop evaluates the neighbours of the current node, so more links means more distance computations per hop. Raising `m` can therefore make queries slightly slower even as recall improves — the opposite direction from what people expect. ## ef_construct — beam width during the build `ef_construct` (default 100) is the size of the dynamic candidate list the builder maintains when it inserts a point and decides which `m` neighbours to keep. A wider beam means the builder considers more candidates and picks better links, producing a graph that supports higher recall at any given search effort. The key property is that this cost is paid **once**, at index build. It does not change the size of the resulting graph — that is fixed by `m` — and it does not change query cost. So `ef_construct` is the cheap knob: if indexing throughput is acceptable, raising it to 200-512 is a low-risk way to buy recall. The trade appears only when you are ingesting continuously and cannot afford the extra CPU in the optimizer. ## Search time is a different knob Neither parameter is per-query. The search-side beam width is `hnsw_ef` in `models.SearchParams`, supplied on the request. Confusing `ef_construct` with `hnsw_ef` is the single most common mistake here: setting `ef_construct` high and expecting each query to get slower or more accurate is wrong — an already-built graph does not consult it again. ## Related fields in the same struct `HnswConfigDiff` also carries `full_scan_threshold` (segments whose vector data is smaller than this many kilobytes are scanned exhaustively instead of traversed — exact results are cheaper than a graph walk at small sizes), `max_indexing_threads` (cap on optimizer parallelism, 0 meaning automatic), and `on_disk` (store the graph itself memory-mapped rather than in RAM, trading latency for footprint). ## Changing the values afterwards An `update_collection` call that alters `m` or `ef_construct` does not rewrite the existing graph in place. It changes the collection's config, and segments are re-indexed by the optimizer in the background as they are touched or rebuilt. On a large collection this is a heavy, sustained CPU and disk load, and query performance during it is a mix of old and new segments. Plan such a change like a migration, not like a config tweak. ## How to choose in practice Measure rather than guess. Hold out a few thousand queries, compute exact ground truth once (a search with `models.SearchParams(exact=True)` gives it directly), then sweep configurations and plot recall@k against p95 latency and resident memory. Start from the defaults; if recall is short, raise `ef_construct` first because it is free at query time, and only reach for `m` when a wider build beam has stopped helping.

  • If recall is too low, which of the two would you raise first and why?
    `ef_construct`, because its cost is one-time indexing CPU: the resulting graph is no bigger and queries are no slower, so a bad outcome costs only build time. `m` is the second move — it permanently increases RAM per point and adds distance computations per hop, so it should be justified by measurement, typically on high-dimensional or strongly clustered embeddings where a wider build beam has stopped helping.
  • What does setting on_disk=True inside HnswConfigDiff change?
    It stores the HNSW graph memory-mapped on disk instead of holding it in RAM. That cuts the resident footprint of the index structure, at the cost of page faults during traversal — each hop may hit storage, so latency becomes sensitive to disk speed and page cache. It is a reasonable trade on NVMe for large, latency-tolerant collections, and a poor one on network storage.
  • You raise m from 16 to 64 on a 100M-point collection. What actually happens on the cluster?
    Nothing instantly: `update_collection` records the new config, and the optimizer re-indexes segments in the background. Expect prolonged CPU and disk load, mixed old/new segment behaviour during the rebuild, and a permanently larger memory footprint afterwards. Treat it as a migration — do it during a low-traffic window, watch resident memory headroom, and validate recall on a held-out query set before and after.

m is how many roads you build out of each town — permanent infrastructure. ef_construct is how many routes the surveyor considers before choosing them: expensive to plan, but it does not widen the roads.

saying these in an interview costs you the question

  • Says ef_construct makes each query slower or more accurate
  • Thinks m or ef_construct can be set per request
  • Assumes raising m is free because it only affects the graph
  • Claims update_collection instantly rebuilds the whole index
  • Confuses ef_construct with hnsw_ef

context

open as a page

Qdrant offers scalar, product, and binary quantization — what does each trade?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Scalar quantization stores int8 components for about 4x compression with small accuracy loss and is the safe default. Product quantization compresses far harder (up to 64x) at real accuracy cost and slower builds. Binary quantization keeps one bit per dimension — 32x, fastest, but only viable for high-dimensional embeddings and only with rescoring.

open as a page

What does hnsw_ef in Qdrant's SearchParams do, and how do you tune it?

level: middleimportance: should knowfreq 58%

basics

~20 s

hnsw_ef is the per-request beam width: how many candidates Qdrant keeps while walking the HNSW graph. Higher values raise recall and latency roughly together, and because it is set per query, different endpoints can use different values.

open as a page

What does Qdrant's optimizer indexing_threshold do, and why lower it during bulk load?

level: seniorimportance: should knowfreq 40%

basics

~20 s

indexing_threshold, in OptimizersConfigDiff, is the segment size in kilobytes above which Qdrant builds an HNSW index; smaller segments are scanned exhaustively. Setting it to 0 during a bulk load suppresses index building so ingestion is not fighting constant rebuilds, then you restore it once.

open as a page

How do oversampling and rescore in Qdrant's QuantizationSearchParams recover recall?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Oversampling makes Qdrant fetch more candidates than requested using cheap quantized distances; rescoring then re-ranks those candidates with the original full-precision vectors and returns the true top results. Together they trade a little extra work for most of the accuracy quantization gave away.

open as a page

How would you fit 50M 1536-dim vectors in Qdrant on a fixed RAM budget?

level: principalimportance: should knowfreq 33%

basics

~20 s

Start from arithmetic: 50M x 1536 x 4 bytes is roughly 300GB of raw vectors plus graph overhead. Push originals to disk with on_disk=True, pin a quantized copy in RAM with always_ram=True, and buy back recall with oversampling and rescoring — then validate against exact search.

open as a page