skip to content

What does Qdrant's optimizer indexing_threshold do, and why lower it during bulk load?

level: seniorimportance: should knowfreq 40%

answer

  1. a segment-size gate, measured in kilobytes
  2. small segments prefer brute force
  3. zero switches index building off
  4. standard trick around a big import
  5. restore it and wait for the optimizer

basics

~20 s

indexing_threshold, in OptimizersConfigDiff, is the segment size in kilobytes above which Qdrant builds an HNSW index; smaller segments are scanned exhaustively. Setting it to 0 during a bulk load suppresses index building so ingestion is not fighting constant rebuilds, then you restore it once.

solid answer

~50 s

Qdrant's optimizer decides per segment whether an HNSW graph is worth building. `models.OptimizersConfigDiff(indexing_threshold=20000)` means segments holding less than that many kilobytes of vector data are searched by brute force, because a full scan of a small segment beats building and walking a graph. During a large ingest, this default causes churn: segments cross the threshold, get indexed, then get merged and re-indexed, so a big share of CPU goes to graphs that are immediately discarded. The standard pattern is `update_collection` with `indexing_threshold=0` before the load — which disables index building entirely — then upload, then set it back to its normal value and let the optimizer build indexes once over settled segments. Searches during the load still work; they are just exhaustive and slow. Watch the collection status go from `grey`/optimizing back to `green` before trusting latency numbers.

code

python · 15 lines
python
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")

client.update_collection(
    collection_name="docs",
    optimizers_config=models.OptimizersConfigDiff(indexing_threshold=0),
)

# ... bulk upload runs here ...

client.update_collection(
    collection_name="docs",
    optimizers_config=models.OptimizersConfigDiff(indexing_threshold=20000),
)

go deeper

for a junior

Know that Qdrant builds vector indexes in the background per segment, and that very small segments are searched by scanning rather than through an index.

for a middle

Explain the threshold as a size gate in kilobytes and why brute force wins below it, and describe setting it to 0 to keep a bulk import from constantly rebuilding graphs.

for a senior

Demonstrate the full operational loop — suppress, load, restore, then confirm via get_collection that indexed vectors match point count before benchmarking or declaring the import done.

for a principal

Own the ingestion strategy: decide segment sizing and indexing policy per collection based on read/write mix, and set the expectation that latency SLOs are measured only on settled collections.

## The segment model A Qdrant collection is physically a set of segments. New points land in a fresh segment; a background optimizer merges small segments into larger ones, vacuums deleted points, and builds vector indexes. `models.OptimizersConfigDiff` is how you influence that machinery, on `create_collection` or `update_collection`. ## What indexing_threshold means `indexing_threshold` is a size in kilobytes of vector data. A segment below it is left unindexed and served by exhaustive scan; above it, the optimizer builds an HNSW graph for that segment. This is not a limitation — it is the correct trade. Brute-force scanning a few thousand vectors is fast, exact, and free of build cost; an HNSW graph over the same data costs CPU to build, adds memory, and returns approximate results. The threshold is simply where the crossover sits. The default is in the tens of megabytes of vector data, which for typical embedding sizes corresponds to tens of thousands of points. ## Why bulk loading fights it During a large ingest, segments are constantly being created, crossing the threshold, being indexed, then being merged into bigger segments — which invalidates the graph just built and triggers another build. A meaningful fraction of ingest-time CPU is spent producing indexes that are discarded minutes later, and that CPU is competing with the write path. The remedy is to tell the optimizer not to bother until the data has settled: 1. `client.update_collection(collection_name="docs", optimizers_config=models.OptimizersConfigDiff(indexing_threshold=0))` — zero disables index building. 2. Upload the data. 3. `client.update_collection(collection_name="docs", optimizers_config=models.OptimizersConfigDiff(indexing_threshold=20000))` — restore, and the optimizer now builds graphs once, over merged segments. During step 2 the collection remains queryable, but every search is a full scan, so latency will be poor and will degrade as data grows. That is the trade you are explicitly accepting. ## Knowing when it is done After restoring the threshold, indexing runs in the background. `client.get_collection(collection_name="docs")` reports the collection's status and counts — including how many points are actually indexed versus merely stored. Benchmarking before that settles produces numbers that describe a half-built index and are worthless. The single most common false alarm in Qdrant operations is "queries are slow after our import" measured while the optimizer is still working. ## The neighbouring knobs `default_segment_number` controls how many segments the optimizer aims for; 0 lets Qdrant choose based on CPU count. Fewer, larger segments give better search performance because a query fans out across segments and merges results — but they are more expensive to optimize and rebuild. More, smaller segments favour write throughput and parallel optimization. `memmap_threshold` is a size in kilobytes above which segment data is memory-mapped rather than held in RAM. It is a memory-pressure control that overlaps with the `on_disk` flags on vectors and on the HNSW graph; on modern versions setting storage explicitly with those flags is the clearer way to express intent. `max_segment_size` caps how large a segment may grow, `flush_interval_sec` governs how often data is flushed, and `deleted_threshold` / `vacuum_min_vector_number` decide when a segment carrying many deletions is rebuilt to reclaim space. That last pair matters for workloads that churn: heavy deletion without vacuuming leaves tombstoned points occupying memory and being traversed. ## Operational guidance Treat these as deployment configuration, not tuning fiddles. Set them deliberately for the collection's workload shape: a write-heavy pipeline wants more segments and a suppressed indexing threshold during loads; a read-heavy serving collection wants fewer, larger, fully indexed segments. Change them through `update_collection` during quiet periods, because every change hands the optimizer more background work, and always confirm the collection has returned to a settled state before drawing conclusions from latency measurements.

  • Can you still search the collection while indexing_threshold is 0?
    Yes. Unindexed segments are served by exhaustive scan, so results are actually exact — they are just slow, and they get slower linearly as the segment grows. That is fine for a load window and unacceptable as a steady state on a large collection. Restore the threshold when the import finishes and let the optimizer build graphs over the merged segments before you measure latency.
  • What does default_segment_number trade off?
    Search performance against optimization cost and write throughput. Fewer, larger segments mean fewer sub-searches to fan out and merge per query, so latency is better — but each segment is more expensive to rebuild or re-index. More, smaller segments parallelise ingestion and optimization well but add per-query overhead. Setting it to 0 lets Qdrant pick based on available CPUs, which is a reasonable default for mixed workloads.
  • A team reports slow queries right after importing 20 million points. What do you check first?
    Whether indexing has finished. Call `get_collection` and compare the indexed vector count against the total point count, and check the collection status — a collection still being optimized is serving a mix of indexed and brute-force segments, and any benchmark taken then describes a half-built index. Only after it settles should you look at `hnsw_ef`, quantization settings or the build parameters.

saying these in an interview costs you the question

  • Thinks indexing_threshold counts points rather than kilobytes
  • Believes unindexed segments cannot be searched at all
  • Benchmarks latency while the optimizer is still indexing
  • Leaves indexing_threshold at 0 after the bulk load finishes
  • Assumes the optimizer never re-indexes a segment once built

context