skip to content

Qdrant

A Rust vector database known for fast filtered search: HNSW combined with payload indexes so metadata filters are applied during traversal rather than after it. Interviewers ask about exactly that, because naive post-filtering is what silently destroys recall.

on this pageshow

questions

18

What are the parts of a Qdrant point, and which id types does it accept?

level: juniorimportance: must knowfreq 80%

answer

  1. id, vector, payload
  2. two accepted id types only
  3. strings are not free-form keys
  4. hash your key into a UUID
  5. upsert replaces, it does not merge

basics

~20 s

A Qdrant point is the unit of storage: an id, one or more vectors, and an arbitrary JSON payload. Ids must be an unsigned integer or a UUID — arbitrary strings are rejected. Upserting an existing id replaces that point.

solid answer

~50 s

A point is what you write and what search returns. It carries three things: an **id**, its **vector**, and a **payload** — free-form JSON attached to the point, used for filtering and for returning metadata with results. In the Python client you build one as `models.PointStruct(id=1, vector=[...], payload={"url": "...", "lang": "en"})` and write it with `client.upsert(collection_name="docs", points=[...])`. The id rule catches people out: Qdrant accepts only unsigned integers or UUIDs, so a natural key like `"doc-42"` is rejected. The usual workaround is to derive a deterministic UUID from your external key (for example `uuid.uuid5`) and keep the original string in the payload. `upsert` is insert-or-replace by id, which makes retries safe but also means a partial point overwrites the whole previous one; to change only metadata use `client.set_payload`, and to change only vectors use `client.update_vectors`.

code

python · 26 lines
python
import uuid
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")
NAMESPACE = uuid.UUID("12345678-1234-5678-1234-567812345678")

external_id = "docs/getting-started.md#chunk-3"
point_id = str(uuid.uuid5(NAMESPACE, external_id))

client.upsert(
    collection_name="docs",
    points=[
        models.PointStruct(
            id=point_id,
            vector=[0.1] * 1536,
            payload={"external_id": external_id, "lang": "en"},
        )
    ],
)

# metadata-only change: vector untouched
client.set_payload(
    collection_name="docs",
    payload={"lang": "de"},
    points=[point_id],
)

go deeper

for a junior

Be able to say that a point is an id plus a vector plus a JSON payload, and that ids must be an unsigned integer or a UUID rather than an arbitrary string.

for a middle

Explain that upsert replaces a point wholesale, and name the targeted alternatives — set_payload for metadata and update_vectors for vectors — plus retrieve and count for reading back.

for a senior

Show that you design ids for idempotent re-ingestion, keep payloads small enough not to bloat every search response, and can explain why deletes do not immediately reclaim disk.

for a principal

Own the id strategy as a system-wide contract: it determines whether reprocessing is safe, whether cross-store joins work, and whether a rebuild can be resumed halfway through.

## The unit of storage Everything in Qdrant is a point. A collection holds points; search returns points; filters select points. A point has exactly three parts. **The id.** A unique identifier within the collection. Qdrant accepts two id types only: an unsigned 64-bit integer, or a UUID (given as its string form, e.g. `"550e8400-e29b-41d4-a716-446655440000"`). Anything else — an arbitrary string like `"doc-42"`, an email address, a float — is rejected by the server. This surprises people who expect a document store's free-form keys. **The vector.** The embedding itself, whose length must match the collection's declared vector size. In a collection configured with named vectors, this is a dict of name to vector rather than a single list. **The payload.** Arbitrary JSON attached to the point: strings, numbers, booleans, arrays, nested objects, geo objects. The payload is what you filter on and what you return alongside search hits, so it typically holds the source text or a pointer to it, plus tenant, timestamp, language, permissions and similar attributes. ## Working with the id restriction Since your application almost certainly has its own keys — a URL, a document id, a chunk identifier — you need a mapping. Two patterns dominate. The first is a deterministic UUID: `uuid.uuid5(namespace, external_key)` produces the same UUID for the same input every time, so re-ingesting a document naturally overwrites its previous point instead of duplicating it. Store the original key in the payload so you can read it back and filter on it. The second is a monotonically increasing integer from your own pipeline, again with the external key in the payload. This is compact and fast, but you have to own the counter, and it makes idempotent re-ingestion harder: if you cannot recompute the same integer for the same document, you will create duplicates. The anti-pattern is a random UUID per write. It works, but the second ingest of the same document silently creates a second copy, and your collection grows with near-duplicate vectors that all match the same query. ## Upsert semantics `client.upsert(collection_name, points=[...])` is insert-or-replace keyed by id. If the id is new, the point is created. If it exists, the new point **replaces** the old one — this is not a merge. Send a point with a vector and no payload, and the previously stored payload is gone. The practical rule: build the complete point every time you upsert, or use the targeted update calls instead. - `client.set_payload(collection_name, payload={...}, points=[ids])` merges the given keys into existing payloads, leaving vectors and other keys alone. - `client.overwrite_payload(...)` replaces the payload wholesale. - `client.delete_payload(...)` removes named keys; `client.clear_payload(...)` empties it. - `client.update_vectors(...)` replaces vectors without touching the payload. The upside of replace-by-id semantics is idempotency: retrying a failed batch after a network error cannot create duplicates, because the same ids land on the same points. That property is what makes bulk loading over an unreliable link tractable. ## Reading points back Two calls fetch points without doing similarity search. `client.retrieve(collection_name, ids=[1, 2, 3])` fetches by id — the direct lookup path, useful for hydrating results or verifying a write. `client.scroll(...)` pages through points in id order and is the tool for exporting or auditing a whole collection. Both take `with_payload` and `with_vectors` flags; vectors are excluded by default because they dominate response size, and you should leave that default alone unless you genuinely need the numbers. `client.count(collection_name, exact=True)` gives an exact point count, which is the fastest sanity check after a load. ## Deleting `client.delete(collection_name, points_selector=models.PointIdsList(points=[1, 2]))` removes points by id. Deletes in Qdrant are initially soft — the point is marked as removed and stops appearing in results immediately, with the space reclaimed later during background optimization. That is why disk usage does not drop the instant you delete, a detail worth mentioning if an interviewer asks why a large delete did not free space. ## Payload discipline Because the payload is schemaless, nothing stops you storing megabytes of text in it. Qdrant will do it, and every search that returns those points now ships that text over the wire. Keep the payload to what you filter on and what you need in the result envelope; if you need large documents, store a key and fetch from your primary store. Payload keys are also case-sensitive and dotted paths address nested fields, so inconsistent key naming across ingest paths quietly breaks the filters that read them later.

  • Your ingest pipeline reprocesses the same documents nightly. How do you avoid duplicate points?
    Derive the point id deterministically from the document's stable external key — `uuid.uuid5` over a namespace plus that key. Re-ingesting then upserts onto the same id and replaces the point rather than adding a second one. Random UUIDs per run are the failure mode, because nothing links the new point to the old one.
  • I upserted a point with just a new vector and my metadata disappeared. Why?
    Because `upsert` replaces the whole point rather than merging fields. The stored payload was dropped along with the old vector. Either send the complete point every time, or use `update_vectors` to change vectors while leaving the payload in place, and `set_payload` for the reverse.
  • Why does disk usage stay flat right after deleting a large number of points?
    Deletes are applied as markers first: the points stop being returned immediately, but their storage is reclaimed later when background optimization rewrites the affected segments. So logical deletion is instant while physical space recovery is deferred.

saying these in an interview costs you the question

  • Thinks any string can be used as a point id
  • Believes upsert merges payload fields instead of replacing the point
  • Uses a fresh random UUID on every re-ingest of the same document
  • Stores whole documents in the payload and wonders why responses are huge
  • Assumes deleting points frees disk space immediately

context

open as a page

In Qdrant, how do you restrict a vector search to points matching a payload condition?

level: juniorimportance: must knowfreq 80%

basics

~10 s

Pass a filter to the search call. In the Python client that is query_points(..., query_filter=models.Filter(must=[models.FieldCondition(key="lang", match=models.MatchValue(value="en"))])). Only points whose payload satisfies the filter come back, still ranked by vector distance.

open as a page

In Qdrant, what do VectorParams size and distance fix at collection creation?

level: middleimportance: must knowfreq 78%

basics

~20 s

VectorParams pins a Qdrant collection's vector dimension and its distance metric for the collection's lifetime. Every upserted point must match that size exactly, and neither value can be changed later — switching embedding models means creating a new collection.

open as a page

In Qdrant, what do hnsw_config's m and ef_construct control, and what do they cost?

level: middleimportance: must knowfreq 72%

basics

~20 s

m is how many graph links Qdrant keeps per vector (default 16); ef_construct is how wide the neighbour search is while building (default 100). Raising m costs RAM permanently, raising ef_construct costs indexing time only.

open as a page

How do must, should, and must_not combine inside a Qdrant Filter?

level: middleimportance: must knowfreq 70%

basics

~20 s

Clauses are ANDed with each other: a point must satisfy every condition in must, at least one in should, and none in must_not. min_should raises the should bar from one to N. Nesting Filter objects inside should expresses OR-of-ANDs.

open as a page

In Qdrant, what does create_payload_index buy you and what breaks without it?

level: middleimportance: must knowfreq 62%

basics

~20 s

A payload index is a per-field, typed structure separate from the vector index. It makes filters selective and lets the query planner estimate how many points a filter matches. Without it results are still correct, but filtering degrades to scanning candidates and the planner picks blindly.

open as a page

Qdrant offers scalar, product, and binary quantization — what does each trade?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Scalar quantization stores int8 components for about 4x compression with small accuracy loss and is the safe default. Product quantization compresses far harder (up to 64x) at real accuracy cost and slower builds. Binary quantization keeps one bit per dimension — 32x, fastest, but only viable for high-dimensional embeddings and only with rescoring.

open as a page

Why does Qdrant apply payload filters during HNSW traversal instead of after the search?

level: seniorimportance: must knowfreq 66%

basics

~20 s

Post-filtering discards results after the fact, so a selective filter can leave a top-10 query with one hit or none. Qdrant instead estimates the filter's cardinality from payload indexes, then either scans the small matching subset exactly or traverses the graph accepting only matching points.

open as a page

When do Qdrant named vectors beat separate collections per embedding?

level: middleimportance: should knowfreq 45%

basics

~20 s

Named vectors let one Qdrant point carry several embeddings under different names, each with its own size and distance, sharing one id and one payload. Use them when the same entity is searched by different representations; use separate collections when the data has different lifecycles.

open as a page

How does Qdrant's scroll API differ from search when exporting all points?

level: middleimportance: should knowfreq 42%

basics

~20 s

Scroll enumerates stored points in id order with no query vector and no similarity scoring, returning a page plus an offset for the next call. Search ranks by distance to a query vector and is bounded by a limit, so it can never enumerate a collection.

open as a page

What does hnsw_ef in Qdrant's SearchParams do, and how do you tune it?

level: middleimportance: should knowfreq 58%

basics

~20 s

hnsw_ef is the per-request beam width: how many candidates Qdrant keeps while walking the HNSW graph. Higher values raise recall and latency roughly together, and because it is set per query, different endpoints can use different values.

open as a page

In Qdrant, why can two must conditions on an array payload match the wrong points?

level: middleimportance: should knowfreq 34%

basics

~20 s

Each condition on an array field is evaluated independently and matches if any element satisfies it, so two conditions can be satisfied by two different elements. To require one element to satisfy both, wrap them in models.NestedCondition with models.Nested.

open as a page

How do Qdrant aliases let you swap in a rebuilt collection with no downtime?

level: seniorimportance: should knowfreq 38%

basics

~20 s

An alias is a name that resolves to a collection and can be used wherever a collection name is accepted. Applying the delete and create operations in one update_collection_aliases call repoints it atomically, so readers move to the rebuilt collection with no gap and the old one stays as rollback.

open as a page

How do you load millions of points into Qdrant without timeouts or duplicates?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Stream the data through the client's batching helpers — upload_points or upload_collection with a batch_size — rather than one giant upsert. Set wait=False so writes are acknowledged rather than awaited, and rely on upsert's replace-by-id semantics to make retries duplicate-free.

open as a page

What does Qdrant's optimizer indexing_threshold do, and why lower it during bulk load?

level: seniorimportance: should knowfreq 40%

basics

~20 s

indexing_threshold, in OptimizersConfigDiff, is the segment size in kilobytes above which Qdrant builds an HNSW index; smaller segments are scanned exhaustively. Setting it to 0 during a bulk load suppresses index building so ingestion is not fighting constant rebuilds, then you restore it once.

open as a page

How do oversampling and rescore in Qdrant's QuantizationSearchParams recover recall?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Oversampling makes Qdrant fetch more candidates than requested using cheap quantized distances; rescoring then re-ranks those candidates with the original full-precision vectors and returns the true top results. Together they trade a little extra work for most of the accuracy quantization gave away.

open as a page

How would you fit 50M 1536-dim vectors in Qdrant on a fixed RAM budget?

level: principalimportance: should knowfreq 33%

basics

~20 s

Start from arithmetic: 50M x 1536 x 4 bytes is roughly 300GB of raw vectors plus graph overhead. Push originals to disk with on_disk=True, pin a quantized copy in RAM with always_ram=True, and buy back recall with oversampling and rescoring — then validate against exact search.

open as a page

When should Qdrant tenants be separated by payload filter rather than by separate collections?

level: principalimportance: should knowfreq 30%

basics

~20 s

Payload partitioning — one collection, a tenant_id in every payload, a keyword index marked is_tenant, and a mandatory filter on every query — scales to many small tenants. Separate collections suit few large tenants needing hard isolation or independent tuning.

open as a page