skip to content

What are the parts of a Qdrant point, and which id types does it accept?

level: juniorimportance: must knowfreq 80%

answer

  1. id, vector, payload
  2. two accepted id types only
  3. strings are not free-form keys
  4. hash your key into a UUID
  5. upsert replaces, it does not merge

basics

~20 s

A Qdrant point is the unit of storage: an id, one or more vectors, and an arbitrary JSON payload. Ids must be an unsigned integer or a UUID — arbitrary strings are rejected. Upserting an existing id replaces that point.

solid answer

~50 s

A point is what you write and what search returns. It carries three things: an **id**, its **vector**, and a **payload** — free-form JSON attached to the point, used for filtering and for returning metadata with results. In the Python client you build one as `models.PointStruct(id=1, vector=[...], payload={"url": "...", "lang": "en"})` and write it with `client.upsert(collection_name="docs", points=[...])`. The id rule catches people out: Qdrant accepts only unsigned integers or UUIDs, so a natural key like `"doc-42"` is rejected. The usual workaround is to derive a deterministic UUID from your external key (for example `uuid.uuid5`) and keep the original string in the payload. `upsert` is insert-or-replace by id, which makes retries safe but also means a partial point overwrites the whole previous one; to change only metadata use `client.set_payload`, and to change only vectors use `client.update_vectors`.

code

python · 26 lines
python
import uuid
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")
NAMESPACE = uuid.UUID("12345678-1234-5678-1234-567812345678")

external_id = "docs/getting-started.md#chunk-3"
point_id = str(uuid.uuid5(NAMESPACE, external_id))

client.upsert(
    collection_name="docs",
    points=[
        models.PointStruct(
            id=point_id,
            vector=[0.1] * 1536,
            payload={"external_id": external_id, "lang": "en"},
        )
    ],
)

# metadata-only change: vector untouched
client.set_payload(
    collection_name="docs",
    payload={"lang": "de"},
    points=[point_id],
)

go deeper

for a junior

Be able to say that a point is an id plus a vector plus a JSON payload, and that ids must be an unsigned integer or a UUID rather than an arbitrary string.

for a middle

Explain that upsert replaces a point wholesale, and name the targeted alternatives — set_payload for metadata and update_vectors for vectors — plus retrieve and count for reading back.

for a senior

Show that you design ids for idempotent re-ingestion, keep payloads small enough not to bloat every search response, and can explain why deletes do not immediately reclaim disk.

for a principal

Own the id strategy as a system-wide contract: it determines whether reprocessing is safe, whether cross-store joins work, and whether a rebuild can be resumed halfway through.

## The unit of storage Everything in Qdrant is a point. A collection holds points; search returns points; filters select points. A point has exactly three parts. **The id.** A unique identifier within the collection. Qdrant accepts two id types only: an unsigned 64-bit integer, or a UUID (given as its string form, e.g. `"550e8400-e29b-41d4-a716-446655440000"`). Anything else — an arbitrary string like `"doc-42"`, an email address, a float — is rejected by the server. This surprises people who expect a document store's free-form keys. **The vector.** The embedding itself, whose length must match the collection's declared vector size. In a collection configured with named vectors, this is a dict of name to vector rather than a single list. **The payload.** Arbitrary JSON attached to the point: strings, numbers, booleans, arrays, nested objects, geo objects. The payload is what you filter on and what you return alongside search hits, so it typically holds the source text or a pointer to it, plus tenant, timestamp, language, permissions and similar attributes. ## Working with the id restriction Since your application almost certainly has its own keys — a URL, a document id, a chunk identifier — you need a mapping. Two patterns dominate. The first is a deterministic UUID: `uuid.uuid5(namespace, external_key)` produces the same UUID for the same input every time, so re-ingesting a document naturally overwrites its previous point instead of duplicating it. Store the original key in the payload so you can read it back and filter on it. The second is a monotonically increasing integer from your own pipeline, again with the external key in the payload. This is compact and fast, but you have to own the counter, and it makes idempotent re-ingestion harder: if you cannot recompute the same integer for the same document, you will create duplicates. The anti-pattern is a random UUID per write. It works, but the second ingest of the same document silently creates a second copy, and your collection grows with near-duplicate vectors that all match the same query. ## Upsert semantics `client.upsert(collection_name, points=[...])` is insert-or-replace keyed by id. If the id is new, the point is created. If it exists, the new point **replaces** the old one — this is not a merge. Send a point with a vector and no payload, and the previously stored payload is gone. The practical rule: build the complete point every time you upsert, or use the targeted update calls instead. - `client.set_payload(collection_name, payload={...}, points=[ids])` merges the given keys into existing payloads, leaving vectors and other keys alone. - `client.overwrite_payload(...)` replaces the payload wholesale. - `client.delete_payload(...)` removes named keys; `client.clear_payload(...)` empties it. - `client.update_vectors(...)` replaces vectors without touching the payload. The upside of replace-by-id semantics is idempotency: retrying a failed batch after a network error cannot create duplicates, because the same ids land on the same points. That property is what makes bulk loading over an unreliable link tractable. ## Reading points back Two calls fetch points without doing similarity search. `client.retrieve(collection_name, ids=[1, 2, 3])` fetches by id — the direct lookup path, useful for hydrating results or verifying a write. `client.scroll(...)` pages through points in id order and is the tool for exporting or auditing a whole collection. Both take `with_payload` and `with_vectors` flags; vectors are excluded by default because they dominate response size, and you should leave that default alone unless you genuinely need the numbers. `client.count(collection_name, exact=True)` gives an exact point count, which is the fastest sanity check after a load. ## Deleting `client.delete(collection_name, points_selector=models.PointIdsList(points=[1, 2]))` removes points by id. Deletes in Qdrant are initially soft — the point is marked as removed and stops appearing in results immediately, with the space reclaimed later during background optimization. That is why disk usage does not drop the instant you delete, a detail worth mentioning if an interviewer asks why a large delete did not free space. ## Payload discipline Because the payload is schemaless, nothing stops you storing megabytes of text in it. Qdrant will do it, and every search that returns those points now ships that text over the wire. Keep the payload to what you filter on and what you need in the result envelope; if you need large documents, store a key and fetch from your primary store. Payload keys are also case-sensitive and dotted paths address nested fields, so inconsistent key naming across ingest paths quietly breaks the filters that read them later.

  • Your ingest pipeline reprocesses the same documents nightly. How do you avoid duplicate points?
    Derive the point id deterministically from the document's stable external key — `uuid.uuid5` over a namespace plus that key. Re-ingesting then upserts onto the same id and replaces the point rather than adding a second one. Random UUIDs per run are the failure mode, because nothing links the new point to the old one.
  • I upserted a point with just a new vector and my metadata disappeared. Why?
    Because `upsert` replaces the whole point rather than merging fields. The stored payload was dropped along with the old vector. Either send the complete point every time, or use `update_vectors` to change vectors while leaving the payload in place, and `set_payload` for the reverse.
  • Why does disk usage stay flat right after deleting a large number of points?
    Deletes are applied as markers first: the points stop being returned immediately, but their storage is reclaimed later when background optimization rewrites the affected segments. So logical deletion is instant while physical space recovery is deferred.

saying these in an interview costs you the question

  • Thinks any string can be used as a point id
  • Believes upsert merges payload fields instead of replacing the point
  • Uses a fresh random UUID on every re-ingest of the same document
  • Stores whole documents in the payload and wonders why responses are huge
  • Assumes deleting points frees disk space immediately

context