skip to content

How do you load millions of points into Qdrant without timeouts or duplicates?

level: seniorimportance: should knowfreq 48%

answer

  1. stream, do not materialize
  2. let the client split batches
  3. acknowledged versus applied
  4. deterministic ids make retries safe
  5. verify with an exact count

basics

~20 s

Stream the data through the client's batching helpers — upload_points or upload_collection with a batch_size — rather than one giant upsert. Set wait=False so writes are acknowledged rather than awaited, and rely on upsert's replace-by-id semantics to make retries duplicate-free.

solid answer

~50 s

Three things carry a bulk load. **Batch**: use `client.upload_points(collection_name, points=<iterable of PointStruct>, batch_size=256, parallel=4)` or `client.upload_collection(...)`, which consume a generator and split it into requests for you, so memory stays flat and no single HTTP call carries a million vectors. **Do not wait**: `wait=False` returns as soon as the operation is accepted into the write-ahead log instead of blocking until it is applied and searchable, which roughly doubles throughput; set `wait=True` only for the final batch or when the next step reads what you just wrote. **Make retries safe**: `upsert` replaces by id, so re-sending a batch after a timeout can never duplicate points — provided your ids are derived deterministically from the source data rather than randomly generated. Beyond that, `QdrantClient(prefer_grpc=True)` cuts serialization overhead noticeably at this volume, and after the load `client.count(collection_name, exact=True)` plus the collection status tell you whether everything landed and whether indexing has caught up.

code

python · 23 lines
python
import uuid
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333", prefer_grpc=True, timeout=120)
NS = uuid.UUID("12345678-1234-5678-1234-567812345678")

def points(rows):
    for row in rows:
        yield models.PointStruct(
            id=str(uuid.uuid5(NS, row["doc_id"])),
            vector=row["embedding"],
            payload={"doc_id": row["doc_id"], "lang": row["lang"]},
        )

client.upload_points(
    collection_name="docs",
    points=points(stream_rows()),
    batch_size=256,
    parallel=4,
    wait=False,
)

print(client.count("docs", exact=True))

go deeper

for a junior

Know that large loads are sent in batches rather than one request, and that the client offers upload_points to do the splitting for you.

for a middle

Explain the wait flag — accepted into the write-ahead log versus applied and searchable — and why upsert's replace-by-id behaviour makes a retried batch safe.

for a senior

Show the production shape: a streaming generator, deterministic ids for resumability, gRPC and a raised timeout, then verification by exact count and collection status before traffic is switched.

for a principal

Own ingest as a repeatable, restartable pipeline: checkpointing, an id scheme that survives reprocessing, and an explicit policy for serving queries from a collection that is still catching up.

## Why the naive version fails The obvious code — build a list of a million `PointStruct`s and pass it to one `upsert` — fails in three ways at once. It holds the entire dataset in client memory; it serializes one enormous request body; and it produces a single call that will hit a request timeout, after which you have no idea how much of the batch was applied. Bulk loading is therefore about streaming, acknowledgement semantics and idempotency. ## The batching helpers The Python client provides two upload calls that exist precisely for this. `client.upload_points(collection_name, points=iterable, batch_size=..., parallel=...)` accepts an **iterable or generator** of `PointStruct` objects and splits it into requests of `batch_size` points. Because it consumes lazily, you can stream straight from a file, a database cursor or an embedding pipeline without materializing everything. `client.upload_collection(collection_name, vectors=..., payload=..., ids=...)` takes parallel sequences — including a numpy array of vectors — which suits an offline job that already has embeddings in a matrix. Both take `parallel`, which spreads the work across worker processes. That helps when the bottleneck is client-side serialization rather than the server, but the data being sent must be picklable, and raising it beyond the server's ability to absorb writes just converts throughput into queueing. Start at 2 to 4 and measure. Batch size is a tradeoff between per-request overhead and blast radius: too small and you pay round-trip cost per point; too large and a single failure or timeout costs you a lot of redone work. Something in the low hundreds is a sane starting point for typical embedding sizes, tuned by watching request latency. ## wait: acknowledged versus applied Every write operation accepts a `wait` flag, and understanding it is the crux of this question. With `wait=False`, the server responds as soon as the operation has been **accepted** — written to the write-ahead log and queued for application. The call returns quickly and the write is durable, but the points are not guaranteed to be visible to a search issued immediately afterwards. With `wait=True`, the response comes only once the operation has been **applied** and the points are searchable. For a bulk load, `wait=False` is what you want: it lets the server pipeline application of writes while the client keeps streaming, and it stops slow segment work from stalling the loader. The Python client's `upsert` defaults to waiting, so this is an explicit choice you make, not a default you inherit. The rule of thumb is to load with `wait=False` and then, before any verification step or before flipping traffic over, do a waiting write or poll the collection until counts and status settle. Tests that write with `wait=False` and immediately assert on search results are a classic flaky-test source. ## Idempotency is the safety net Because `upsert` is keyed by point id and replaces rather than appends, retrying a batch that timed out is safe: the ids that already landed are simply overwritten with identical content, and the ones that did not are created. There is no duplicate-key error to handle and no compensating delete to write. This property is only real if your ids are **deterministic functions of the source data**. Derive them from the document's stable key — for instance a UUIDv5 over `"<doc_id>#<chunk_index>"` — and a resumed or re-run job converges on the same collection. Generate a random UUID per point instead, and every retry doubles part of your corpus with near-identical vectors that all match the same queries. This is the single most common way a bulk load quietly corrupts relevance. The same reasoning makes the job restartable: because re-writing a point is harmless, a crashed loader can restart from the last checkpointed offset — or even from the beginning — without a cleanup step. ## Transport and verification At this volume, transport matters. `QdrantClient(url=..., prefer_grpc=True)` uses the gRPC interface, which serializes large float arrays far more efficiently than JSON over REST and measurably raises sustained ingest rate. Also set a generous client `timeout` — the default is tuned for interactive queries, not for a batch that the server may take seconds to accept under load. When the load finishes, verify rather than assume. `client.count(collection_name, exact=True)` gives the exact point count to compare against your source. `client.get_collection(name)` returns the collection status along with `points_count` and `indexed_vectors_count`; a collection that is still working through the ingest reports a non-green status, and search served during that window may have lower recall than the finished state. Sampling a few known ids with `client.retrieve` confirms that payloads landed in the shape you expected — a check that catches serialization bugs no counter will. ## Failure modes to name in an interview Wrong-dimension vectors reject the entire request, so a single bad row can fail a whole batch — validate shapes before sending or narrow the batch to find the offender. Payloads that embed whole documents inflate every request and every later search response. And a loader that logs nothing per batch leaves you unable to answer "where did it stop", which is the question you will actually be asked at 3am.

  • Your integration test writes points and then immediately searches for them, and it fails about one run in five. What is going on?
    The write almost certainly used `wait=False`, so it returned once the operation was accepted rather than applied, and the search ran before the points became searchable. Fix it by writing that test fixture with `wait=True`, or by polling until the expected count is reached. Do not fix it with a sleep.
  • Halfway through a 10-million-point load the job crashes. What do you do?
    Restart it — ideally from the last checkpointed offset, but a full re-run is also correct. Because upsert replaces by id and the ids are derived deterministically from the source rows, re-writing points that already landed is a no-op in effect. Afterwards compare `count(exact=True)` with the source row count.
  • Does raising parallel always increase ingest throughput?
    No. It helps when the client is the bottleneck — serialization and request construction across worker processes. Once the server is saturated, more parallelism just deepens queues and raises latency and memory pressure without moving the ingest rate. Measure the rate at 1, 2 and 4 and stop where the curve flattens.

saying these in an interview costs you the question

  • Sends a single upsert containing the entire dataset
  • Uses random UUIDs so retried batches create duplicate points
  • Thinks wait=False means the write may be lost
  • Searches immediately after a wait=False write and expects results
  • Declares the load finished without comparing an exact count to the source

context