skip to content

How does Qdrant's scroll API differ from search when exporting all points?

level: middleimportance: should knowfreq 42%

answer

  1. no query vector involved
  2. ordered walk, not ranking
  3. returns a page plus an offset
  4. loop until the offset is None
  5. vectors are excluded by default

basics

~20 s

Scroll enumerates stored points in id order with no query vector and no similarity scoring, returning a page plus an offset for the next call. Search ranks by distance to a query vector and is bounded by a limit, so it can never enumerate a collection.

solid answer

~50 s

`client.scroll(collection_name="docs", limit=1000, with_payload=True, with_vectors=False)` returns a tuple: the page of points, and a `next_page_offset`. You loop, passing that offset back in as `offset=`, until it comes back `None`. There is no query vector involved — points come out ordered by id, so the traversal is complete and stable rather than ranked. That is the difference that matters: similarity search returns the `limit` best matches for a vector and has no notion of "the rest", so raising `top_k` to a huge number to dump a collection is both wrong and expensive. Scroll accepts `scroll_filter` to restrict which points are visited, and an `order_by` on a payload field if you need a different order — that requires an index on the field. Keep `with_vectors=False` unless you actually need the embeddings, since vectors dominate response size; the common uses are export, backfill, re-embedding and auditing.

code

python · 19 lines
python
from qdrant_client import QdrantClient

client = QdrantClient(url="http://localhost:6333")

offset = None
total = 0
while True:
    points, offset = client.scroll(
        collection_name="docs",
        limit=1000,
        with_payload=["doc_id", "lang"],
        with_vectors=False,
        offset=offset,
    )
    total += len(points)
    if offset is None:
        break

print(total, client.count("docs", exact=True).count)

go deeper

for a junior

Know that scroll pages through stored points without a query vector, and that it returns both the points and an offset to continue with.

for a middle

Explain the loop-until-offset-is-None pattern, the role of limit, with_payload and with_vectors, and why search with a giant limit is not an enumeration tool.

for a senior

Show you use scroll for real jobs — re-embedding, backfill, reconciliation — with checkpointed offsets for resumability, and that you know it is a live walk rather than a consistent snapshot.

for a principal

Own the choice between a live scroll and a snapshot for bulk data movement, weighing consistency guarantees, load on a serving cluster, and how much of the corpus you can afford to reprocess after a failure.

## Two different jobs Similarity search answers "which points are most like this vector?" — it takes a query vector, walks the index, and returns the top matches with scores. It is inherently bounded and inherently ranked. There is no meaningful continuation of it: asking for results 10,000 to 11,000 of an approximate nearest-neighbour search is neither cheap nor stable. Scroll answers a different question: "give me the points, all of them, in a defined order." No query vector, no scoring, no index traversal — just an ordered walk over stored points. Reaching for search to enumerate a collection is a category error, and it is one interviewers probe deliberately. ## The call and its cursor `client.scroll(...)` returns a **tuple** of `(points, next_page_offset)`. The first element is the page; the second is the value you feed back as `offset=` on the following call. When `next_page_offset` is `None`, you have reached the end. ``` offset = None while True: points, offset = client.scroll("docs", limit=1000, offset=offset) handle(points) if offset is None: break ``` The offset is a point id, not an opaque server-side session — nothing expires, and nothing has to be released. That means a scroll can be checkpointed: persist the last offset, and a crashed export resumes from there rather than starting over. It also means the paging is keyset-style rather than numeric-skip, so page N does not get more expensive than page 1. Because the walk is ordered by id and the offset is an id, concurrent writes behave predictably: a point inserted with an id before your current position will not be visited, and one inserted after it will. A scroll is not a transactional snapshot of the collection, and you should say so when asked — for a consistent point-in-time copy you take a snapshot instead. ## The parameters that matter `limit` is the page size. Bigger pages mean fewer round trips and larger responses; a few hundred to a few thousand is typical, tuned by payload size. `with_payload` and `with_vectors` control what comes back. Payload is included by default, vectors are not — and that default is right. Vectors are the bulk of the bytes, and most scroll use cases (auditing metadata, backfilling a field, exporting ids) do not need them. Turn `with_vectors=True` on only for genuine vector export, and drop the page size when you do. `with_payload` also accepts a list of field names, so you can fetch just the two keys you need instead of a fat payload. `scroll_filter` restricts the walk to points matching a condition, which is how you enumerate a subset — one tenant's points, or everything ingested before a date — without pulling the whole collection and discarding most of it client-side. `order_by` changes the traversal order to a payload field rather than id. It requires that field to have an index, and it is what you use when the export must be chronological rather than id-ordered. ## Related read calls Scroll is one of three ways to read without searching. `client.retrieve(collection_name, ids=[...])` fetches specific points by id — the right call when you already know what you want, and much cheaper than scrolling to find them. `client.count(collection_name, exact=True)` returns the point count; use it to know how big a scroll will be, or to verify one finished. And a collection snapshot is the right tool when you need a consistent copy rather than a live walk. ## Where scroll shows up in real work **Re-embedding.** Scroll the source collection with payload, run the new model over the stored text, and upsert into a new collection. Because scroll is resumable by offset, a job that dies at 60% picks up where it stopped. **Backfill.** A new payload field needs computing for existing points: scroll ids and the inputs, compute, then `set_payload` in batches. **Audit and reconciliation.** Compare the ids present in Qdrant against your primary datastore to find drift — points whose source rows were deleted, or rows never ingested. **Export and migration.** Dump points to a file for transfer to another environment when a snapshot is not appropriate — for instance when the target has a different collection configuration. ## Anti-patterns The headline one is `query_points(..., limit=1000000)` to "get everything". It forces an enormous approximate search, returns results in similarity order relative to an arbitrary vector, and gives no guarantee of completeness. Second is scrolling with `with_vectors=True` when only payload is needed, which multiplies transfer volume for nothing. Third is discarding `next_page_offset` and trying to page with an incrementing counter — the offset is an id, and an integer skip is not what this API takes.

  • Is a scroll a consistent snapshot of the collection?
    No. It is a live ordered walk, so writes that happen while it runs may or may not be seen depending on where their ids fall relative to your current offset. If you need a point-in-time consistent copy — for backup or migration — take a collection snapshot instead and read from that.
  • How do you resume an export that crashed halfway through?
    Persist the last `next_page_offset` you successfully processed and pass it as `offset` on restart. The offset is a point id rather than a server-side session, so nothing has expired and no state was held on the server while you were gone.
  • Why is with_vectors left off by default?
    Because vectors dominate response size — a thousand 1536-dimension float vectors is several megabytes per page, while the payload fields you usually want are a fraction of that. Most scroll workloads audit or backfill metadata and never need the embeddings, so paying for them by default would be the wrong trade.

saying these in an interview costs you the question

  • Uses search with a huge limit to dump an entire collection
  • Ignores next_page_offset and pages with an incrementing integer
  • Assumes scroll gives a transactionally consistent snapshot
  • Requests vectors on every page when only payload is needed
  • Thinks scroll ranks results by similarity

context