When should you use Chroma's collection.get() instead of collection.query()?
answer
- One ranks, one retrieves
- No query point, no distances
- Pagination lives on only one of them
- Flat lists versus nested lists
- Enumerating with a huge n_results is the smell
basics
~20 sUse get() for exact retrieval: fetch by ids, or scan by where and where_document filters with limit and offset. It computes no embedding and returns no distances. Use query() when the question is similarity and you need ranked nearest neighbours.
solid answer
~50 s`get()` and `query()` answer different questions. `collection.get(ids=["a", "b"])` is a direct lookup, and `collection.get(where={"source": "faq"}, limit=100, offset=0)` is a filter scan — no query vector, no embedding call, no ranking, and no `distances` in the response. Results come back as flat lists rather than the per-query nesting `query()` uses, and `limit`/`offset` give you pagination for walking a collection. `query()` always needs a query point, embeds `query_texts` if that is what you gave it, and returns records ordered by distance. So: "show me every chunk from this document" or "fetch these five ids the LLM cited" is `get()`; "what is most similar to this question" is `query()`. Reaching for `query()` with a huge `n_results` just to enumerate a filtered subset is the common anti-pattern — it pays for an embedding and a ranking you then ignore.
code
python · 13 lines# Exact retrieval by id — no embedding, no ranking
cited = collection.get(ids=["chunk-17", "chunk-18"], include=["documents"])
print(cited["documents"]) # flat list, not nested
# Paging through a filtered slice
offset = 0
while True:
page = collection.get(where={"source": "faq"}, limit=500, offset=offset)
if not page["ids"]:
break
offset += len(page["ids"])
print("total in collection:", collection.count())go deeper
Know that get() fetches by id or by filter with no similarity involved, and query() finds nearest neighbours to a query point and returns distances.
Be able to name the concrete differences: get() has limit and offset for pagination, returns flat lists, and never calls the embedding function; query() nests results per query and always ranks.
Point out the cost angle — using query() to enumerate a filtered subset pays for an embedding call and the slowest shape of constrained search, then discards the ordering it bought.
Frame the split as intent in the API: describable requests go through similarity search, nameable requests through exact retrieval. Blurring them is what produces pipelines that are both expensive and quietly incorrect.
## Two different questions Chroma gives you two read paths, and the choice is not stylistic: - `query()` answers **"what is nearest to this point?"** It requires a query point (`query_texts` or `query_embeddings`), it ranks by distance, and `n_results` caps how many neighbours come back. - `get()` answers **"which records match these exact criteria?"** It takes `ids`, `where`, and `where_document`, has no notion of a query point, and returns matches with no ranking. Both accept `include` to trim the payload, and both accept the same filter syntax. The differences that matter are the query point, the ordering, and the response shape. ## What get() gives you that query() does not **Id lookup.** `get(ids=["chunk-17", "chunk-18"])` is the natural way to resolve identifiers back to content. In a RAG pipeline this is how you fetch the chunks a downstream component referenced, or how you show a source document after the user clicks a citation. **Pagination.** `limit` and `offset` let you walk a collection or a filtered slice in pages. That is how you export data, audit what was ingested, or build an admin view. `query()` has no offset — it has `n_results`, which is a top-K cap, and "the next 100 nearest" is not a thing you can ask for. **No embedding cost.** `get()` never calls the embedding function. With a hosted embedding model, that difference is a network round trip and real money per call, so using `query()` to enumerate records is directly wasteful. **A flat response.** `get()` returns `{"ids": [...], "documents": [...], "metadatas": [...]}` — one level. `query()` nests one level per query because its inputs are batched. Code that copies an indexing pattern between the two and keeps a stray `[0]` silently reads only the first record, and it is a bug that survives review easily. ## What get() cannot do It cannot rank. There is no semantic ordering, no distance, and no "closest first" — you get matches in the store's own order. If relevance matters at all, `get()` is the wrong tool. It also has no default cap the way `query()` does with `n_results=10`. A `get()` with a broad `where` and no `limit` can pull a very large amount of data into memory, so treat `limit` as mandatory on anything that is not an id lookup, and page with `offset`. ## The anti-pattern to name The mistake worth calling out in an interview is `query(query_texts=["anything"], n_results=10000, where={...})` used to enumerate a filtered subset. It embeds a meaningless query string, forces a ranked nearest-neighbour search over a heavily constrained set — the exact shape that is slowest — and then the caller ignores the ordering entirely. `get(where=..., limit=..., offset=...)` expresses the intent directly and does none of that work. The inverse mistake exists too: using `get(where_document={"$contains": "..."})` as a search feature. That is a substring test with no ranking and no semantic matching, so it misses paraphrases entirely and returns hits in arbitrary order. It is a filter, not a search. ## Related helpers Two small methods round this out. `collection.count()` returns the number of records without pulling any of them — the right way to check "did my ingest land?". `collection.peek()` returns the first handful of records for eyeballing structure during development. Neither is a substitute for `get()` with a filter, but both save you from writing one. ## A rule of thumb If you can name the records you want — by id, or by an exact predicate — use `get()`. If you can only describe them — "about connection timeouts" — use `query()`. The moment you find yourself sorting or truncating `query()` results by something other than distance, or ignoring the distances altogether, you probably wanted `get()`.
- How do you page through every record matching a filter?Use `get()` with `where`, a fixed `limit`, and an advancing `offset` — for example limit 500 and offset 0, 500, 1000, until a page comes back short. `query()` offers no offset; `n_results` is a top-K cap, so "the next 500 nearest" cannot be expressed. Always set a limit on a broad get, or you may pull the whole slice into memory.
- Why is get() cheaper than query() beyond just skipping the ranking?`get()` never invokes the collection's embedding function. With `query_texts` and a hosted embedding model, every call is a network round trip and a per-token charge before any search happens. For enumeration workloads that cost is pure waste, and it is often the dominant term in the latency of the call.
- Is where_document with $contains a reasonable substitute for keyword search?No. It is a literal substring filter with no tokenisation, no stemming and no scoring, so it misses paraphrases and morphological variants and returns matches unordered. It is useful to narrow a candidate set alongside a vector query, but presenting its output as search results gives users an arbitrary ordering and silent misses.
saying these in an interview costs you the question
- Using query() with a huge n_results to list records
- Expecting distances back from get()
- Indexing get() results with a leading [0] like query()
- Running a broad get() with no limit
- Treating where_document $contains as ranked keyword search