skip to content

What does Chroma's include parameter control in query(), and what comes back by default?

level: middleimportance: should knowfreq 55%

answer

  1. Chooses fields, not rows
  2. Ids always ride along
  3. One field is opt-in for size reasons
  4. Query default includes distances, get does not
  5. Lower distance means closer

basics

~20 s

include selects which per-record fields the response carries: documents, metadatas, distances, embeddings, uris, data. Query defaults to documents, metadatas and distances; embeddings are never returned unless asked for. Ids always come back and need not be requested.

solid answer

~50 s

`include` is a payload selector, not a filter — it changes what each result row carries, never which rows you get. For `query()` the default is `["documents", "metadatas", "distances"]`; `get()` defaults to `["documents", "metadatas"]` and has no distances because there is no query point. `"embeddings"` is deliberately excluded from both defaults: returning full float vectors for every hit can dwarf the rest of the payload, so you opt in only when you actually need the numbers. Ids are always present in the response and are not something you request. Trimming `include` down to what the caller uses — often just `["documents", "distances"]` for a RAG prompt — is a cheap win over an HTTP client, where every unused field is serialised and shipped. The `distances` values are distances, so lower is closer, and their scale depends on the distance space the collection was created with.

code

python · 12 lines
python
# Trim the payload to exactly what the caller consumes
res = collection.query(
    query_texts=["retry budget"],
    n_results=5,
    include=["documents", "distances"],
)

for doc, dist in zip(res["documents"][0], res["distances"][0]):
    print(round(dist, 4), doc[:60])

# Vectors are opt-in, and get() returns flat lists, not nested ones
vecs = collection.get(ids=["doc-1"], include=["embeddings"])["embeddings"]

go deeper

for a junior

Know that include picks which fields come back — documents, metadatas, distances — and that embeddings are left out unless you ask for them. Ids are always there.

for a middle

Be able to state the two defaults (query adds distances, get does not) and explain that include changes the payload, never the result set, and that lower distance means closer.

for a senior

Show the operational angle: over an HTTP client, unrequested fields are serialised and shipped, so trimming include on a hot retrieval path is a real latency and bandwidth win.

for a principal

Own the point that a distance is not a calibrated relevance score. Any thresholding policy has to be derived empirically per collection and per embedding model, and revisited whenever either changes.

## What include is for A Chroma result row can carry several parallel pieces of data: the id, the document text, the metadata dictionary, the distance to the query point, the raw embedding, and (for multimodal collections) uris and data. `include` chooses which of these travel back. It is orthogonal to `where` and `where_document`: those decide *which records* qualify, `include` decides *what each record brings with it*. Setting `include=["distances"]` does not remove records — it just leaves the documents out of the payload. ## The defaults, and the one deliberate omission `query()` returns documents, metadatas and distances by default. `get()` returns documents and metadatas; there is no distance because `get` has no query vector to measure against. Embeddings are in neither default, and that is a considered choice. A collection using a 1536-dimensional model returns roughly 1536 floats per hit. Ask for twenty hits and the vector payload is far larger than every document and metadata field combined. Most callers — a RAG prompt builder, a UI result list — never look at the vectors, so shipping them is pure waste. You request `"embeddings"` when you genuinely need the numbers: computing your own re-ranking, clustering the neighbourhood, debugging a suspected model mismatch, or migrating vectors elsewhere. Ids are a special case: they are always in the response and are not something you put in `include`. Every result row is addressable, always. ## Reading distances The field is named `distances`, not `scores`, and the naming is the whole point: **smaller means closer**. Sorting ascending gives you best-first, which is already the order Chroma returns. Two follow-on facts matter in practice. First, the numeric scale depends on the distance space chosen when the collection was created — the default space is a squared-L2 distance, and cosine-configured collections produce a different range. So a raw threshold like "drop anything above 0.8" is only meaningful once you know the space, and it does not transfer between collections configured differently. Second, distances are not calibrated probabilities. Converting them into a confidence percentage for a UI is a fabrication unless you have empirically mapped distances to relevance on your own data. The defensible use of the number is *relative*: comparing hits within one result set, or watching the distribution shift over time as a regression signal. ## The response shape `query()` returns a dictionary whose values are nested one level per query — `res["documents"][0]` is the document list for the first query, `res["distances"][0][2]` the distance of the third hit for the first query. Fields you did not include are still present as keys but hold no data, so code that blindly indexes `res["embeddings"][0][0]` after trimming `include` fails at read time rather than at call time. The response also reports which fields were included, so defensive code can check rather than assume. `get()` returns flat lists instead — one level, because there is no batch of queries. Code that copies an indexing pattern from `query` to `get` and keeps the `[0]` is a common and confusing bug: it silently takes the first element instead of the whole list. ## When trimming actually matters Against an in-process `PersistentClient`, trimming `include` saves object construction and little else. Against an HTTP server, every unused field is serialised, pushed over the wire and parsed by the client — and that is where a 20-hit query returning unnecessary embeddings shows up as latency and bandwidth on a hot path. The habit worth having is to name exactly the fields the caller consumes: a retrieval step feeding an LLM prompt usually needs `["documents"]` plus `["distances"]` if it thresholds, and nothing more. The inverse habit is worth naming too: do not reach for `"embeddings"` to "check that things are working". `collection.count()` and a small `peek()` answer that question far more cheaply.

  • Why are embeddings excluded from the default include list?
    Size. A hit from a 1536-dimensional collection carries about 1536 floats, so twenty hits ship far more vector data than document text — and almost no caller reads it. Chroma makes it opt-in so the common path stays cheap. You request it when you need the numbers themselves: custom re-ranking, clustering, debugging a suspected model mismatch, or exporting vectors.
  • How does the response shape differ between query() and get()?
    `query()` nests one level per query because query_texts and query_embeddings are batched: `res["documents"][0]` is the first query's list. `get()` returns flat lists — there is no batch dimension. Copying a `[0]` index from query code into get code silently takes the first record instead of all of them, which is a quiet and very common bug.
  • Can you turn a distance into a confidence score for the UI?
    Not honestly, without calibration. Distances depend on the collection's configured distance space and on the embedding model's geometry; they are not probabilities. Use them relatively — comparing hits inside one result set, or monitoring the distribution for drift. If you need a threshold, derive it empirically from labelled examples on your own data and re-derive it whenever the model changes.

saying these in an interview costs you the question

  • Thinking include filters which records are returned
  • Expecting embeddings in the response by default
  • Adding 'ids' to include to get identifiers back
  • Treating distances as similarity scores where higher is better
  • Reusing a distance threshold across differently configured collections

context