In Chroma, when do you pass query_embeddings instead of query_texts to collection.query()?
answer
- Two mutually exclusive query inputs
- One embeds for you, one does not
- Collection's embedding function versus your vector
- Dimensionality must match the collection
- Batched input means nested result arrays
basics
~20 squery_texts hands raw strings to Chroma, which embeds them with the collection's embedding function. query_embeddings takes vectors you already computed, so Chroma skips embedding entirely. Use it when the vector comes from elsewhere or is cached.
solid answer
~50 s`collection.query()` accepts either `query_texts` or `query_embeddings` — one or the other, not both. With `query_texts=["how do I tune GC"]`, Chroma calls the collection's embedding function on each string and searches with the result; that is the convenient path and it guarantees the query is embedded by the same model that embedded the stored documents. With `query_embeddings=[[0.1, 0.2, ...]]`, you supply vectors directly, which is what you want when the embedding was produced outside the process (a hosted model, an image or multimodal encoder), when you cache query vectors to avoid paying for repeated embedding calls, or when you are searching a collection whose vectors you inserted precomputed. The vector's dimensionality must match the collection's, otherwise Chroma rejects the query. Both parameters take a list, so one call can run a batch of queries; the response arrays are nested one level per query.
code
python · 17 linesimport chromadb
client = chromadb.EphemeralClient()
collection = client.get_or_create_collection("docs")
collection.add(
ids=["a", "b"],
documents=["tuning the garbage collector", "configuring the load balancer"],
)
# Chroma embeds the string with the collection's embedding function
by_text = collection.query(query_texts=["how do I tune GC"], n_results=2)
# You supply the vector; Chroma embeds nothing
vec = collection.get(ids=["a"], include=["embeddings"])["embeddings"][0]
by_vector = collection.query(query_embeddings=[vec], n_results=2)
print(by_text["ids"][0], by_vector["ids"][0])go deeper
Know that query_texts lets Chroma embed the string for you and query_embeddings takes a vector you already have, and that n_results caps how many neighbours come back.
Be ready to explain that the embedding function attached to the collection is what runs on query_texts, and that both parameters are batched, so results come back nested one list per query.
Show that you think about where embedding cost lands: query_texts pays a model forward pass on every request, and caching query vectors can be the cheapest latency win in a hot retrieval path.
Own the rule that one embedding model and version governs both ingest and query. Choosing the precomputed-vector path shifts that guarantee from the library to your own pipeline, and that contract needs to be written down.
## The two entry points A Chroma similarity search is one method, `collection.query()`, with two mutually exclusive ways to express the query point: - `query_texts=["..."]` — a list of raw strings. Chroma runs them through the embedding function that was attached to the collection when it was created or fetched, and searches with the resulting vectors. - `query_embeddings=[[...]]` — a list of numeric vectors that you computed yourself. Chroma embeds nothing and searches with exactly what you gave it. Supply one of them. Supplying both is an error, and supplying neither leaves Chroma with nothing to search for. ## Why query_texts is the default habit The great convenience of Chroma is that a collection carries its own embedding function, so "search by meaning" is a one-liner with no model plumbing at the call site. It also removes the single most common correctness bug in vector search: querying with a vector produced by a *different* model than the one that embedded the corpus. Cosine or L2 distance between vectors from two different embedding spaces is arithmetically valid and semantically meaningless, so the query returns plausible-looking nonsense with no error anywhere. When both sides go through the collection's embedding function, that mismatch cannot happen. ## Why you would reach for query_embeddings Several real situations make the text path the wrong one: - **The embedding is produced outside this process.** A separate service, a batch job, or another language runtime already computed the vector; re-embedding locally would need the same model loaded twice. - **The query is not text.** Image, audio, or multimodal search means the query vector comes from an encoder that does not take a string at all. The stored vectors were inserted precomputed; the query must be too. - **Caching and cost.** Hosted embedding APIs cost money and add network latency on every query. If a small set of queries repeats (canned prompts, autocomplete, evaluation suites), embedding once and reusing the vector removes that latency from the hot path. - **Experimentation.** Perturbing a vector, averaging several query vectors, or feeding a centroid computed from example documents are all things you can only do when you hold the numbers. ## Shape of the call and the response Both parameters are lists, which means one round trip can carry a batch of queries. `n_results` caps how many nearest records each query returns; it defaults to 10. Because the input is a batch, the response arrays are nested one level per query: `res["ids"][0]` is the id list for the first query, `res["distances"][1]` the distances for the second. A frequent beginner bug is indexing `res["documents"]` directly and being surprised by a list of lists. If the collection holds fewer records than `n_results`, you simply get fewer results — that is not an error condition. ## Failure modes **Dimension mismatch.** A hand-supplied `query_embeddings` whose length differs from the collection's vector width is rejected. This is the loud, easy failure. **Model mismatch.** The silent one. Nothing checks that your externally computed vector came from the same model as the stored vectors — dimensionality can coincide while the spaces differ completely. If you take the `query_embeddings` route, own the discipline of pinning the same model and version on both the write and read paths. **Assuming query_texts is free.** Every call with `query_texts` runs an embedding forward pass, either locally or against a remote API. At high query rates that is often the dominant term in latency, not the nearest-neighbour search itself. Measuring the split is what tells you whether caching query vectors is worth the complexity. ## Practical guidance Start with `query_texts` — it is fewer moving parts and it forecloses the model-mismatch bug. Move to `query_embeddings` deliberately, when a measurement or an architectural constraint demands it, and when you do, write down which model and version produced the vectors on both sides.
- What is the shape of the dictionary that query() returns for a batch of two queries?Each key holds one inner list per query, in input order. With two `query_texts`, `res["ids"]` is a list of two id lists, and `res["distances"][1]` holds the distances for the second query. Flattening this without indexing the query dimension is a classic first-day bug, and it hides itself when you only ever send one query.
- What happens if n_results is larger than the number of records in the collection?You get every record that qualifies and no error — the result lists are simply shorter than `n_results`. It is a cap, not a demand. Older clients logged a notice about requesting more results than the index holds; treat the short list as normal, and check `collection.count()` if you expected more data to be there.
- Why is querying with a vector from a different model than the stored one dangerous?Distance is still computable, so nothing errors — you just get meaningless neighbours. Two embedding models can share a dimensionality while placing concepts in completely unrelated geometry. The bug surfaces as "retrieval quality is bad" rather than as an exception, which is why pinning one model and version across ingest and query is a hard rule.
saying these in an interview costs you the question
- Passing both query_texts and query_embeddings in one call
- Assuming query_texts skips an embedding model call
- Reading res['documents'][0] as a single document string
- Believing matching dimensionality means matching embedding model
- Thinking a short result list means the query failed