skip to content

Chroma

The lightweight embedding database used for prototypes and local RAG: collections, an embedding function, and query-by-text with no infrastructure to run. It comes up as the starting point, and then as the migration question of what changes when you outgrow it.

on this pageshow

questions

17

In Chroma, what happens to raw text passed to collection.add(documents=...)?

level: juniorimportance: must knowfreq 76%

answer

  1. Text goes in, vectors come out
  2. The collection owns the embedding step
  3. ids are mandatory, lists are positional
  4. Default is all-MiniLM-L6-v2, 384 dims
  5. Duplicate id warns, does not overwrite

basics

~20 s

Chroma runs each document through the collection's embedding function and stores the resulting vector next to the text, its id and its metadata. ids are required and must be unique; documents, metadatas and embeddings are parallel lists.

solid answer

~40 s

A Chroma collection bundles storage with an embedding function, so `add(ids=[...], documents=[...])` is enough: Chroma calls the collection's embedding function on the document strings, in your process, and stores the vector, the original text, the id and any metadata as one record. If you never chose a function, Chroma attaches its default — an ONNX build of `all-MiniLM-L6-v2` producing 384-dimensional vectors — which downloads the model file on first use. `ids` is the only always-required argument, and every list you pass must be the same length and aligned by position. `add()` is not an update: supplying an id that already exists makes Chroma warn and keep the existing record rather than overwrite it. If you already have vectors, pass `embeddings=` and the embedding function is not called at all.

code

python · 12 lines
python
import chromadb

client = chromadb.EphemeralClient()
collection = client.get_or_create_collection(name="docs")

collection.add(
    ids=["doc-1", "doc-2"],
    documents=["Chroma stores embeddings.", "Collections bundle text and vectors."],
    metadatas=[{"source": "guide", "page": 1}, {"source": "guide", "page": 2}],
)

print(collection.count())

go deeper

for a junior

Be able to say that a collection carries its own embedding function, that add() needs ids plus either documents or embeddings, and that the lists line up by position.

for a middle

Explain that embedding runs client-side in your process before the write, that the default is a 384-dimensional MiniLM ONNX model downloaded on first use, and that add() warns on a duplicate id instead of overwriting.

for a senior

Show that you plan ingestion as a batched, retryable job: chunk sizes, idempotent ids derived from content, count() checks after each run, and awareness that a hosted embedding function turns every add into a rate-limitable network call.

for a principal

Own the decision of what the id and metadata schema should be before a single document is written, since both are effectively immutable across a corpus, and decide whether Chroma holds the only copy of the text or is a rebuildable derived store.

## What a collection actually is In Chroma, a *collection* is not just a table of vectors. It is a named bundle of four things that travel together: the vectors, the original document text, arbitrary JSON-ish metadata per record, and an **embedding function** — the callable that turns a string into a vector. That last part is what makes Chroma feel lightweight: because the collection knows how to embed, you can hand it raw text and never touch a model yourself. You get a collection with `get_or_create_collection(name=...)`, optionally passing `embedding_function=` and a `metadata=` dict for index settings. ## The shape of add() `collection.add()` takes several parallel arguments: - `ids` — a list of strings. **Required on every call.** Chroma does not generate ids for you; if you want UUIDs or content hashes, you produce them. - `documents` — the raw text, one string per id. Optional if you supply `embeddings`. - `embeddings` — precomputed vectors, one list of floats per id. Optional if you supply `documents`. - `metadatas` — one dict per id, whose values are scalars (string, int, float, bool). Optional. All supplied lists must be the same length and are matched **by position**, not by any key inside them. A single off-by-one in how you build these lists silently attaches the wrong metadata to the wrong text — a bug no exception will catch for you, so build the lists in one loop rather than three. ## Where the embedding happens The embedding call runs **in your process, client-side**, before anything is written. With the default function, that means an ONNX Runtime session in your Python process; with a hosted function such as `OpenAIEmbeddingFunction`, it means an HTTP call per batch out to the provider. Two consequences follow. First, a big `add()` is not a cheap write. Its latency is dominated by embedding, not by storage, and a hosted function makes it a network operation that can rate-limit or fail halfway. Batch your adds into chunks of a few hundred to a few thousand documents rather than one giant call, and be prepared to retry a chunk. Second, the very first call with the default embedding function downloads the model artifact to a local cache. In a container or a CI job with no outbound network, that is where an otherwise-correct script fails — and the error is about a model download, not about Chroma. ## Precomputed embeddings If you pass `embeddings=`, Chroma stores exactly those vectors and never calls the embedding function for that write. You may still pass `documents=` alongside them; the text is then stored purely as payload so results can carry it back. This is the path you take when a separate pipeline (a batch job, another service, a GPU box) already produced the vectors. The first write fixes the collection's dimensionality. Every later write must match it, or Chroma rejects the batch with a dimension-mismatch error. That check is on the *number* of dimensions only — it cannot tell you that two 384-dimensional vectors came from two different models, which is the far more dangerous mistake. ## add() is not upsert Re-running an ingestion script over the same corpus does not refresh it. Chroma treats an id that is already present as a duplicate: it warns and leaves the stored record alone, so an edited document keeps its stale vector. If your ingestion is meant to be re-runnable, use the upsert path or delete the ids first; if ids are meant to be new every time, derive them from something genuinely unique. A content hash makes a good id precisely because it makes re-ingestion idempotent — an unchanged document produces the same id and is skipped, while an edited one produces a new id. ## Metadata rules of thumb Metadata values must be scalars. Nested dicts and lists are not stored, so flatten (`"tags": "a,b"` or one boolean column per tag) before writing. Because metadata is what later filtering keys off, decide the field names during ingestion — retrofitting a field means rewriting the records that lack it. ## What to check after ingesting `collection.count()` tells you how many records landed, which immediately catches a partially-failed batch or a run that silently skipped duplicates. `collection.peek()` returns a small sample of records so you can eyeball that documents, metadata and ids line up the way you intended.

  • Why is re-running an ingestion script over an edited corpus not enough to refresh Chroma?
    Because `add()` is not an update. Chroma sees an id that already exists, warns, and keeps the stored record — so the edited document keeps its old vector and old text. Either upsert, delete the ids first, or derive ids from a content hash so edited documents produce new ids and unchanged ones are genuinely no-ops.
  • You call add() with 500 documents and it takes twenty seconds. Where is the time going?
    Almost all of it is the embedding step, which runs client-side before the write. With the default ONNX model it is local CPU inference; with a hosted function it is HTTP round-trips to the provider. Storage is a small fraction. Batch into chunks, run embedding concurrently if the provider allows, and treat a failed chunk as retryable.
  • What kinds of values can go into metadatas, and what happens to nested structures?
    Only scalars — strings, ints, floats, booleans. Nested dicts and lists are not supported, so flatten before writing: join tags into a delimited string, or use one boolean field per tag. Pick the field names during ingestion, because adding a field later leaves every already-written record without it.

saying these in an interview costs you the question

  • Thinks Chroma generates ids when you omit them
  • Believes add() overwrites an existing id like an update
  • Assumes embedding happens server-side or lazily at query time
  • Passes metadatas as nested dicts or lists
  • Thinks documents and metadatas are matched by key, not position

context

open as a page

In Chroma, what is the difference between EphemeralClient and PersistentClient?

level: juniorimportance: must knowfreq 80%

basics

~20 s

EphemeralClient keeps collections in memory only and loses everything when the process exits. PersistentClient(path="./chroma") writes to that directory and reloads it on the next run. Plain chromadb.Client() is ephemeral, which is why prototype data disappears.

open as a page

In Chroma, when do you pass query_embeddings instead of query_texts to collection.query()?

level: juniorimportance: must knowfreq 78%

basics

~20 s

query_texts hands raw strings to Chroma, which embeds them with the collection's embedding function. query_embeddings takes vectors you already computed, so Chroma skips embedding entirely. Use it when the vector comes from elsewhere or is cached.

open as a page

Why must a Chroma collection keep the same embedding_function for its whole life?

level: middleimportance: must knowfreq 68%

basics

~20 s

Distances are only meaningful between vectors from the same model. Mixing embedding functions in one Chroma collection puts incomparable vectors in one index, so nearest-neighbour results become arbitrary — and if the dimensions happen to match, nothing errors.

open as a page

In Chroma's query(), how do the where and where_document filters differ?

level: middleimportance: must knowfreq 72%

basics

~20 s

where filters on the structured metadata dictionary attached to each record, using operators like $eq, $in and $gte. where_document filters on the stored document text itself, with $contains and $not_contains substring matching. They are separate arguments and combine with AND.

open as a page

What breaks when two processes open the same Chroma PersistentClient path?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A persistent Chroma client is embedded, so each process gets its own SQLite handle and its own in-memory copy of the vector index. Writes by one process are invisible to the other's loaded index, and concurrent writers hit SQLite locking. Run the Chroma server and use HttpClient instead.

open as a page

When should you use Chroma's collection.get() instead of collection.query()?

level: juniorimportance: should knowfreq 50%

basics

~20 s

Use get() for exact retrieval: fetch by ids, or scan by where and where_document filters with limit and offset. It computes no embedding and returns no distances. Use query() when the question is similarity and you need ranked nearest neighbours.

open as a page

In Chroma, what does the hnsw:space collection metadata set, and what is its default?

level: middleimportance: should knowfreq 52%

basics

~10 s

It sets the distance metric the collection's index uses: "l2" (the default, squared Euclidean), "ip" (inner product) or "cosine". You pass it in the metadata dict at creation, and it cannot be changed afterwards.

open as a page

What does Chroma write inside a PersistentClient path directory on disk?

level: middleimportance: should knowfreq 52%

basics

~20 s

A chroma.sqlite3 file holds collections, ids, documents, metadata and the embedding records; alongside it sit UUID-named subdirectories, one per vector segment, containing the binary HNSW index files. Both parts belong to one database and must be copied together.

open as a page

How do you run Chroma as a server and connect clients to it?

level: middleimportance: should knowfreq 45%

basics

~20 s

Start the server with chroma run --path ./chroma_data --host 0.0.0.0 --port 8000, or run the chromadb/chroma container with its data directory on a mounted volume. Applications then use chromadb.HttpClient(host=..., port=...) and call heartbeat() to verify connectivity.

open as a page

What does Chroma's include parameter control in query(), and what comes back by default?

level: middleimportance: should knowfreq 55%

basics

~20 s

include selects which per-record fields the response carries: documents, metadatas, distances, embeddings, uris, data. Query defaults to documents, metadatas and distances; embeddings are never returned unless asked for. Ids always come back and need not be requested.

open as a page

In Chroma, how do collection.update() and collection.upsert() differ for unknown ids?

level: middleimportance: should knowfreq 45%

basics

~20 s

upsert() writes the record either way: it modifies an existing id or inserts a new one. update() only modifies records that already exist, so ids that are not there are not created. For idempotent re-ingest, upsert is the safe call.

open as a page

How do you store vectors from your own embedding pipeline in a Chroma collection?

level: seniorimportance: should knowfreq 44%

basics

~10 s

Pass embeddings= to add() alongside the ids; the collection's embedding function is then never called. The first write fixes the collection's dimensionality, and any later batch of a different width is rejected.

open as a page

How do you back up and move a Chroma persistent store between machines?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Chroma has no online backup command, so you quiesce writes or stop the process, copy the whole data directory atomically (SQLite file plus index directories), and restore it into the same or a newer Chroma version. Copying a live directory risks an inconsistent snapshot.

open as a page

In Chroma, what happens to results and latency when a where filter matches very few documents?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Chroma resolves the filter to a set of allowed ids and searches only within it, so you get up to n_results genuine matches rather than a short post-filtered list. The cost is latency: as the allowed set shrinks toward a tiny fraction of the collection, the search degrades toward scanning it.

open as a page

How would you switch embedding models for a Chroma collection already serving production traffic?

level: principalimportance: should knowfreq 36%

basics

~20 s

Treat it as a rebuild, not an edit. Re-embed the corpus from its source of truth into a new collection with the new model, evaluate retrieval quality against the old one on a fixed query set, then cut the application over by name and delete the old collection.

open as a page

When are Chroma's tenants and databases enough to isolate multiple customers?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Tenants and databases are namespacing, not isolation: they scope collection names inside one store but share the same process, memory, disk and index resources. They suit internal separation of environments or teams, not untrusted customers needing enforced boundaries or independent capacity.

open as a page