skip to content

In Chroma, what happens to raw text passed to collection.add(documents=...)?

level: juniorimportance: must knowfreq 76%

answer

  1. Text goes in, vectors come out
  2. The collection owns the embedding step
  3. ids are mandatory, lists are positional
  4. Default is all-MiniLM-L6-v2, 384 dims
  5. Duplicate id warns, does not overwrite

basics

~20 s

Chroma runs each document through the collection's embedding function and stores the resulting vector next to the text, its id and its metadata. ids are required and must be unique; documents, metadatas and embeddings are parallel lists.

solid answer

~40 s

A Chroma collection bundles storage with an embedding function, so `add(ids=[...], documents=[...])` is enough: Chroma calls the collection's embedding function on the document strings, in your process, and stores the vector, the original text, the id and any metadata as one record. If you never chose a function, Chroma attaches its default — an ONNX build of `all-MiniLM-L6-v2` producing 384-dimensional vectors — which downloads the model file on first use. `ids` is the only always-required argument, and every list you pass must be the same length and aligned by position. `add()` is not an update: supplying an id that already exists makes Chroma warn and keep the existing record rather than overwrite it. If you already have vectors, pass `embeddings=` and the embedding function is not called at all.

code

python · 12 lines
python
import chromadb

client = chromadb.EphemeralClient()
collection = client.get_or_create_collection(name="docs")

collection.add(
    ids=["doc-1", "doc-2"],
    documents=["Chroma stores embeddings.", "Collections bundle text and vectors."],
    metadatas=[{"source": "guide", "page": 1}, {"source": "guide", "page": 2}],
)

print(collection.count())

go deeper

for a junior

Be able to say that a collection carries its own embedding function, that add() needs ids plus either documents or embeddings, and that the lists line up by position.

for a middle

Explain that embedding runs client-side in your process before the write, that the default is a 384-dimensional MiniLM ONNX model downloaded on first use, and that add() warns on a duplicate id instead of overwriting.

for a senior

Show that you plan ingestion as a batched, retryable job: chunk sizes, idempotent ids derived from content, count() checks after each run, and awareness that a hosted embedding function turns every add into a rate-limitable network call.

for a principal

Own the decision of what the id and metadata schema should be before a single document is written, since both are effectively immutable across a corpus, and decide whether Chroma holds the only copy of the text or is a rebuildable derived store.

## What a collection actually is In Chroma, a *collection* is not just a table of vectors. It is a named bundle of four things that travel together: the vectors, the original document text, arbitrary JSON-ish metadata per record, and an **embedding function** — the callable that turns a string into a vector. That last part is what makes Chroma feel lightweight: because the collection knows how to embed, you can hand it raw text and never touch a model yourself. You get a collection with `get_or_create_collection(name=...)`, optionally passing `embedding_function=` and a `metadata=` dict for index settings. ## The shape of add() `collection.add()` takes several parallel arguments: - `ids` — a list of strings. **Required on every call.** Chroma does not generate ids for you; if you want UUIDs or content hashes, you produce them. - `documents` — the raw text, one string per id. Optional if you supply `embeddings`. - `embeddings` — precomputed vectors, one list of floats per id. Optional if you supply `documents`. - `metadatas` — one dict per id, whose values are scalars (string, int, float, bool). Optional. All supplied lists must be the same length and are matched **by position**, not by any key inside them. A single off-by-one in how you build these lists silently attaches the wrong metadata to the wrong text — a bug no exception will catch for you, so build the lists in one loop rather than three. ## Where the embedding happens The embedding call runs **in your process, client-side**, before anything is written. With the default function, that means an ONNX Runtime session in your Python process; with a hosted function such as `OpenAIEmbeddingFunction`, it means an HTTP call per batch out to the provider. Two consequences follow. First, a big `add()` is not a cheap write. Its latency is dominated by embedding, not by storage, and a hosted function makes it a network operation that can rate-limit or fail halfway. Batch your adds into chunks of a few hundred to a few thousand documents rather than one giant call, and be prepared to retry a chunk. Second, the very first call with the default embedding function downloads the model artifact to a local cache. In a container or a CI job with no outbound network, that is where an otherwise-correct script fails — and the error is about a model download, not about Chroma. ## Precomputed embeddings If you pass `embeddings=`, Chroma stores exactly those vectors and never calls the embedding function for that write. You may still pass `documents=` alongside them; the text is then stored purely as payload so results can carry it back. This is the path you take when a separate pipeline (a batch job, another service, a GPU box) already produced the vectors. The first write fixes the collection's dimensionality. Every later write must match it, or Chroma rejects the batch with a dimension-mismatch error. That check is on the *number* of dimensions only — it cannot tell you that two 384-dimensional vectors came from two different models, which is the far more dangerous mistake. ## add() is not upsert Re-running an ingestion script over the same corpus does not refresh it. Chroma treats an id that is already present as a duplicate: it warns and leaves the stored record alone, so an edited document keeps its stale vector. If your ingestion is meant to be re-runnable, use the upsert path or delete the ids first; if ids are meant to be new every time, derive them from something genuinely unique. A content hash makes a good id precisely because it makes re-ingestion idempotent — an unchanged document produces the same id and is skipped, while an edited one produces a new id. ## Metadata rules of thumb Metadata values must be scalars. Nested dicts and lists are not stored, so flatten (`"tags": "a,b"` or one boolean column per tag) before writing. Because metadata is what later filtering keys off, decide the field names during ingestion — retrofitting a field means rewriting the records that lack it. ## What to check after ingesting `collection.count()` tells you how many records landed, which immediately catches a partially-failed batch or a run that silently skipped duplicates. `collection.peek()` returns a small sample of records so you can eyeball that documents, metadata and ids line up the way you intended.

  • Why is re-running an ingestion script over an edited corpus not enough to refresh Chroma?
    Because `add()` is not an update. Chroma sees an id that already exists, warns, and keeps the stored record — so the edited document keeps its old vector and old text. Either upsert, delete the ids first, or derive ids from a content hash so edited documents produce new ids and unchanged ones are genuinely no-ops.
  • You call add() with 500 documents and it takes twenty seconds. Where is the time going?
    Almost all of it is the embedding step, which runs client-side before the write. With the default ONNX model it is local CPU inference; with a hosted function it is HTTP round-trips to the provider. Storage is a small fraction. Batch into chunks, run embedding concurrently if the provider allows, and treat a failed chunk as retryable.
  • What kinds of values can go into metadatas, and what happens to nested structures?
    Only scalars — strings, ints, floats, booleans. Nested dicts and lists are not supported, so flatten before writing: join tags into a delimited string, or use one boolean field per tag. Pick the field names during ingestion, because adding a field later leaves every already-written record without it.

saying these in an interview costs you the question

  • Thinks Chroma generates ids when you omit them
  • Believes add() overwrites an existing id like an update
  • Assumes embedding happens server-side or lazily at query time
  • Passes metadatas as nested dicts or lists
  • Thinks documents and metadatas are matched by key, not position

context