skip to content

How do you store vectors from your own embedding pipeline in a Chroma collection?

level: seniorimportance: should knowfreq 44%

answer

  1. Hand Chroma the numbers directly
  2. Embedding function is skipped on that write
  3. First write pins the width
  4. Text search still needs a matching function
  5. Watch the accidental default model download

basics

~10 s

Pass embeddings= to add() alongside the ids; the collection's embedding function is then never called. The first write fixes the collection's dimensionality, and any later batch of a different width is rejected.

solid answer

~50 s

Supply the vectors directly: `add(ids=[...], embeddings=[[...], ...])`, optionally with `documents=` so the original text rides along as payload and `metadatas=` for filtering. When `embeddings` is present Chroma stores exactly those numbers and skips the embedding function entirely. Two things then become your responsibility. First, dimensionality — the first write pins the collection's width and later mismatches are rejected, but nothing verifies the vectors all came from the same model. Second, the read path: a collection whose text side is never used still needs a matching embedding function if anyone is going to search it by text, otherwise every search has to arrive as a precomputed query vector. Worth knowing that if you name no embedding function at all, Chroma attaches its default, which will fetch an ONNX MiniLM model on first text use — surprising in a container that has no business downloading models.

code

python · 19 lines
python
import chromadb

client = chromadb.PersistentClient(path="./chroma")

# Explicitly no embedding function: this collection is a pure index.
collection = client.get_or_create_collection(
    name="vectors_only",
    embedding_function=None,
    metadata={"hnsw:space": "cosine", "embedding_model": "internal-encoder-v3"},
)

chunk = {
    "ids": ["a1", "a2"],
    "embeddings": [[0.1] * 768, [0.2] * 768],
    "documents": ["first chunk text", "second chunk text"],
    "metadatas": [{"source": "manual"}, {"source": "manual"}],
}
collection.add(**chunk)
assert collection.count() == 2

go deeper

for a junior

Know that passing embeddings= to add() stores your vectors as-is and skips the collection's embedding function, and that ids are still required.

for a middle

Explain that the first write fixes dimensionality, that documents passed alongside vectors are unchecked payload, and that text search still requires an embedding function matching the model that produced the stored vectors.

for a senior

Show the operational shape: batched writes with deterministic ids, count() reconciliation, an explicit decision about whether an embedding function is attached, and awareness that omitting one silently attaches the default.

for a principal

Own where embedding runs in the architecture — a single upstream service versus per-consumer inference — and treat the vector width and model id as versioned contract decisions fixed before the first byte is written.

## Why you would do this at all Chroma's convenience feature is that a collection can embed text for you. Production pipelines often want the opposite arrangement, for reasons that have nothing to do with Chroma: - **Embedding is the expensive step.** It belongs on hardware you chose — a GPU worker, a batch job, an autoscaled service — not inside whatever process happens to be writing to the store. - **You want one embedding path.** If a Spark job, a streaming consumer and an API server all embed, having each construct its own function is three chances to drift. One upstream service that emits vectors is one path. - **The vectors already exist.** A migration, a re-embedding backfill, or a model that lives behind an internal API all produce vectors before Chroma ever sees the text. - **Reproducibility.** Vectors computed once and stored can be replayed into a new index without re-running inference. ## The write ``` collection.add( ids=[...], embeddings=[[...], [...]], documents=[...], # optional payload metadatas=[...], # optional, for filtering ) ``` When `embeddings` is present it wins: the collection's embedding function is not invoked for that call. `documents` is then purely payload — text stored so results can return it — and carries no obligation to correspond to the vectors, which is a footgun worth naming out loud: if your pipeline chunks text differently from what you store as the document, results will show text that does not match what was actually embedded. Keep the chunk you embedded and the document you store identical unless you have a deliberate reason otherwise. All lists are positional. Vector *i* belongs to id *i*, document *i* and metadata *i*, and a misaligned zip is a silent corruption. Build the four lists in one pass. ## Dimensionality is the only guardrail The first successful write fixes the collection's width. Every subsequent batch must match or it is rejected with a dimension-mismatch error. That is the entire structural contract — Chroma has no notion of which model produced a vector, so two 768-dimension models write into the same collection without complaint and quietly destroy the ranking. Recording the model id in the collection metadata at creation time is the cheap mitigation. A related consequence: if you truncate or reduce vector width (some models allow shortened outputs), pick the width before the first write. There is no reshaping an existing collection. ## What still needs an embedding function Supplying your own vectors does not automatically mean the collection has no embedding function. The choices are: - **Attach the matching function anyway.** Best when you also want text search against the collection. The write path uses your precomputed vectors; the read path can accept raw query text and embed it with the same model. Consistency is on you: the function you attach must be the model your pipeline used. - **Attach none deliberately.** Then every search must arrive as a query vector your pipeline produced. This is the cleanest arrangement when Chroma is a pure index behind your own retrieval service, and it removes any chance of a mismatched read path. - **Attach nothing by accident.** This is the case to avoid. Say nothing and Chroma attaches its default MiniLM function, which fetches the ONNX model on first use. In a locked-down container that is a startup failure at an odd moment, and if the widths happen to line up it is worse: text queries get embedded by a model that never touched your corpus. Be explicit either way; "I didn't pass one" is not a decision. ## Metric still matters Bring-your-own vectors do not change the index configuration story. The collection's distance space is chosen at creation and should match what your model was trained for — cosine for the common sentence-embedding families. You do not need to pre-normalise for the cosine space; it accounts for magnitude itself. If you *have* normalised upstream, inner product is a legitimate alternative that ranks identically for unit vectors. ## Operational shape of a bulk load A backfill of any size is a batched job, and the batching now sits on your side of the wall: - Chunk writes into a few hundred to a few thousand records so a failure costs one chunk, not the run. - Make ids deterministic — a content or source-key hash — so a retried chunk is a no-op rather than a duplicate storm, remembering that `add()` on an existing id warns and keeps the old record rather than refreshing it. - Check `count()` against the number of records your pipeline emitted after each phase. A quiet gap is usually duplicate ids being skipped. - Sample with `peek()` and confirm that a stored document really is the text whose vector sits beside it. ## When the extra machinery is not worth it For a prototype, a local RAG demo, or a corpus of a few thousand documents, letting the collection embed is simply less code and fewer ways to be wrong. Bring-your-own vectors earns its keep when embedding is expensive, shared across consumers, or already happening somewhere else — not as a default posture.

  • If you supply your own vectors, is there any reason to still attach an embedding function?
    Yes — if anyone will search the collection with raw text. The write path uses your vectors while the read path embeds the incoming query, so the attached function must be the same model your pipeline used. If Chroma sits purely behind your own retrieval service that produces query vectors, attaching none is cleaner and removes the mismatch risk entirely.
  • You pass documents alongside precomputed embeddings. What relationship does Chroma enforce between them?
    None. The document is stored as payload and is never checked against the vector, so a pipeline that chunks text one way and stores a different string will happily return text that does not correspond to what was embedded. Keep the embedded chunk and the stored document identical unless you have a deliberate reason not to.
  • What goes wrong if you simply omit embedding_function on a bring-your-own-vectors collection?
    Chroma attaches its default MiniLM function. In a network-restricted container that surfaces as a model-download failure the first time text is embedded. Worse, if your vectors happen to be 384 dimensions wide, a text query gets embedded by a model that never saw your corpus and returns confident nonsense. Decide explicitly rather than by omission.
  • Do you need to normalise vectors before writing them if the collection uses cosine space?
    No — the cosine space accounts for magnitude when computing distance, so raw model output is fine. Normalising upstream is still a reasonable choice, and if you do it consistently you can use inner product instead, which ranks identically for unit-length vectors while skipping redundant norm computation.

saying these in an interview costs you the question

  • Thinks Chroma re-embeds or validates vectors you supply
  • Assumes documents and embeddings are checked for correspondence
  • Believes the collection's width can be changed after the first write
  • Omits embedding_function and is surprised by a model download
  • Retries a failed bulk chunk expecting add() to overwrite duplicates

context