skip to content

Why must a Chroma collection keep the same embedding_function for its whole life?

level: middleimportance: must knowfreq 68%

answer

  1. One model, one coordinate space
  2. Chroma only validates the dimension count
  3. Same width, different model — silent failure
  4. Loud error is the lucky case
  5. Fix is re-embedding, not repair

basics

~20 s

Distances are only meaningful between vectors from the same model. Mixing embedding functions in one Chroma collection puts incomparable vectors in one index, so nearest-neighbour results become arbitrary — and if the dimensions happen to match, nothing errors.

solid answer

~50 s

A collection's embedding function decides the coordinate space every stored vector lives in. Two models — even two versions of the same model — place the same sentence in completely different positions, so a distance computed between a vector from model A and a query embedded by model B is noise. Chroma cannot detect this: its only structural check is dimensionality. Swap `all-MiniLM-L6-v2` (384 dims) for a 1536-dimension model and you get a loud dimension-mismatch error on the next write, which is the *lucky* case. Swap it for a different 384-dimension model and everything keeps working while retrieval quality quietly collapses. Current Chroma persists the embedding-function configuration with the collection so a later handle rebuilds the same function instead of falling back to the default, but that protects the accidental case, not a deliberate model change: switching models means re-embedding the corpus into a fresh collection.

code

python · 14 lines
python
import chromadb
from chromadb.utils import embedding_functions

# Define once, import everywhere: ingestion, query service, backfill.
EMBEDDER = embedding_functions.SentenceTransformerEmbeddingFunction(
    model_name="all-MiniLM-L6-v2"
)

client = chromadb.PersistentClient(path="./chroma")
collection = client.get_or_create_collection(
    name="docs",
    embedding_function=EMBEDDER,
    metadata={"embedding_model": "all-MiniLM-L6-v2"},
)

go deeper

for a junior

Remember that a collection is tied to one embedding function and that mixing models makes search results wrong. Say plainly that changing models means re-embedding.

for a middle

Explain that the vector space is model-specific, that Chroma's only check is dimensionality, and that two same-width models therefore fail silently while differing widths fail loudly.

for a senior

Demonstrate the operational discipline: one shared embedding-function definition, the model id recorded in collection metadata, and a golden query set run after ingestion to catch a mixed collection before users do.

for a principal

Own the position that the embedding model is a versioned interface with the same weight as a schema — pinned, recorded, and changed only through a rebuild-and-cutover process, never through an edit to a config value.

## Vectors are only comparable within one model An embedding model is a learned map from text into a space of a few hundred to a few thousand dimensions. The space has no external meaning: the axes are whatever training produced, and two independently trained models put the same sentence in unrelated places. Similarity search works because *all* the vectors being compared came out of the same map, so proximity in that space corresponds to similarity in meaning. The moment a collection contains vectors from two different maps, that property is gone. A query embedded by model B is compared against records embedded by model A, and the ranking that comes back is not "less accurate" — it is meaningless, driven by incidental geometry rather than semantics. This is why a Chroma collection's embedding function is effectively part of its identity, not a runtime option. ## Why Chroma cannot catch it for you Chroma validates one structural property: dimensionality. The first write fixes the collection's vector width, and later writes that disagree are rejected with a dimension-mismatch error. That check is coarse. - **Dimensions differ** (384 vs 1536): you get a hard failure on the next write or query. This is the good outcome — the system stops and tells you. - **Dimensions match** (two different 384-dimension models, or two revisions of one model): every call succeeds. The index accepts the vectors, queries return results, distances come back as plausible-looking numbers. Nothing in the API surface hints that half your corpus is unreachable. The only symptom is that retrieval quality degrades, which you will discover through user complaints or an evaluation set, not a stack trace. The second case is the one that reaches production, and it is why "the collection has one embedding function, forever" is treated as a rule rather than a guideline. ## How the function gets attached You pass it at creation: ``` get_or_create_collection(name="docs", embedding_function=ef) ``` If you pass nothing, Chroma attaches its default — an ONNX build of `all-MiniLM-L6-v2` at 384 dimensions. That default is a genuine trap in older code: a script that created the collection with `OpenAIEmbeddingFunction` and a later script that fetched it without naming a function used to end up with two different functions on one collection, with no error if the dimensions lined up. Current Chroma closes that hole by persisting the embedding-function configuration as part of the collection, so obtaining the collection later reconstructs the same function rather than silently defaulting. For built-in functions this works out of the box. A **custom** embedding function has to be serializable for this to hold — Chroma's `EmbeddingFunction` interface exposes a `name()`, a `get_config()` and a `build_from_config()` for exactly this purpose. A hand-rolled class that only implements `__call__(self, input)` still works when you pass the instance explicitly, but it is not something Chroma can rebuild from stored configuration, so you must keep passing it. Either way, the safe habit is to construct the embedding function in exactly one place in your codebase and have every path — ingestion, query service, backfill job, notebook — import that one object. The failure mode is not a bad line of code; it is two files that each look correct in isolation. ## Same model name is not the same model Even pinning the model *name* is not always enough. A hosted embedding endpoint can be revised behind a stable name, and a locally cached model can differ between a laptop and a container image. Recording the model identifier (and, where you have it, a version or revision) in the collection's metadata costs nothing and turns "which model produced these vectors?" from archaeology into a lookup. ## Detecting a mixed collection There is no built-in audit. The practical detections are behavioural: - Keep a small golden set of queries with known-good expected documents and run it after any ingestion or deployment. A mixed collection tanks it immediately. - Watch the distance distribution of top results. Vectors from a foreign model tend to sit at distances clustered differently from native ones, so a sudden shift in the score profile is a signal. - Compare `count()` against the number of documents you believe you embedded with the current function. ## The correct response Once a collection is mixed, there is no in-place repair worth attempting: you cannot tell from the stored data which vector came from which model. The fix is to re-embed from your source of truth into a fresh collection. This is precisely why the original text should live somewhere you can replay — an object store, a database, a document pipeline — rather than only inside Chroma.

  • Which is worse in practice: a dimension mismatch or two same-width models on one collection?
    Two same-width models, without question. A dimension mismatch is rejected on the next write, so the problem surfaces in seconds with an unambiguous error. Two 384-dimension models coexist happily: every call succeeds, distances look normal, and the only symptom is degraded retrieval quality that shows up as vague user complaints weeks later.
  • How would you detect that a collection has been silently mixed?
    Behaviourally, since there is no stored provenance to inspect. Keep a golden set of queries with known-good expected results and run it after every ingestion or deploy — a mixed collection fails it sharply. Watching the distance distribution of top hits also helps, since foreign-model vectors sit at a noticeably different score profile.
  • If a custom embedding function has to be re-created every time you obtain the collection, what does that imply for your code layout?
    Construct it in exactly one module and import that object everywhere — ingestion, query service, backfill jobs, notebooks. Chroma can rebuild built-in functions from stored configuration, but a hand-rolled class that only implements `__call__` cannot be reconstructed, so every code path that touches the collection has to supply the identical instance.
  • Why is recording the model identifier in collection metadata worth the effort?
    Because vectors carry no provenance. Six months later, deciding whether a collection needs re-embedding means knowing which model produced it, and a hosted endpoint can be revised behind an unchanged name. A model id and revision in the collection metadata turns that question into a lookup instead of an inference from behaviour.

Two embedding models are two different cities that both have a Main Street. Giving a courier an address from one city's map and a destination from the other's produces a confident delivery to the wrong place.

saying these in an interview costs you the question

  • Thinks vectors from different models are comparable if dimensions match
  • Believes Chroma validates that the embedding model is unchanged
  • Says you can swap the embedding function and re-index in place
  • Treats a dimension-mismatch error as the main risk
  • Creates the embedding function separately in ingestion and query code

context