skip to content

How would you switch embedding models for a Chroma collection already serving production traffic?

level: principalimportance: should knowfreq 36%

answer

  1. Vectors cannot be translated between models
  2. Rebuild in parallel, never in place
  3. Source text must live outside the store
  4. Measure quality before you cut over
  5. Version the collection name, flip by config

basics

~20 s

Treat it as a rebuild, not an edit. Re-embed the corpus from its source of truth into a new collection with the new model, evaluate retrieval quality against the old one on a fixed query set, then cut the application over by name and delete the old collection.

solid answer

~50 s

There is no in-place path: a collection's vectors are only comparable within one model, and if the new model has a different width, the collection physically cannot hold them. So the plan is a parallel build. Create a second collection — a versioned name such as `docs_v2` — configured with the new embedding function and the metric that model expects, and backfill it from the *source text*, which is why the corpus must live somewhere replayable rather than only inside Chroma. Budget the backfill honestly: it is N documents of inference, with real cost and hours of wall time for a hosted model. Before cutover, run a fixed evaluation set against both collections and compare, because a newer model is not automatically better for your domain. Then flip the collection name in configuration, keep the old one for a rollback window, and delete it once you are confident. Record the model id in the collection metadata so the next person can tell what produced which vectors.

go deeper

for a junior

Know that changing the embedding model means every vector must be recomputed, and that this is done by building a new collection rather than editing the existing one.

for a middle

Explain why in-place swapping produces a mixed, meaningless index, why a width change is rejected outright, and how a versioned collection name plus a config-driven cutover keeps the application simple.

for a senior

Show the operational plan: resumable batched backfill with deterministic ids, count reconciliation, handling of writes arriving during the window, a rollback copy, and disk headroom for two collections at once.

for a principal

Own the framing that the embedding model is a pinned interface with a recurring migration cost, insist the go/no-go rests on measured retrieval quality against a fixed evaluation set, and weigh a wider model's permanent memory cost against that measured gain.

## Why this is a rebuild The question is not really about Chroma's API — it is about the fact that stored vectors carry no way back to text. A vector is the output of one specific model; you cannot translate it into another model's space, so "switch the model" always means "recompute every vector". Chroma makes this concrete in two ways: the collection's dimensionality is fixed by its first write, and nothing in the store records which model produced which vector, so a partial swap yields a collection whose distances are meaningless with no error to tell you. That rules out the tempting shortcuts — passing the new embedding function to the existing collection, or re-adding only changed documents. The first only affects how new text is embedded; the second leaves a mixed index. ## Prerequisite: the source of truth is not Chroma The rebuild is only possible if you can replay the corpus. If Chroma holds the sole copy of your documents you are dependent on reading them back out and hoping the stored text is exactly what was embedded — plausible, but fragile, and it fails outright for anything where the stored document differed from the embedded chunk. The durable arrangement is that the vector store is a *derived* index: object storage, a relational store, or the upstream document pipeline holds canonical text, and Chroma can be dropped and rebuilt at any time. Establishing that is worth doing before you need it, and a model migration is usually the moment teams discover they have not. ## The plan **1. Provision a new collection.** Versioned name (`docs_v2`), the new embedding function attached, and a distance metric matching what the new model was trained for — do not assume the old collection's metric carries over, since models differ in whether cosine or dot product is their native scoring. Put the model identifier in the collection metadata. **2. Size the backfill before starting it.** Multiply corpus size by per-document embedding cost and by throughput. For a hosted embedding API this is a real invoice and a real rate limit; for local inference it is GPU hours. Anything large is a resumable batch job with deterministic ids (content hashes work well), checkpointing, and per-chunk retry — not a script you run in a terminal and hope survives. **3. Backfill.** Chunk the writes, reconcile `count()` against what your job emitted at each phase, and sample records to confirm text and vector correspond. Deterministic ids make a retried chunk harmless. **4. Evaluate before you commit.** This is the step most often skipped and the one that justifies the whole exercise. Keep a fixed set of representative queries with known-good expected documents, and score both collections on it — recall at k, or whatever your product actually cares about. A larger or newer model is not automatically better on your domain, and the dimension increase you are paying for in storage and latency has to show up as measured quality. If it does not, the honest outcome is to abandon the migration, and having run the comparison is what makes that decision defensible. **5. Cut over.** The application should already read its collection name from configuration, so the flip is a config change and a restart, not a code deploy. A shadow-read period — serving from the old collection while querying the new one and logging the differences — buys real confidence when the traffic justifies it. **6. Keep the old collection for a rollback window,** then `delete_collection` it. Two full copies of the corpus is the cost of a safe migration; on a single-node Chroma deployment, check that the disk can hold both before you start, because the new model may also be wider than the old. ## The tradeoffs a lead actually owns - **Cost versus measured gain.** Re-embedding is the dominant expense and it recurs every time you change models. That argues for changing models rarely and deliberately, and for pinning the model version rather than tracking a moving alias. - **Dimensionality has a running cost.** A wider model raises memory and index size for as long as it is deployed, not just during the migration. That is an ongoing bill against a one-off quality measurement. - **The embedding model is an interface.** Treat it like a schema version: pinned, recorded in collection metadata, changed only through this process. Teams that treat it as a tunable config value end up with mixed collections and no way to prove it. - **Freshness during the window.** If documents are still arriving while you backfill, either dual-write new documents into both collections or replay everything written after your backfill's start watermark before cutting over. Deciding which *before* you start avoids a gap nobody notices until a recent document is unfindable. ## What good looks like afterwards A collection whose metadata names the model that built it, an application that reads the collection name from config, a source of truth outside the vector store, and an evaluation set that can be re-run on demand. Those four things turn the next model change from a project into a routine.

  • Why can't you re-embed in place by passing the new embedding function to the existing collection?
    Because that only changes how *new* text is embedded. Every stored vector still comes from the old model, and the two spaces are not comparable, so the collection becomes mixed and rankings become noise. If the new model is a different width, writes are rejected outright — which is actually the safer failure, since the same-width case fails silently.
  • Documents keep arriving during a multi-hour backfill. How do you avoid a gap?
    Decide upfront between two approaches: dual-write every new document into both collections during the window, or record a watermark at backfill start and replay everything written after it before cutting over. Either works; choosing neither leaves recently-added documents missing from the new collection with no error to reveal it.
  • How do you justify not migrating after building the new collection?
    With the evaluation numbers. If recall at k on your fixed query set does not improve, the migration buys nothing while adding ongoing memory and index cost for a wider model, plus the re-embedding bill. Having run the comparison turns 'we decided against it' into a measured decision rather than inertia, and the new collection can simply be deleted.
  • What makes the next model migration cheaper than this one?
    Four things established during this one: canonical text living outside the vector store so any rebuild is replayable, the collection name read from configuration so cutover is a config flip, the model id recorded in collection metadata so provenance is a lookup, and a reusable evaluation set. With those in place the next change is a routine batch job rather than a project.

saying these in an interview costs you the question

  • Suggests changing the embedding function on the existing collection
  • Assumes a newer or larger model is automatically better
  • Plans no evaluation step before cutover
  • Forgets documents arriving during the backfill window
  • Assumes Chroma can supply the source text for re-embedding

context