In Chroma, how do collection.update() and collection.upsert() differ for unknown ids?
answer
- One creates missing ids, one does not
- Deterministic ids make replay safe
- Only the fields you pass are written
- Passing documents means re-embedding
- Re-chunking leaves orphans behind
basics
~20 supsert() writes the record either way: it modifies an existing id or inserts a new one. update() only modifies records that already exist, so ids that are not there are not created. For idempotent re-ingest, upsert is the safe call.
solid answer
~50 sBoth take `ids` plus any of `documents`, `metadatas`, `embeddings`, and both only touch the fields you pass. The difference is what happens when an id is not present: `upsert()` inserts it, `update()` does not create it — the id simply is not written. That makes `upsert()` the right call for any re-ingest pipeline, because re-running it over a mixed batch of new and changed chunks converges to the correct state without you having to know which ids already exist, and without the duplicate-id problem you would get from a plain add. A subtlety on both: if you pass `documents` without `embeddings`, Chroma re-embeds the new text with the collection's embedding function, so a metadata-only correction should omit `documents` entirely to avoid paying for embeddings you did not need. Deletion is the mirror image — `delete(ids=[...])` for known ids, or `delete(where={...})` to remove everything matching a filter.
code
python · 9 lines# upsert writes whether or not the id exists
collection.upsert(
ids=["handbook-v3::0", "handbook-v3::1"],
documents=["retry budgets explained", "timeout defaults"],
metadatas=[{"doc_id": "handbook-v3"}, {"doc_id": "handbook-v3"}],
)
# Metadata-only fix: omit documents so nothing is re-embedded
collection.update(ids=["handbook-v3::0"], metadatas=[{"reviewed": True}])go deeper
Know that upsert() inserts an id that is not there yet while update() only modifies records that already exist, and that both write only the fields you pass.
Be able to explain that passing documents without embeddings triggers a re-embed with the collection's embedding function, so metadata-only fixes should omit the document field entirely.
Show the pipeline judgment: deterministic ids plus upsert make replay idempotent, and re-chunking needs a source-scoped delete first or old chunks linger as retrievable orphans.
Own the sync contract between the source of truth and the store — id derivation, the metadata field that scopes a document's records, and a counted, snapshotted procedure for any filter-based deletion.
## The three write paths Chroma has `add`, `update` and `upsert`, and the distinction is entirely about how each treats an id that already exists — or does not. - **`add`** is for records you believe are new. Re-adding an id that already exists is not a clean way to refresh it. - **`update`** modifies existing records only. Pass an id that is not in the collection and no record is created for it. - **`upsert`** writes either way: existing ids are modified, missing ids are inserted. All three accept `ids` together with `documents`, `metadatas`, and `embeddings`, as parallel lists of the same length. ## Why upsert is the default for pipelines Almost every real ingest is re-run: a source document changes, a chunking strategy is tweaked, a backfill replays a range. If your ids are **deterministic** — derived from source identity and chunk position, not random — then `upsert` makes the whole pipeline idempotent. Replay it and you converge to the intended state regardless of what was there before, without querying first to sort ids into new-versus-existing buckets, and without the race that check-then-write introduces when two workers process overlapping batches. That one design decision, deterministic ids plus upsert, removes most of the operational pain of keeping a vector store in sync with a source of truth. Random ids plus `add` produces silent duplicates: the same content stored twice, both copies retrieved, and a RAG prompt padded with redundant context that crowds out other material. ## Only the fields you pass are touched A point people miss: these calls are partial. `update(ids=["a"], metadatas=[{"reviewed": True}])` changes metadata and leaves the document and vector alone. That is exactly what you want for a re-tagging job. The corollary is the cost trap. If you pass `documents` **without** `embeddings`, Chroma must produce a vector for the new text, so it calls the collection's embedding function. Include the document field "just to be safe" in a metadata-only correction across 200,000 records and you have bought 200,000 embedding calls for nothing — with a hosted model, that is a real invoice and a long wall-clock time. Pass only what actually changed. The reverse case matters too: if the text changed, you **must** pass the new document (or a freshly computed embedding), or the stored vector still represents the old text and retrieval silently points at stale content. ## Deletion `delete` mirrors the read API. `collection.delete(ids=[...])` removes known records. `collection.delete(where={"source": "legacy"})` removes everything matching a metadata filter, and `where_document` works too. Filter-based deletion is the sharp edge in the whole API, because a filter that is broader than you think deletes correspondingly more, and there is no undo. The discipline is simple and worth stating in an interview: run the identical filter through `get(where=..., include=[])` first, look at the count, and only then convert it to a delete. On a collection that matters, take a copy of the store first — a local Chroma database is a directory, so a filesystem copy is a perfectly good pre-deletion snapshot. ## Re-ingest as a whole A source document that has been re-chunked usually produces a *different number* of chunks than before. Upserting the new chunks refreshes the ones whose ids repeat but leaves orphans behind — chunk 8 through 12 of the old version still sit in the collection, still retrievable, now representing text that no longer exists. The correct sequence is delete-then-write, scoped by a metadata field that identifies the source: `delete(where={"doc_id": "handbook-v3"})` followed by upserting the new chunks. Tagging every record with its source id at ingest time is what makes that possible, which is another reason to treat filterable metadata as a schema decision rather than an afterthought. ## The short version Use deterministic ids. Use `upsert` for writes so replay is safe. Pass only the fields that changed, and remember that passing `documents` means paying to re-embed. Delete by source-scoped filter before re-chunking, and always count the filter before you run it.
- You only need to fix a metadata field on 200,000 records. What must you avoid passing?`documents`. If you pass document text without embeddings, Chroma re-embeds it with the collection's embedding function — 200,000 needless embedding calls, which against a hosted model is real cost and a long run. Pass `ids` and `metadatas` only; the stored document and vector stay untouched. Update and upsert are both partial writes, so omitting a field leaves it alone.
- Why do deterministic ids matter so much for a re-ingest pipeline?They make upsert idempotent. Derive the id from source identity plus chunk position and replaying the pipeline converges to the right state whatever was there before — no check-then-write race between workers, no duplicates. Random ids force you to add, which stores the same content twice; both copies then get retrieved and crowd redundant context into the prompt.
- A source document is re-chunked from 12 chunks into 8. What breaks if you only upsert?Chunks 9 through 12 from the old version remain in the collection and stay retrievable, representing text that no longer exists. Upsert refreshes ids it writes and knows nothing about the ones it did not. The fix is to delete by a source-scoped filter first — `delete(where={"doc_id": ...})` — then write the new chunks, which requires that every record carries its source id in metadata.
- What is your safety procedure before running delete(where=...)?Run the identical filter through `get(where=..., include=[])` and check the count against what you expect; a filter that is broader than intended deletes correspondingly more and there is no undo. On anything that matters, snapshot first — a local Chroma store is a directory, so a filesystem copy is a perfectly adequate pre-deletion backup.
saying these in an interview costs you the question
- Expecting update() to insert an id that does not exist
- Using random ids and re-adding to refresh content
- Passing documents on a metadata-only correction
- Assuming upsert removes chunks that no longer exist
- Running delete(where=...) without counting the matches first