skip to content

How do you back up and move a Chroma persistent store between machines?

level: seniorimportance: should knowfreq 34%

answer

  1. no dump command, only a directory
  2. two halves, one atomic copy
  3. stop the writers first
  4. upgrades migrate the format one way

basics

~20 s

Chroma has no online backup command, so you quiesce writes or stop the process, copy the whole data directory atomically (SQLite file plus index directories), and restore it into the same or a newer Chroma version. Copying a live directory risks an inconsistent snapshot.

solid answer

~50 s

Backup in Chroma is a filesystem operation, because the store is a directory rather than a service with a dump API. The safe procedure is: stop the process — or at minimum stop all writes — then copy the entire directory, both `chroma.sqlite3` and every UUID index subdirectory, as one unit. Copying while a process is writing can capture the SQLite file and the index at different points, which restores as an inconsistent store with no error to warn you. On the target machine, restore into a directory of the same path shape and open it with a Chroma version at least as new as the one that wrote it; upgrades apply schema migrations on first startup and there is no supported downgrade, so keep a pre-upgrade copy as the rollback. Volume snapshots work if the snapshot is atomic and taken with writes quiesced. For anything critical, treat the pipeline that built the corpus as the real disaster-recovery plan and the directory copy as the fast path.

code

bash · 6 lines
bash
# 1. stop the owner of the directory (embedded app, or the chroma run server)
# 2. copy the whole store as one unit
tar -czf chroma-backup-$(date +%F).tar.gz -C /srv chroma_data
# 3. restore elsewhere and verify before trusting it
tar -xzf chroma-backup-2026-08-19.tar.gz -C /srv
python -c "import chromadb; c=chromadb.PersistentClient(path='/srv/chroma_data'); print([(x.name, x.count()) for x in c.list_collections()])"

go deeper

for a junior

Know that a Chroma store is a directory and that backing it up means copying the whole directory, not one file inside it, while nothing is writing.

for a middle

Explain why the copy has to be atomic — the SQLite file and the index directories are only consistent as a pair — and that restoring means opening the directory with a compatible Chroma version.

for a senior

Give the full procedure with quiescing and verification, name version skew as one-way, and describe migrating between deployment shapes as a data move rather than a config change.

for a principal

Own recovery as a design property: price the rebuild-from-source path against the restore path, insist that model version, chunking and id scheme are recorded as first-class metadata, and make the choice between them a decision already made before the outage.

## Why backup is a filesystem problem here Most databases give you an online backup or dump command that produces a consistent artefact while the server runs. Chroma does not: an embedded `PersistentClient` store and a `chroma run` server's store are both just a directory containing `chroma.sqlite3` and one UUID-named index directory per vector segment. So backup means copying that directory, and the entire difficulty is getting a consistent copy of two halves that are only meaningful together. ## The safe procedure 1. **Stop writes.** Ideally stop the process that owns the directory. If you cannot, at least stop ingestion so nothing is calling `add`, `upsert` or `delete` while you copy. 2. **Copy the whole directory.** Everything under the path, not just the SQLite file. `tar` or `rsync` the directory as a unit; a partial copy restores as a store whose index does not match its metadata. 3. **Verify the copy.** On the target, open it with `PersistentClient(path=...)` and check `count()` on each collection and run a known query. A restore you have never opened is a hypothesis, not a backup. 4. **Restart the source.** On a container platform, the equivalent is a volume snapshot — acceptable if the snapshot is genuinely atomic at the block level and taken with writes quiesced. A snapshot of a busy volume is the same torn-copy problem wearing different clothes. ## What makes a restore fail **Partial copies.** The most common cause. Somebody backs up `chroma.sqlite3` because it looks like "the database" and leaves the index directories behind. **Live copies.** `rsync` over a running store finishes without complaint and produces something that opens, queries, and returns subtly wrong results. **Version skew.** The on-disk format is coupled to the Chroma version. Opening a directory with a newer Chroma applies migrations to it on first startup; that step is effectively one-way, and the older version is not guaranteed to open the migrated directory afterwards. Two rules follow: pin the Chroma version alongside the data volume, and take a copy of the directory *before* an upgrade, because that copy is your only rollback. **Path assumptions.** Restoring into a different absolute path is fine — the client takes the path as an argument — but application configuration, container mounts and any tooling that hard-codes the old path have to move with it. ## Migration between deployment shapes Moving from embedded to server mode is a data move, not just a constructor change. Copy the directory to where the server will keep it, start `chroma run --path <that directory>`, and repoint the applications at `HttpClient`. Nothing imports a local store into a running server for you. Moving to a different vector database entirely is a different exercise, because there is no cross-vendor dump format. You read the records back out with `get(include=["embeddings", "documents", "metadatas"])`, paginating with `limit` and `offset`, and write them into the target. This is worth planning for early: keep the ids stable and meaningful, keep the source text and its chunking parameters somewhere outside the vector store, and record which embedding model produced the vectors. Without that last detail, exported embeddings are unusable — you cannot mix vectors from two models in one index, and you cannot regenerate matching ones for new documents. ## The plan that actually survives an incident For most RAG systems the strongest recovery story is not the directory copy at all: it is the ability to rebuild. If the source documents live in object storage or a relational database, the chunking is deterministic and pinned, and the embedding model and its version are recorded, then a lost Chroma store is a re-ingestion job — expensive in embedding cost and wall-clock time, but guaranteed correct and independent of on-disk format compatibility. Treat the directory copy as the fast path that saves hours, and the rebuild pipeline as the guarantee. Measure the rebuild's cost and duration before you need it, so "restore from backup or rebuild" is a decision you have already priced rather than one made at 3am. ## What interviewers are checking They want to hear that you know there is no online backup API, that the copy must be whole-directory and quiesced, and that version compatibility is one-way. The senior signal is treating the rebuild pipeline as part of the recovery design rather than trusting a file copy alone, and the ability to say what you would need to have recorded — model, chunking, ids — for either path to work.

  • What is wrong with rsyncing a live Chroma directory to another host every hour?
    It completes without error and can still produce a torn snapshot: the SQLite file and the index directories are copied at different instants, so the restored store's metadata and vectors disagree. Nothing validates that on open. Either quiesce writes for the copy, take an atomic volume snapshot, or accept that the artefact is best-effort and keep a rebuild path as the real guarantee.
  • You need to restore last week's backup but have since upgraded Chroma. Does it open?
    Usually yes in that direction — a newer Chroma opens an older directory and applies its schema migrations on first startup. The unsupported direction is the reverse: once a directory has been migrated by a newer version, an older Chroma is not guaranteed to open it. That is why the pre-upgrade copy matters as your rollback artefact.
  • What would you need recorded to rebuild a Chroma corpus from scratch after losing the volume?
    The source documents themselves in durable storage, the exact chunking parameters, the embedding model identity and version, the distance space each collection was created with, and the id scheme so references from other systems still resolve. Missing the model version is the killer: vectors from a different model are not comparable, so a partial rebuild silently degrades retrieval quality.

saying these in an interview costs you the question

  • Backing up only chroma.sqlite3 and calling it a backup
  • Copying the directory while the service is still accepting writes
  • Assuming an older Chroma can open a directory a newer version migrated
  • Never test-restoring the backup on another machine
  • Exporting embeddings without recording which model produced them

context