skip to content

What breaks when two processes open the same Chroma PersistentClient path?

level: seniorimportance: must knowfreq 50%

answer

  1. embedded engine, not a shared service
  2. each process caches its own index
  3. one worker in dev, four in prod
  4. a server process must own the directory

basics

~20 s

A persistent Chroma client is embedded, so each process gets its own SQLite handle and its own in-memory copy of the vector index. Writes by one process are invisible to the other's loaded index, and concurrent writers hit SQLite locking. Run the Chroma server and use HttpClient instead.

solid answer

~50 s

`PersistentClient` is an embedded database, not a shared one. Two OS processes pointed at the same directory are two independent database engines over the same files. Two things go wrong. First, correctness: each process loads the collection's HNSW index into its own memory, so documents added by worker A are not in worker B's index — queries return stale results with no error, and the staleness is per-worker, which makes it look intermittent. Second, contention: SQLite serialises writers with file locks, so concurrent writes surface as lock timeouts, and in the worst case the two processes' views of metadata and index diverge. This bites the moment you deploy a Gunicorn or Uvicorn app with more than one worker, because it works perfectly in development with one process. The fix is to make exactly one process own the storage: run `chroma run --path ./chroma_data --host 0.0.0.0 --port 8000` and have every worker connect with `chromadb.HttpClient`.

code

python · 10 lines
python
import chromadb

# Embedded: this process owns ./chroma_data. Do not run two of these.
local = chromadb.PersistentClient(path="./chroma_data")

# Shared: one server owns the data, every worker is a stateless client.
# Server side:  chroma run --path ./chroma_data --host 0.0.0.0 --port 8000
remote = chromadb.HttpClient(host="localhost", port=8000)
remote.heartbeat()
col = remote.get_or_create_collection("docs")

go deeper

for a junior

Remember the rule of thumb: one persistent Chroma directory, one process. If more than one process needs the data, something has to run as a server.

for a middle

Explain why — the client is embedded, so each process has its own engine and its own in-memory index — and name both symptoms, stale reads and SQLite write locking.

for a senior

Diagnose it from the outside: intermittent, per-worker inconsistent search results with no exception in the logs, appearing the day the service went to multiple workers. Then give the fix and what the migration actually involves.

for a principal

Treat it as a deployment-topology decision made too early: an embedded store bakes single-process ownership into the architecture, so decide up front whether the service is one writer forever, or plan the server-mode boundary before horizontal scaling forces it under incident pressure.

## Embedded means one owner Chroma's `PersistentClient` embeds the whole database inside your Python process — the SQLite store, the vector index, and the query engine. That is what makes it so pleasant to start with: no server, no ports, no ops. It also means the process is the database. The directory on disk is state that one engine is managing, not a shared protocol two engines can speak at once. Inside a single process, constructing `PersistentClient` on the same path more than once is fine — Chroma hands back a shared instance for the same path and settings, so a module that builds a client in two places does not create two engines. Across processes, nothing coordinates them. ## The two distinct failures **Stale reads, silently.** When a collection is first queried, its HNSW index is materialised in memory. Worker A adds fifty documents; its own in-memory index now contains them. Worker B, which loaded the index earlier, has no idea — nothing invalidates its copy. A user who hits worker A sees the new documents and a user who hits worker B does not. Because load balancers spread requests, the bug presents as "search results are inconsistent between refreshes", which sends people hunting for a caching bug in their own application. There is no exception anywhere. **Write contention and divergence.** SQLite permits many readers but one writer, enforced with file locks. Two Chroma processes writing concurrently produce lock waits and "database is locked" style failures under load. Worse, the SQLite half and the index half of the store are maintained by whichever process performed the write; a second process holding its own materialised index has no path to learn about the change. Over time the two processes' notions of what a collection contains drift apart, and whichever one is restarted last wins. Add network filesystems and it gets worse still. Pointing a persistent path at NFS or an S3-backed FUSE mount so that "both containers can see it" is a classic mistake: SQLite's locking is unreliable on many network filesystems, so you lose even the protection that a local disk gives you. A Chroma persistent path belongs on a local disk or a single-attach block volume. ## Where this hits in practice The canonical incident is a FastAPI or Flask service that builds a `PersistentClient` at import time and is deployed under Gunicorn with `--workers 4`. Development runs one worker and is flawless. Production forks four, each of which now owns a copy of the engine. Container platforms reproduce it another way: scale a deployment to two replicas that mount the same volume, and you have two engines again — this time on different machines. Threads within one process are a different question and are largely fine: one engine, one index, with Chroma serialising as needed. The line to hold in your head is *processes*, not threads or requests. ## Fixing it The supported fix is client-server mode. Run one Chroma server that owns the directory — `chroma run --path ./chroma_data --host 0.0.0.0 --port 8000`, or the `chromadb/chroma` container with the data directory on a mounted volume — and have every application process use `chromadb.HttpClient(host=..., port=...)` (or `AsyncHttpClient` in async code). Now there is exactly one engine, one index in one memory space, and the workers are stateless clients. The collection API you were already calling does not change, which makes the migration mostly a constructor swap plus moving the existing directory to wherever the server runs. If you cannot run a server, the alternatives are all forms of restoring single ownership: run the service with a single worker; or split the roles so that one dedicated writer process owns the store and the readers query it through that process. A weaker but occasionally sufficient pattern for read-mostly corpora is to build the store offline, ship it as a read-only artefact baked into the image, and have each replica open its own private copy — no writes, no sharing, and a redeploy is how content is updated. ## What interviewers are checking This is the question that separates people who have deployed Chroma from people who have run the quickstart. They want the mental model — embedded engine, one owner — and both symptoms, especially the silent one. Naming the multi-worker web server as the place it appears, and `chroma run` plus `HttpClient` as the fix, is the complete answer.

  • Is it safe for multiple threads in one process to share a PersistentClient?
    Yes, in the sense that matters here — one process means one engine and one in-memory index, so threads do not create the divergence that separate processes do. Chroma handles the necessary serialisation internally. Heavy concurrent writes from many threads will still contend on the single writer underneath, so throughput is bounded, but you do not get stale-per-thread reads.
  • Could you put the persistent path on a shared network volume so two containers can use it?
    No. Beyond the stale-index problem, which the filesystem cannot solve, SQLite's file locking is unreliable on NFS and similar network filesystems, so you lose the one guard against concurrent writers. Keep a persistent path on local disk or a single-attach block volume and share through the Chroma server instead.
  • How would you make a read-only Chroma store work across several replicas without a server?
    Build the store once in an offline job, then treat the directory as an immutable artefact — bake it into the image or copy it into each replica's own local path at startup. Every replica opens its private copy, nothing writes at runtime, so there is no coordination problem. Updating the corpus becomes a rebuild and redeploy rather than a live write.

saying these in an interview costs you the question

  • Assuming SQLite file locking makes concurrent Chroma processes safe
  • Expecting one process to see another's writes after a reload
  • Putting the persistent directory on NFS so replicas can share it
  • Blaming intermittent stale results on application caching
  • Thinking more Gunicorn workers scale an embedded Chroma store

context