skip to content

How do you run Chroma as a server and connect clients to it?

level: middleimportance: should knowfreq 45%

answer

  1. the package ships a server, not just a client
  2. the constructor is the only code change
  3. a container without a volume forgets
  4. nothing guards the port by default

basics

~20 s

Start the server with chroma run --path ./chroma_data --host 0.0.0.0 --port 8000, or run the chromadb/chroma container with its data directory on a mounted volume. Applications then use chromadb.HttpClient(host=..., port=...) and call heartbeat() to verify connectivity.

solid answer

~50 s

Chroma ships a server in the same package. `chroma run --path ./chroma_data --host 0.0.0.0 --port 8000` starts a process that owns that directory and serves the same API over HTTP; the official `chromadb/chroma` image does the same thing in a container. Client code changes only at construction: `chromadb.HttpClient(host="chroma", port=8000)` — or `AsyncHttpClient` in asyncio code — and every collection call after that is identical to the embedded API. `client.heartbeat()` is the readiness check. The two deployment mistakes to call out are storage and exposure. If the container's data path is not a mounted volume, the store dies with the container, which is the same "my data vanished" bug one layer up. And the server is not authenticated by default, so anything that can reach the port can read and delete every collection — bind it to a private network and put auth in front of it before it is reachable from anywhere else.

code

bash · 1 line
bash
chroma run --path ./chroma_data --host 0.0.0.0 --port 8000

go deeper

for a junior

Know that chroma run --path ... starts a server and that application code connects with chromadb.HttpClient(host, port), with the collection API unchanged afterwards.

for a middle

Explain that the server owns the same directory format an embedded client would create, that the code change is only the constructor, and that the container needs a mounted volume or the data is ephemeral.

for a senior

Lead with the operational risks: an unauthenticated endpoint that can delete collections, single-owner storage that must not be replicated over one volume, and network semantics that make batching and idempotent upserts matter.

for a principal

Own the boundary decision — when the store stops being a library inside the app and becomes a service with its own availability, backup and access-control story, and what the team owes it once the prototype is on the critical path.

## Starting the server Installing the Chroma Python package also installs the `chroma` CLI. `chroma run --path ./chroma_data --host 0.0.0.0 --port 8000` starts a long-running server process that opens the given directory — exactly the same on-disk store an embedded `PersistentClient` would create — and exposes the API over HTTP. `--host 0.0.0.0` binds all interfaces, which is what you need inside a container; on a developer machine, leaving it on localhost is the safer default. In container form, the official image is `chromadb/chroma`. It runs the same server; your job is to publish the port and, critically, to mount a volume at the data directory the server is configured to use. A container without a volume is a persistent store with an ephemeral lifetime — every `docker run` of a fresh container starts empty, and the first time someone recreates the container to change an environment variable, the corpus is gone. ## Connecting ``` client = chromadb.HttpClient(host="chroma", port=8000) client.heartbeat() col = client.get_or_create_collection("docs") ``` That is the entire migration from embedded mode: the constructor. `get_or_create_collection`, `add`, `query`, `get`, `delete` all behave as before, because the client presents one interface over both transports. `chromadb.AsyncHttpClient` is the awaitable equivalent for asyncio services. `heartbeat()` is the cheap liveness probe — use it in a container healthcheck or at startup so a misconfigured host or port fails loudly rather than on the first user query. What does *not* migrate is the data. The server owns its own directory; pointing clients at it does not import whatever an earlier `PersistentClient` wrote locally. Moving an existing prototype means copying that directory to the server's storage while nothing is running. ## The things that bite in deployment **Persistence of the volume.** Say it twice because it fails twice: on the client side by choosing an ephemeral client, and on the server side by not mounting a volume. Both look fine until a restart. **No authentication by default.** A freshly started Chroma server accepts any request that reaches it. There is no default credential, and the API includes deleting collections. Treat the port as internal-only: bind it to a private network or an overlay network between your services, never publish it to the internet, and put a token-authentication configuration or an authenticating reverse proxy in front of it if anything less trusted can reach it. `chromadb.HttpClient` accepts a `headers` dictionary, which is how a token or proxy credential is attached from the client side. **One server, one store.** Running two server replicas over the same volume recreates the multi-process problem the server was supposed to solve — the server is a single owner, not a distributed cluster. Scale the application tier horizontally; keep the Chroma server a single process against a single-attach volume, and size that machine for the resident index. **Client-side work still happens client-side.** In server mode the Python client is still where documents are turned into embeddings unless the server is configured to do it, so each application container may still need the embedding model available to it. That matters for image size, cold start, and CPU sizing of the application tier — it is easy to assume that "moving to a server" moved the embedding cost too. **Network semantics arrive with the transport.** Every call is now an HTTP request that can time out, be retried, or fail mid-batch. Large `add` calls that were a local function call become large request bodies; batch them, and make ingestion idempotent by using stable ids with `upsert` so a retry after a timeout does not duplicate records. ## When to make the move Go to server mode when more than one process needs the data, when storage must outlive the application container, when the ingestion job and the serving app are separate deployables, or when you want to restart the app without paying to reload a large index. Stay embedded when a single process owns everything and the simplicity is worth more than the flexibility — that is a real and defensible choice for a CLI tool, a batch job, or a desktop application. ## What interviewers are checking They want the command, the client constructor, and — more than either — the two operational gotchas: the volume and the open port. A candidate who describes the migration as "just swap the constructor" without mentioning that the data has to be moved and the endpoint has to be protected is answering from the README.

  • How do you verify from a client that the server is actually reachable and healthy?
    Call `client.heartbeat()` right after constructing the `HttpClient` — it is a cheap round trip that fails immediately on a wrong host, port, or a server that is not up yet. Wire it into your service's startup check or the container healthcheck so misconfiguration surfaces at deploy time rather than on a user's first query.
  • Can you run two Chroma server replicas behind a load balancer over the same volume?
    No — that reintroduces two engines over one store, with the same stale-index and write-contention problems as two embedded clients. The Chroma server is a single-owner process, not a clustered database. Scale the stateless application tier and keep one server against one volume, sizing that host for the resident index.
  • What changes about error handling when you move from PersistentClient to HttpClient?
    Calls become network calls, so timeouts, partial failures and retries are now yours to handle. Keep `add`/`upsert` batches modest so a request body does not blow a timeout, and use stable ids with `upsert` so retrying an ambiguous failure is idempotent rather than duplicating records.

saying these in an interview costs you the question

  • Publishing the Chroma server port publicly because it has no login page
  • Running the container without mounting a volume for the data directory
  • Believing HttpClient automatically migrates an existing local persistent store
  • Running several server replicas over one shared volume for high availability
  • Assuming server mode moves embedding computation off the application process

context