skip to content

When are Chroma's tenants and databases enough to isolate multiple customers?

level: principalimportance: nice to knowfreq 26%

answer

  1. a namespace is not a boundary
  2. who enforces the scoping?
  3. one process, one memory, one blast radius
  4. fine for environments, thin for customers

basics

~20 s

Tenants and databases are namespacing, not isolation: they scope collection names inside one store but share the same process, memory, disk and index resources. They suit internal separation of environments or teams, not untrusted customers needing enforced boundaries or independent capacity.

solid answer

~60 s

Every Chroma client operates within a tenant and a database — defaulting to `DEFAULT_TENANT` and `DEFAULT_DATABASE` — and you can pass `tenant=` and `database=` to any client constructor to scope it elsewhere, creating them through the admin client. What that buys is a clean namespace: two customers can each have a collection called `documents` without colliding, and a client bound to one database cannot see another's collections by name. What it does not buy is isolation in any operational sense. All tenants live in the same SQLite store, the same process, the same memory, and compete for the same index cache and CPU; there is no per-tenant quota, no per-tenant authentication in the open-source server, and no per-tenant backup or restore. So they are the right tool for separating dev from prod, or one team's corpora from another's inside a trusted deployment. For untrusted customers who need enforced access boundaries, independent capacity, or the ability to delete or export one customer's data cleanly, the boundary has to be a separate store or a separate deployment.

code

python · 12 lines
python
import chromadb
from chromadb import DEFAULT_TENANT

# Scope a client to one customer's database within a shared store.
# The mapping from authenticated principal -> database name is YOUR enforcement point.
client = chromadb.HttpClient(
    host="localhost",
    port=8000,
    tenant=DEFAULT_TENANT,
    database="customer_acme",
)
print([c.name for c in client.list_collections()])  # only customer_acme's collections

go deeper

for a junior

Know that clients default to DEFAULT_TENANT and DEFAULT_DATABASE, and that passing tenant= and database= scopes which collections a client can see.

for a middle

Explain that this is namespacing: it prevents name collisions and scopes listing, but everything still shares one store, one process and one set of resources.

for a senior

Separate namespacing from isolation concretely — no per-tenant authentication in the open-source server, no quotas, one shared blast radius, and no per-tenant backup or atomic delete — and say where enforcement therefore has to live.

for a principal

Own the tradeoff: price shared-namespace simplicity against per-customer deployments, define the trigger that promotes a customer to a dedicated store, and design the data path so that promotion stays a configuration change rather than a rewrite.

## What tenants and databases actually are Chroma organises collections in a two-level namespace above them: a tenant contains databases, and a database contains collections. Both have defaults — `chromadb.DEFAULT_TENANT` and `chromadb.DEFAULT_DATABASE` — which is why most code never mentions them. Any client constructor accepts `tenant=` and `database=`, and an admin client (`chromadb.AdminClient`) creates them with `create_tenant` and `create_database`. Once a client is bound, its `list_collections`, `get_or_create_collection` and everything else operate only within that database. That is genuinely useful. Name collisions disappear, listing is scoped, and an accidental `delete_collection` in one database cannot touch another's. It is the same kind of value a schema gives you in a relational database. ## The gap between namespacing and isolation The mistake is reading "multi-tenant" as "tenant-isolated". Consider what a tenant boundary would have to provide for real customer separation, and what Chroma's provides: **Access control.** A namespace only helps if something enforces who may bind to it. The open-source Chroma server does not authenticate per tenant; access control is configured for the server as a whole, so any client that can authenticate to the server can typically name any tenant and database it likes. The enforcement therefore lives in *your* application — you decide, per request, which tenant the client is scoped to. That is a perfectly workable design, but it means one bug in that mapping is a cross-customer data leak, and no layer beneath you will catch it. If your compliance story requires the datastore itself to enforce the boundary, this design does not deliver it. **Resource isolation.** All tenants share one process. The HNSW indexes for whichever collections are being queried live in the same memory; one customer ingesting ten million vectors evicts nothing politely and can push the process into swap or OOM for everyone. There are no per-tenant quotas, rate limits, or memory ceilings. Noisy-neighbour behaviour is unmediated. **Blast radius.** One directory, one SQLite file, one server process. A corrupted store, a bad upgrade, or an out-of-disk condition takes down every tenant at once. Availability is shared whether you want it to be or not. **Data lifecycle.** "Delete everything for customer X" is a loop over their collections, not an atomic operation — and "export everything for customer X" is a read-out through the query API. If you have deletion deadlines or portability obligations, you have to build and test that machinery yourself, and prove it, because there is no per-tenant dump or drop that does it for you. ## Where the line sits Use tenants and databases when the parties are mutually trusted and the value you want is organisational: separating environments, keeping teams' corpora from colliding, or giving each feature its own namespace within one application. In those cases the shared process is not a risk, it is the point — one deployment to run, one backup to take. Separate stores — a directory or a server per customer — when any of these are true: customers are untrusted or contractually isolated; one customer's volume could starve the others; a customer needs independent backup, restore, or deletion; different customers need different upgrade windows. The cost is real and worth stating: N deployments to operate, N sets of memory that cannot be pooled, and per-customer operational overhead that only pays off above a certain contract value. That is exactly the tradeoff a lead is expected to reason about out loud. A hybrid is common and defensible: many small customers share one deployment with database-per-customer namespacing and application-enforced scoping, while large or regulated customers get dedicated stores. The important part is that the promotion path exists and that the application already routes by tenant, so moving a customer to a dedicated deployment is a configuration change rather than a rewrite. ## Designing so the decision stays reversible Whatever you choose first, keep the customer identifier a first-class part of the data path: derive the tenant and database from the authenticated principal in one place, never from a request parameter, and put the customer id in record metadata as well so a mis-scoped read is at least detectable after the fact. Keep ingestion re-runnable per customer so a dedicated store can be populated from source rather than migrated file-by-file. Those two habits are what make "we outgrew shared namespacing" a planned migration instead of an emergency. ## What interviewers are checking This is a judgment question with no single right answer. The weak answer is "Chroma supports multi-tenancy, so use tenants". The strong one distinguishes namespacing from isolation across at least three axes — enforcement, resources, blast radius — states plainly where the application has to carry the enforcement, and prices the alternative rather than asserting that separate deployments are simply safer.

  • If the application enforces tenant scoping, how do you keep that from becoming a leak?
    Derive the tenant and database from the authenticated principal in exactly one place — never from a client-supplied parameter — and make the scoped client the only way any code path reaches Chroma. Add the customer id to record metadata too, so a mis-scoped read can be detected in audit rather than passing silently, and test the boundary explicitly with a cross-tenant read attempt in CI.
  • What would make you promote one customer from a shared namespace to a dedicated store?
    Volume that makes their working set dominate the shared index memory; a contractual or regulatory demand for isolated storage, independent deletion or a separate backup; or an availability commitment that cannot tolerate sharing a blast radius with everyone else. The trigger should be defined in advance so the migration is scheduled work rather than an incident response.
  • Does giving each customer their own collection instead of their own database change the isolation story?
    Only cosmetically. Collections in the same database still share the process, memory, disk and failure domain, and the scoping is still enforced by your application choosing the collection name. It is a slightly flatter namespace with the same properties — convenient, and equally not a security boundary.

saying these in an interview costs you the question

  • Calling Chroma tenants a security boundary between customers
  • Expecting per-tenant quotas or rate limits to exist
  • Assuming deleting a tenant's data is one atomic operation
  • Taking the tenant name from a request parameter instead of the session
  • Believing separate tenants get separate backups or upgrade windows

context