skip to content

What does Chroma write inside a PersistentClient path directory on disk?

level: middleimportance: should knowfreq 52%

answer

  1. look inside the directory, not just the API
  2. one relational file, one binary index tree
  3. copy the whole thing or nothing
  4. chroma.sqlite3 plus UUID segment directories

basics

~20 s

A chroma.sqlite3 file holds collections, ids, documents, metadata and the embedding records; alongside it sit UUID-named subdirectories, one per vector segment, containing the binary HNSW index files. Both parts belong to one database and must be copied together.

solid answer

~50 s

A Chroma persistent directory has two halves. `chroma.sqlite3` is the system-of-record: it stores the tenant/database and collection registry, plus every record's id, document text, metadata and embedding. Beside it, Chroma keeps one directory per vector segment, named with the segment's UUID, holding the serialised HNSW index binaries for that collection's vectors. The index directory is derived data — it exists so queries do not have to rebuild a graph from SQLite on every start — but the two halves are only consistent as a pair. Three practical consequences follow. Copying `chroma.sqlite3` alone gives you a store whose indexes do not match. SQLite is what grows with document text and metadata, so a corpus of large documents can make the file much bigger than the vectors themselves. And deletes mark records rather than shrinking the file, so disk usage does not fall when you delete a collection's contents.

code

bash · 8 lines
bash
ls -R ./chroma
# ./chroma:
# chroma.sqlite3
# 4b3a9e2c-0f21-4d55-9c8a-7e6f1b2d3c44
#
# ./chroma/4b3a9e2c-0f21-4d55-9c8a-7e6f1b2d3c44:
# (binary HNSW index files for one vector segment)
du -sh ./chroma

go deeper

for a junior

Know that a persistent Chroma store is a directory you can look inside, that it contains a SQLite file plus index data, and that the whole directory is the database.

for a middle

Explain the split: SQLite holds collections, ids, documents, metadata and embeddings; the UUID directories hold the serialised vector index. Be able to say why copying one without the other is broken.

for a senior

Bring the operational consequences: whole-directory backup atomicity, disk dominated by document text, deletes that do not reclaim space, and RAM sized by the resident index rather than the file.

for a principal

Own the lifecycle question — the on-disk format is coupled to the Chroma version, migrations run on first startup and are effectively one-way, so version pinning and a pre-upgrade copy belong in the deployment design, not in an incident runbook.

## The two halves of the directory When you construct `chromadb.PersistentClient(path="./chroma")`, Chroma creates and manages everything under that path. Listing it shows a shape like: - `chroma.sqlite3` — a single SQLite database file. - one or more directories named by a UUID — the persisted vector index for a segment. SQLite carries the relational side of Chroma: which tenants and databases exist, which collections exist and what metadata they carry (including index settings such as the distance space chosen at creation), and the records themselves — id, document text, metadata key/values, and the embedding. The UUID directories carry the vector-search side: the serialised HNSW graph files that let a query find nearest neighbours without a linear scan. ## Why the split matters operationally **They are one database, not two artefacts.** The index directories reference records that SQLite defines. Backing up or moving only the `.sqlite3` file, or only the index directories, produces a store that is at best missing its index and at worst inconsistent. Any copy, restore, or container-volume plan must treat the whole directory as the atomic unit — and copy it while nothing is writing. **Size is driven by the text, not just the vectors.** People size a Chroma deployment by multiplying vector count by dimensions by four bytes and are then surprised by the disk usage. Chroma stores the full document text you passed to `add(documents=...)` plus every metadata field in SQLite. For a RAG corpus of chunked documents, that text can easily exceed the embedding bytes. If you already have the text in another store and do not need Chroma to return it, you can keep documents out and store only a reference in metadata. **Deletes do not shrink the file.** SQLite reuses freed pages internally but does not return them to the filesystem by default. Deleting a large collection frees space for future inserts inside the same file, but the file on disk stays the same size. A long-lived store that churns — reindexing a corpus nightly, say — can therefore grow monotonically even though the logical row count is flat. Monitor the directory size, not the record count, and plan for occasional compaction or a rebuild into a fresh directory when the ratio drifts. **Memory follows the index, not the file.** The HNSW graph for a collection is loaded into memory to serve queries. That is what makes Chroma fast, and it is also the real capacity limit of a single-node deployment: the working set of vectors for the collections you actually query has to fit in RAM, whatever the disk footprint. A machine with plenty of disk and a small heap will still fall over on a large collection. ## What you should not do with the files Do not open `chroma.sqlite3` with your own SQLite client and write to it. The schema is internal, it changes across Chroma versions, and Chroma applies its own migrations on startup — hand edits will diverge from the index directories with no error to warn you. Reading it for diagnostics (row counts, collection listing) is a defensible last resort, but the supported path for every mutation is the client API. Equally, do not point two Chroma processes at the same directory expecting shared state. Each process opens its own SQLite handle and materialises its own copy of the HNSW index; writes from one are not reflected in the other's loaded index, and concurrent writers hit SQLite's locking. Sharing means running the Chroma server and having every process use `HttpClient`. ## Version coupling The on-disk format is tied to the Chroma version that wrote it. Upgrading Chroma applies schema migrations to `chroma.sqlite3` on the first startup against that directory. That is a one-way door: a directory written by a newer Chroma is not guaranteed to open under an older one. Pin the Chroma version in the same place you pin the data volume, and take a copy of the directory before an upgrade so you have a rollback that does not depend on downgrade support. ## What interviewers are checking They want to know whether you have looked inside the directory or only read the quickstart. The strong answer names both halves, explains that the index is derived but must be copied with the metadata, and draws at least one operational consequence — backup atomicity, text-dominated size, non-shrinking deletes, or the RAM-bound index.

  • You deleted half a collection's records and the directory did not get smaller. Why?
    SQLite marks the pages free for reuse inside the file rather than returning them to the filesystem, so the `chroma.sqlite3` file keeps its size and simply has room for future inserts. The index directory likewise does not compact on delete. If you genuinely need the space back, rebuild into a fresh directory and swap, or compact the SQLite file out of band while nothing is running.
  • How should you size RAM for a persistent Chroma store?
    By the vectors you query, not by disk. The HNSW index for a queried collection is held in memory, so roughly vector count times dimensions times four bytes, plus graph overhead, plus whatever the process needs. Document text and metadata stay in SQLite and do not have to be resident. A large disk footprint dominated by text is fine; a large resident index is what runs you out of memory.
  • Is it safe to query chroma.sqlite3 directly for a report?
    Reading is possible but unsupported — the schema is internal and changes across versions, so a report built on it will break on upgrade. Writing to it is genuinely unsafe: the vector index directories will not reflect your change and there is no error to tell you the two halves diverged. Use `get()` with the client for anything you need to depend on.

saying these in an interview costs you the question

  • Saying Chroma stores everything in flat parquet or JSON files
  • Backing up only chroma.sqlite3 and expecting queries to work
  • Sizing the deployment from vector bytes and ignoring stored document text
  • Editing the SQLite file by hand to fix data
  • Assuming a delete frees disk space immediately

context