skip to content

In Chroma, what does the hnsw:space collection metadata set, and what is its default?

level: middleimportance: should knowfreq 52%

answer

  1. It picks the distance metric
  2. Three legal values, one default
  3. The default is not cosine
  4. Baked into the index at creation
  5. Smaller distance means closer, scales differ

basics

~10 s

It sets the distance metric the collection's index uses: "l2" (the default, squared Euclidean), "ip" (inner product) or "cosine". You pass it in the metadata dict at creation, and it cannot be changed afterwards.

solid answer

~50 s

`hnsw:space` is a collection-level setting passed in the `metadata` dict when the collection is created, and it picks the distance function the index scores with: `"l2"` (squared Euclidean, the default), `"ip"` (inner product) or `"cosine"`. It matters because most sentence-embedding models are trained and evaluated under cosine similarity, so leaving the default in place scores them under a metric they were not built for — results are usually still plausible, which is what makes it easy to miss. The setting is fixed at creation: it is baked into the index, so you cannot flip it later, and calling `get_or_create_collection` with a different space on an existing collection just returns the existing collection with its original metric. Chroma also returns *distances*, not similarities — smaller is closer, and under cosine space the value is 1 minus cosine similarity.

code

python · 10 lines
python
import chromadb

client = chromadb.PersistentClient(path="./chroma")
collection = client.get_or_create_collection(
    name="docs_cosine",
    metadata={"hnsw:space": "cosine"},
)

# The stored value is the truth, not the argument you passed.
print(collection.metadata)

go deeper

for a junior

Know that hnsw:space chooses the distance metric, that the accepted values are l2, ip and cosine, and that l2 is what you get if you say nothing.

for a middle

Explain why the default often mismatches sentence-embedding models trained for cosine, that the metric is fixed at creation, and that results come back as distances where smaller is closer.

for a senior

Show that you set the metric in the provisioning path and verify the stored value, because get_or_create_collection silently returns an existing collection with its original metric no matter what you pass.

for a principal

Own the standard that distance thresholds are tuned per metric on real data and recorded with the space they belong to, so retrieval cut-offs do not travel between collections as unexamined constants.

## What the setting does A Chroma collection stores its index configuration in the `metadata` dict you supply at creation. The key `hnsw:space` names the distance function used both when the index is built and when a query is scored. Three values are accepted: - `"l2"` — squared Euclidean distance. The default when you say nothing. - `"ip"` — inner product (dot product), expressed so that smaller is closer. - `"cosine"` — cosine distance, i.e. 1 minus cosine similarity. You set it like this: ``` client.get_or_create_collection(name="docs", metadata={"hnsw:space": "cosine"}) ``` ## Why the default is worth thinking about Almost every general-purpose text-embedding model in common use — including the `all-MiniLM-L6-v2` model behind Chroma's default embedding function — is trained with a cosine objective and published with cosine-similarity benchmarks. Its vectors are meaningful in *direction*; their magnitude is an artefact. Squared Euclidean distance is sensitive to magnitude as well as direction. For vectors that happen to be near-uniform in length, L2 and cosine rank very similarly, which is exactly why the mismatch survives casual testing: your demo queries return sensible documents, so nothing looks broken. Where the two diverge is on the tail — long documents, unusual inputs, anything whose vector norm drifts from the pack — and that tail is where retrieval quality is actually decided. The practical rule: match the metric to what your embedding model was trained for. For the common sentence-transformer and hosted text-embedding families, that is cosine. Set it deliberately rather than inheriting the default. Inner product is the right choice when a model is explicitly documented as a dot-product model, or when you have already normalised your vectors — for unit-length vectors, inner product and cosine rank identically, and inner product avoids re-computing norms. ## It is fixed at creation The space is compiled into the index structure, not evaluated per query, so it is not editable after the fact. There is no "change the metric" operation. Changing it means creating a new collection with the right metadata and re-adding the records. The subtler consequence involves `get_or_create_collection`. When the named collection already exists, that call is a *get*: you receive the existing collection, with the metric it was created with. Passing `metadata={"hnsw:space": "cosine"}` to a collection that was born with the L2 default does not upgrade it and does not raise. Code that looks entirely correct — every call site names cosine — can be operating on an L2 index because the very first run, months earlier, forgot the metadata. If you need certainty, inspect the collection's stored metadata rather than trusting the argument you passed. This also means the setting belongs in whatever provisioning step creates the collection, not scattered across every module that obtains a handle to it. ## Distances, not similarities Chroma returns distances in query results, and the direction is consistent across all three spaces: **smaller means closer**. The scale, however, is not. - Under `"cosine"`, the value is 1 − cosine similarity, so it runs from 0 (identical direction) to 2 (opposite), with 1 meaning orthogonal. - Under `"l2"`, it is the *squared* Euclidean distance, unbounded above and quadratic in the actual geometric distance. - Under `"ip"`, it is derived from the dot product and is not bounded to a friendly range at all. So any threshold you hard-code — "drop results with distance above 0.35" — is meaningful only for the space it was tuned in. Moving a collection from L2 to cosine silently invalidates every such constant, and a threshold copied out of a blog post written for a different setup is worse than no threshold. Tune cut-offs against your own data, and record which space they were tuned under. ## Other index settings live in the same dict `hnsw:space` is one of a family of index keys on the same metadata dict; the graph-construction parameters (`hnsw:M`, `hnsw:construction_ef`) and the search-time breadth (`hnsw:search_ef`) sit beside it and follow the usual build-cost-versus-recall tradeoffs. Two practical notes. First, the same dict also holds *your* application metadata, so keep your own keys clearly distinct from the `hnsw:` prefix. Second, the space is the one setting where the default is genuinely likely to be wrong for your model, whereas the graph parameters have defaults that are fine until you have measured a reason to move them.

  • You need cosine on a collection that was created with the default. What do you do?
    Create a new collection with `metadata={"hnsw:space": "cosine"}` and re-add the records; the metric is fixed at creation and there is no in-place change. Note that `get_or_create_collection` on the existing name will simply return the old L2 collection and ignore the metadata you passed, so use a new name and cut over.
  • When would inner product be the right choice over cosine?
    When the model is documented as a dot-product model, or when your vectors are already unit-normalised. For unit-length vectors inner product and cosine produce identical rankings, so `"ip"` avoids redundant norm computation. Choosing it for vectors of varying magnitude, though, lets long-vector records dominate results regardless of direction.
  • Why is a hard-coded distance threshold risky across collections?
    Because the scale differs per space. Cosine distance is bounded 0 to 2, squared L2 is unbounded and quadratic, and inner-product values have no friendly range at all. A cut-off tuned under one metric is meaningless under another, so tune against your own data and record which space it was tuned in.

saying these in an interview costs you the question

  • Assumes cosine is Chroma's default distance metric
  • Thinks the metric can be changed on an existing collection
  • Believes get_or_create_collection updates metadata on an existing collection
  • Reads returned distances as similarity scores where higher is better
  • Copies a distance threshold across collections with different metrics

context