skip to content

In Chroma's query(), how do the where and where_document filters differ?

level: middleimportance: must knowfreq 72%

answer

  1. Two filters, two different record fields
  2. Metadata dictionary versus document text
  3. Mongo-style operators on one side
  4. Substring $contains on the other
  5. Scalar-only metadata values

basics

~20 s

where filters on the structured metadata dictionary attached to each record, using operators like $eq, $in and $gte. where_document filters on the stored document text itself, with $contains and $not_contains substring matching. They are separate arguments and combine with AND.

solid answer

~40 s

They target two different fields of the same record. `where` reads the `metadatas` dict: `where={"source": "docs", "year": {"$gte": 2024}}` keeps only records whose metadata satisfies those conditions, using `$eq`, `$ne`, `$gt`, `$gte`, `$lt`, `$lte`, `$in`, `$nin`, composed with `$and` and `$or`. A bare `{"key": value}` is shorthand for `$eq`. `where_document` reads the stored `documents` text and does literal substring matching with `$contains` and `$not_contains` — no tokenisation, no stemming, no relevance scoring, so it is a filter and not a keyword search engine. Passing both narrows on both conditions together. Metadata values must be scalars — string, int, float, or bool — so there is no filtering on list-valued fields; the usual workaround is to flatten tags into separate boolean keys or an encoded string.

code

python · 11 lines
python
res = collection.query(
    query_texts=["connection pool exhaustion"],
    n_results=5,
    where={
        "$and": [
            {"source": {"$in": ["handbook", "faq"]}},
            {"year": {"$gte": 2024}},
        ]
    },
    where_document={"$not_contains": "deprecated"},
)

go deeper

for a junior

Remember which filter reads which field: where looks at the metadata dictionary, where_document looks at the document text with $contains. Both are optional arguments to query().

for a middle

Be able to write a compound filter from memory — $in, $gte, an explicit $or — and to say that metadata values must be scalars, so list-valued tags need flattening into separate keys.

for a senior

Point out that where_document filters but never scores, so it cannot stand in for hybrid keyword-plus-vector relevance, and that inconsistent metadata types silently drop records from range filters.

for a principal

Own the ingest-side decision: the filterable metadata schema is a contract fixed at write time. Treat filter fields as a small, typed, enforced set, because a filter that silently misses records is worse than one that errors.

## Two filters over two different fields Every record Chroma stores is an id, a vector, an optional document string, and an optional metadata dictionary. The query API gives you one filter per non-vector field: - `where` filters the **metadata dictionary**. - `where_document` filters the **document text**. Both are optional, both narrow the candidate set before you get results, and when you pass both, a record must satisfy both. ## The where syntax Metadata filtering uses a small Mongo-flavoured operator set: - Equality: `$eq`, `$ne` - Ordering: `$gt`, `$gte`, `$lt`, `$lte` - Membership: `$in`, `$nin` - Composition: `$and`, `$or`, each taking a list of clauses A bare mapping is sugar: `where={"source": "handbook"}` means `{"source": {"$eq": "handbook"}}`. Several keys in one dict are combined with AND, so `{"source": "handbook", "year": {"$gte": 2024}}` requires both. When you need OR, you must say it explicitly with `$or`. The hard constraint people trip over is the **value type**. Chroma metadata values are scalars only: `str`, `int`, `float`, `bool`. You cannot store `{"tags": ["a", "b"]}` and you therefore cannot filter "has tag a". The two standard workarounds are to explode the list into boolean keys (`{"tag_a": True, "tag_b": True}`), which filters cleanly, or to store a delimited string and reach for `where_document`-style substring logic, which is fragile. The boolean-key approach is generally the right one; it costs metadata width but keeps filters exact. A second silent trap is **type consistency across records**. If some records store `{"year": 2024}` and others `{"year": "2024"}`, a range filter matches only the numeric ones and quietly drops the rest. Normalise types at write time. ## The where_document syntax `where_document={"$contains": "deadlock"}` keeps records whose document text contains that literal substring. `$not_contains` inverts it. You can compose these with `$and` / `$or` to express things like "mentions retry but not deprecated". What matters is what this is *not*. It is a substring test, not an inverted-index keyword search: there is no tokenisation, no stemming, no case-normalisation guarantee to lean on, and — crucially — **no contribution to the ranking**. A document that mentions the term forty times ranks exactly the same as one that mentions it once; the ordering is still purely vector distance among the surviving records. If the interview question is really "how do I blend keyword relevance with vector relevance", the honest answer is that this filter does not do that, and Chroma is not the tool you would reach for. ## How they interact with the search Chroma resolves the filters against its metadata store to obtain the set of ids that qualify, and restricts the nearest-neighbour search to that set. Two consequences follow. First, `n_results` counts *matching* records, so a filter never causes you to receive fewer results than exist in the filtered set — unlike systems that filter after taking a fixed candidate pool and then hand back a short list. Second, the cost of a query grows with how much of the collection the filter excludes, because the restricted search loses the benefit of the graph structure as the allowed set shrinks. The same two filters are accepted by `collection.get()` and by `collection.delete()`. `get(where=..., where_document=...)` is a pure filter scan with no query vector at all, which is the right tool for "show me every record from this source". `delete(where=...)` deletes everything that matches, which is powerful and unforgiving — run it as a `get` first and look at the count. ## Designing metadata for filtering Because filters are exact and typed, the quality of your filtering is decided at ingest time, not query time. Decide the small set of fields you will actually filter on — source, tenant, document type, timestamp bucket, language — and write them consistently on every record. Free-form metadata that varies by ingestion path produces filters that appear to work and silently miss data, which is far worse than a filter that errors.

  • How would you express "source is handbook OR faq, and the text does not mention deprecated"?
    `where={"$or": [{"source": {"$eq": "handbook"}}, {"source": {"$eq": "faq"}}]}` — or more compactly `{"source": {"$in": ["handbook", "faq"]}}` — combined with `where_document={"$not_contains": "deprecated"}`. The two arguments are ANDed together, so both conditions must hold. Keys inside a single where dict also AND, which is why the OR has to be spelled out explicitly.
  • Does a where_document $contains match improve a record's ranking?
    No. It only decides membership in the candidate set; ordering among survivors is still purely embedding distance. A record mentioning the term once and one mentioning it fifty times rank identically. That is the difference between a filter and a hybrid scoring system, and it is the honest limit to state if someone asks Chroma to do keyword relevance.
  • You need to filter on a tags list. What do you do?
    Chroma metadata values must be scalars, so you cannot store or filter a list. Flatten it: write one boolean key per tag, such as `{"tag_billing": True}`, and filter with `$eq`. That keeps the filter exact at the cost of metadata width. Encoding tags into one delimited string and substring-matching works but produces false positives on shared prefixes.

saying these in an interview costs you the question

  • Treating where_document as a scored keyword search
  • Storing a list as a metadata value
  • Expecting multiple keys in one where dict to OR
  • Mixing int and string types for the same metadata key
  • Believing where_document ranks results by term frequency

context