In Haystack, how do you express AND/OR conditions in a metadata filter?
answer
- Two node shapes, nothing else
- Three keys or two keys
- Nesting replaces operator precedence
- Metadata lives behind a prefix
- Dollar-sign operators are the old version
basics
~10 sHaystack filters are nested dicts of two node shapes. A comparison node is {"field", "operator", "value"}; a logical node is {"operator": "AND"|"OR"|"NOT", "conditions": [...]} holding further nodes. Metadata fields must be written as meta.<key>.
solid answer
~50 sA Haystack filter is a tree built from exactly two node types. A **comparison** node has three keys — `field`, `operator`, `value` — with operators `==`, `!=`, `>`, `>=`, `<`, `<=`, `in` and `not in`. A **logical** node has `operator` set to `"AND"`, `"OR"` or `"NOT"` and a `conditions` list holding further nodes of either kind, so arbitrary nesting is just logical nodes inside logical nodes. Metadata keys must be addressed with a `meta.` prefix — `"meta.year"`, not `"year"` — while top-level Document fields such as `content` are referenced directly. The same dict is accepted by `document_store.filter_documents(filters=...)` and by retrievers' `filters` input, so you can develop and verify a filter directly against the store before wiring it into a query pipeline. The MongoDB-style `{"year": {"$eq": 2024}}` shape belongs to Haystack 1.x and is not this format.
code
python · 18 linesfrom haystack import Document
from haystack.document_stores.in_memory import InMemoryDocumentStore
store = InMemoryDocumentStore()
store.write_documents([
Document(content="Q3 results", meta={"year": 2024, "type": "report"}),
Document(content="Q3 recap", meta={"year": 2024, "type": "blog"}),
Document(content="Q3 2019", meta={"year": 2019, "type": "report"}),
])
filters = {
"operator": "AND",
"conditions": [
{"field": "meta.year", "operator": ">=", "value": 2024},
{"field": "meta.type", "operator": "in", "value": ["report", "memo"]},
],
}
print(len(store.filter_documents(filters=filters))) # 1go deeper
Memorise the two dict shapes and the meta. prefix. Being able to write a two-condition AND filter correctly on a whiteboard is the whole bar at this level.
Explain nesting, the full operator list, and that in takes a list on the value side. Recognise the 1.x dollar-sign syntax as legacy and be able to translate it.
Show that you debug filters with filter_documents before blaming retrieval, know that stores translate filters into their own query languages with edge-case differences, and treat filterable fields as a schema decision.
Own the security angle: mandatory scoping conditions belong in the outermost AND, assembled by shared code rather than by each caller, so no query path can accidentally widen the tenant or permission scope.
## Two node types, arbitrarily nested Everything in a Haystack filter is one of two dict shapes. **Comparison node** — exactly three keys: `{"field": "meta.type", "operator": "==", "value": "article"}` **Logical node** — exactly two keys: `{"operator": "AND", "conditions": [ <node>, <node>, ... ]}` Because `conditions` accepts nodes of either kind, nesting is free: an `OR` of two `AND`s is just a logical node whose conditions are logical nodes. There is no special syntax for grouping and no operator precedence to remember — the structure *is* the precedence. ## The operators Comparison operators are `==`, `!=`, `>`, `>=`, `<`, `<=`, `in`, `not in`. The membership operators take a list as their `value`: `{"field": "meta.type", "operator": "in", "value": ["report", "memo"]}` means "the document's type is one of these". Note the direction — it reads as *field in value*, so `value` is always the collection. Logical operators are `"AND"`, `"OR"`, `"NOT"`, written uppercase as strings. ## The meta. prefix This is the single most common mistake. `Document` has top-level attributes (`content`, `id`, `score`, `embedding`) and a `meta` dict for everything you attached. Metadata fields must be addressed through that dict: `"meta.year"`, `"meta.file_path"`, `"meta.tenant_id"`. Writing `{"field": "year", ...}` is asking about a top-level Document field named `year`, which does not exist — depending on the store you get either an error or, worse, an empty result set that looks like "retrieval found nothing" rather than "your filter is malformed". The fastest way to catch this: run the filter directly against the store with `filter_documents(filters=...)` and check the count before you ever put it in a pipeline. `filter_documents` performs no ranking, no embedding and no retrieval — it is pure metadata selection, which makes it the right debugging tool. ## Where filters are used The same dict is accepted in three places: `DocumentStore.filter_documents(filters=...)` for direct selection, retrievers' `filters` run-time input, and — on most retrievers — a `filters` constructor argument that acts as a default. Where both a constructor default and a run-time value exist, the run-time value is what applies to that run, so the constructor form is best reserved for an invariant such as a tenant scope you never want overridden by accident. ## What the store does with it Haystack defines the filter format; each `DocumentStore` implementation translates it into its backend's own query language — Elasticsearch/OpenSearch query DSL, a Qdrant filter, a SQL `WHERE` clause for pgvector, and so on. `InMemoryDocumentStore` evaluates it in Python. Two consequences follow. First, portability is good but not absolute: comparison semantics on missing keys, on `None`, and on nested meta paths can differ between backends, so a filter that behaves one way in tests against the in-memory store deserves a check against the real one. Second, filter performance is entirely the backend's problem — a filter over an unindexed metadata field in a database-backed store is a scan, which is why filterable fields want to be low-cardinality scalars rather than long free text. ## Legacy syntax Haystack 1.x used a MongoDB-flavoured shape: `{"type": {"$eq": "article"}}`, with `$and`, `$or`, `$in`. Haystack 2 replaced it with the explicit `field`/`operator`/`value` form described here, and 3.x keeps that form. Most training material and half the blog posts on the internet still show the old shape, so recognising and rejecting it is a genuine interview discriminator. If you meet the old syntax in a codebase you are migrating, the translation is mechanical: each `{key: {"$op": v}}` becomes a comparison node with `meta.` prepended to the key, and each `$and`/`$or` becomes a logical node. ## Practical shape Filters are plain Python dicts, so build them programmatically — a helper that assembles the `conditions` list from whatever scoping the request carries (tenant, language, date window) is far more maintainable than string-formatting a filter, and it makes the mandatory conditions impossible to forget. Keep the mandatory security scope as the outermost `AND` so that no caller-supplied condition can widen it.
- How do you filter on a top-level Document field rather than metadata?Reference it without the prefix — `{"field": "content", "operator": "!=", "value": ""}`. The `meta.` prefix exists precisely to distinguish the Document's own attributes (`content`, `id`, `score`, `embedding`) from keys in its `meta` dict. Getting the prefix wrong usually yields an empty result set rather than an error, which is why it is worth verifying with `filter_documents` first.
- How do you express "type is one of report or memo"?Use a single comparison node with the `in` operator and a list value: `{"field": "meta.type", "operator": "in", "value": ["report", "memo"]}`. It reads as *field in value*, so the collection is always on the `value` side. An `OR` of two `==` nodes is equivalent but more verbose and does not scale to a long list.
- Does the same filter behave identically across every document store?Mostly, but not guaranteed. Haystack defines the format and each store translates it into its own backend query language, so edge cases — missing keys, null values, nested meta paths, type coercion between strings and numbers — can differ. Validate filters against the store you actually deploy on, not only against InMemoryDocumentStore in tests.
- Why is filter_documents useful even though it does no retrieval?It applies only the metadata predicate: no embedding, no ranking, no top_k truncation. That makes it the right tool for answering "is my filter wrong, or is my retrieval wrong?" — if `filter_documents` returns zero documents, the filter is the bug. It is also how you find the ids of a source's chunks before deleting them.
saying these in an interview costs you the question
- Writing MongoDB-style $eq or $and operators from Haystack 1.x
- Omitting the meta. prefix on metadata field names
- Putting the collection on the field side of an in comparison
- Assuming a filter matched when an empty result was really a bad field name
- Expecting nested AND/OR to need special grouping syntax