skip to content

In Qdrant, how do you restrict a vector search to points matching a payload condition?

level: juniorimportance: must knowfreq 80%

answer

  1. metadata lives beside the vector
  2. one object, many call sites
  3. query_filter on the search call
  4. FieldCondition = key plus match rule
  5. filters gate eligibility, not score

basics

~10 s

Pass a filter to the search call. In the Python client that is query_points(..., query_filter=models.Filter(must=[models.FieldCondition(key="lang", match=models.MatchValue(value="en"))])). Only points whose payload satisfies the filter come back, still ranked by vector distance.

solid answer

~40 s

Every Qdrant point carries a JSON `payload` alongside its vector, and search calls take a filter over that payload. In the Python client you build `models.Filter(...)` out of `models.FieldCondition` objects, each naming a payload `key` plus a condition — `match=models.MatchValue(value=...)` for an exact value, `match=models.MatchAny(any=[...])` for a value in a list, `range=models.Range(gte=..., lt=...)` for numbers, `models.GeoRadius` for coordinates. You hand it to `query_points` as `query_filter`. The same `Filter` object is reused by other calls under different parameter names: `scroll(scroll_filter=...)`, `count(count_filter=...)`, and `delete(points_selector=models.FilterSelector(filter=...))`. Filtering is purely boolean — it decides which points are eligible, never how they are scored, so ranking stays driven by the distance metric alone.

code

python · 19 lines
python
from qdrant_client import QdrantClient, models

client = QdrantClient(url="http://localhost:6333")

response = client.query_points(
    collection_name="articles",
    query=[0.1] * 384,
    query_filter=models.Filter(
        must=[
            models.FieldCondition(key="lang", match=models.MatchValue(value="en")),
            models.FieldCondition(key="views", range=models.Range(gte=100)),
        ]
    ),
    limit=10,
    with_payload=True,
)

for point in response.points:
    print(point.id, point.score, point.payload)

go deeper

for a junior

Be able to write a filtered search from memory: build a Filter with FieldCondition entries, pass it as query_filter, and name MatchValue, MatchAny and Range as the everyday conditions.

for a middle

Explain that the same Filter object drives search, scroll, count and delete, and that filters are boolean-only so they never influence the similarity score.

for a senior

Show that you verify filters operationally — run count with exact=True to learn a filter's true cardinality before blaming the search for returning too few results.

for a principal

Own the convention that filter construction is centralised rather than hand-rolled per call site, since the filter is also the security boundary for anything partitioned by payload.

## What a payload is A point in Qdrant is three things: an id, one or more vectors, and a `payload` — an arbitrary JSON object such as `{"lang": "en", "views": 1200, "tags": ["ml", "db"], "location": {"lat": 51.5, "lon": -0.1}}`. The payload is what filters operate on. It is stored with the point, returned when you ask for it (`with_payload=True`), and can be indexed separately from the vector. ## Attaching a filter to a search The modern search entry point is `query_points`. It takes the query vector as `query`, the number of results as `limit`, and the payload filter as `query_filter`: ``` client.query_points(collection_name="articles", query=vec, query_filter=flt, limit=10) ``` The returned object exposes `.points`, a list of scored points with `id`, `score`, and (when requested) `payload`. The result contains only points that satisfy the filter, ordered by similarity to the query vector. ## The building blocks A `Filter` holds lists of conditions. The workhorse condition is `FieldCondition`, which names a payload key and one matching rule: - **`match=MatchValue(value=x)`** — the field equals `x`. Works for strings, integers and booleans. - **`match=MatchAny(any=[a, b])`** — the field equals any listed value; the equivalent of SQL `IN`. - **`match=MatchExcept(**{"except": [a, b]})`** — the field equals none of the listed values. - **`match=MatchText(text="...")`** — substring or token match; needs a text payload index to behave as a full-text match. - **`range=Range(gt=, gte=, lt=, lte=)`** — numeric bounds; `DatetimeRange` does the same for RFC 3339 timestamps. - **Geo conditions** — `GeoRadius(center=GeoPoint(lat=, lon=), radius=meters)`, `GeoBoundingBox(top_left=, bottom_right=)`, `GeoPolygon(...)` over a payload field holding `lat`/`lon`. Besides `FieldCondition` there are whole-point conditions: `IsEmptyCondition` (the key is missing or an empty array), `IsNullCondition` (the key holds JSON null), and `HasIdCondition(has_id=[...])` to restrict to specific point ids. ## Arrays behave as "any element" If the payload value is an array, a condition matches when **any** element matches. `tags` = `["ml", "db"]` satisfies `MatchValue(value="db")`. This is convenient and occasionally surprising when you combine two conditions over the same array — different elements may satisfy each one. ## Filters do not touch the score This trips up people arriving from search engines. In Qdrant a filter is a hard predicate: a point is eligible or it is not. There is no boost, no partial credit, no blending of a metadata match into the similarity score. If you want metadata to influence ranking you have to do it yourself after retrieval, or encode it in the vector. ## The same filter everywhere One of the nicer parts of the design is that the `Filter` type is universal. The identical object drives: - similarity search (`query_points(query_filter=...)`), - paging through the collection without a vector (`scroll(scroll_filter=...)`), - cardinality checks (`count(count_filter=..., exact=True)`), - bulk mutation — `delete(points_selector=FilterSelector(filter=...))`, `set_payload`, `delete_payload`. That makes it easy to sanity-check a filter: run `count` with it and see how many points it actually matches before wondering why a search returned three results. ## Practical notes Filters work without any payload index — correctness never depends on one — but performance and the query planner's strategy do, so any field you filter on regularly should get an index. Keys are dotted paths for nested objects (`"author.country"`) and use `field[]` syntax to reach into arrays of objects (`"items[].price"`). And ask for the payload explicitly with `with_payload=True` (or a list of keys) if you need it back; otherwise you get ids and scores only.

  • If the payload field is an array of strings, what does a MatchValue condition on it mean?
    It matches when any element of the array equals the value. Qdrant treats array payload fields as "contains" semantics for every condition type, so `tags: ["ml", "db"]` satisfies `MatchValue(value="db")`. There is no separate contains operator, and no way to require that the whole array equals a list.
  • How would you find points where a payload key is missing entirely?
    Use `models.IsEmptyCondition(is_empty=models.PayloadField(key="lang"))`, which matches points where the key is absent, holds null, or holds an empty array. If you specifically want the JSON null value and not absence, use `models.IsNullCondition` instead. A `must_not` around a `MatchValue` will not find these, because a missing key simply fails the condition.
  • Does a filter change the similarity scores of the points that survive it?
    No. Qdrant filters are strictly boolean eligibility predicates: a point either passes or is excluded, and the surviving points are ranked purely by the collection's distance metric. There is no boosting or score blending from payload matches, so any metadata-driven reranking has to happen in your own code after the query returns.

The filter is a guest list, not a scoring judge: it decides who is allowed into the room, and the vector distance alone decides who stands closest to you.

saying these in an interview costs you the question

  • Thinking filters boost ranking rather than exclude points
  • Expecting a SQL-style WHERE string instead of the Filter object
  • Assuming payload filtering requires a separate metadata store
  • Believing an array field needs a special contains operator
  • Passing the filter positionally instead of as query_filter

context