skip to content

How do you build and combine where filters in the Weaviate v4 Python client?

level: middleimportance: must knowfreq 68%

answer

  1. conditions are objects, not dicts
  2. start from Filter.by_property
  3. & and | — never and / or
  4. all_of and any_of for dynamic lists
  5. the inverted index has to know the property

basics

~10 s

Build conditions with the Filter class — Filter.by_property("category").equal("news") — and combine them with the & and | operators or with Filter.all_of([...]) and Filter.any_of([...]). Pass the result as the filters argument to any query method.

solid answer

~40 s

In the v4 client, filters are objects, not dictionaries. You start from `Filter.by_property("name")` and apply a comparison: `equal`, `not_equal`, `greater_than`, `greater_or_equal`, `less_than`, `less_or_equal`, `like` for wildcard text matching, `contains_any` / `contains_all` for array properties, and `is_none`. There are also `Filter.by_id()`, `Filter.by_creation_time()` and `Filter.by_update_time()` for object metadata. Boolean composition uses Python operators — `f1 & f2` for AND, `f1 | f2` for OR, and `~f` for negation — or the explicit `Filter.all_of([...])` and `Filter.any_of([...])` when you build conditions in a loop. The same filter object goes to `near_text`, `near_vector`, `near_object` or to `fetch_objects` when you want a pure structured query with no vector search at all. Filtering resolves against the inverted index, so the properties you filter on must be indexed for filtering in the collection's schema.

code

python · 15 lines
python
from weaviate.classes.query import Filter

articles = client.collections.get("Article")

f = (
    Filter.by_property("category").contains_any(["news", "opinion"])
    & Filter.by_property("word_count").greater_than(500)
    & ~Filter.by_property("status").equal("draft")
)

# constrained vector search
hits = articles.query.near_text(query="climate policy", limit=10, filters=f)

# same condition, no vector search at all
rows = articles.query.fetch_objects(filters=f, limit=100)

go deeper

for a junior

Know that you build a condition with Filter.by_property(name) followed by a comparison such as equal, and pass it as the filters argument to a query. Recall that & means AND and | means OR.

for a middle

Explain the full comparison vocabulary including like, contains_any and is_none, why the and / or keywords silently drop a condition, and that filters work on fetch_objects with no vector search involved.

for a senior

Show that filters resolve against the inverted index, so schema-time indexing decisions determine whether a filter is fast, and treat validating a filter with a read before reusing it for a bulk delete as standard practice.

for a principal

Own the modelling tradeoff: cross-reference filters cost a traversal per candidate and usually signal a value that should be denormalised onto the searched object, and filter shape drives which properties must be indexed — a schema decision made long before the query is written.

## Filters are objects, not dictionaries The defining change in the v4 Python client is that a filter is a typed object built with a fluent API rather than a nested dict. That matters practically: your editor completes the operator names, a typo fails at construction time instead of producing a query that silently matches nothing, and filters are ordinary Python values you can build, name, store and pass around. ```python from weaviate.classes.query import Filter articles = client.collections.get("Article") response = articles.query.near_text( query="climate policy", limit=10, filters=Filter.by_property("category").equal("news"), ) ``` ## The comparison vocabulary From `Filter.by_property("prop")` you reach: - `equal(value)` and `not_equal(value)` — exact match on text, numbers, booleans, dates. - `greater_than`, `greater_or_equal`, `less_than`, `less_or_equal` — ranges over numbers and dates. - `like(pattern)` — wildcard text matching, where `*` covers any run of characters and `?` a single one. Useful for prefix matching; expensive if you lead with a wildcard. - `contains_any([...])` and `contains_all([...])` — for array-valued properties, asking whether the array intersects your list or covers it entirely. `contains_any` is also the idiomatic way to express "category is one of these". - `is_none(True)` — matches objects where the property was never set, which is distinct from matching an empty string. Beyond properties there are metadata entry points: `Filter.by_id()` for the object UUID, and `Filter.by_creation_time()` / `Filter.by_update_time()` for timestamp ranges, which give you "only documents ingested since yesterday" without storing a timestamp property of your own. ## Composition Two equivalent styles exist and both appear in real code. Operator style reads well for fixed conditions: ```python f = ( Filter.by_property("category").equal("news") & Filter.by_property("word_count").greater_than(500) & ~Filter.by_property("status").equal("draft") ) ``` The usual Python pitfall applies: `&` and `|` are the operators, not the `and` / `or` keywords. Writing `f1 and f2` does not build a conjunction — Python evaluates truthiness and hands you one of the two objects, so the query runs with half the constraint and no error anywhere. Also remember `&` binds tighter than comparison operators, so parenthesise generously. List style suits filters assembled dynamically: ```python conditions = [Filter.by_property("category").equal(c) for c in chosen] f = Filter.any_of(conditions) ``` `Filter.all_of` and `Filter.any_of` take a list and are the natural fit when the number of conditions depends on user input. Building a list and folding it is cleaner than accumulating with `&` in a loop, and it makes the empty case explicit — an empty condition list should mean "no filter", so pass `None` rather than an empty combinator. ## Where filters can be attached The same object is accepted by every read path: - `query.near_text(..., filters=f)` and its `near_vector` / `near_object` siblings — constrained vector search. - `query.fetch_objects(filters=f, limit=100)` — a purely structured query with no vector involved. This is the right call for "show me every draft article", and reaching for a vector search with a dummy query vector instead is a common beginner mistake. - Aggregation and deletion paths accept filters too, so a bulk delete reuses exactly the condition you validated with a read. That last pattern is worth internalising as a safety habit: before a bulk delete, run the identical filter through `fetch_objects` and inspect what comes back. ## Schema is a prerequisite Filters are evaluated against Weaviate's inverted index, not by scanning objects. Properties therefore have to be indexed for filtering in the collection definition. Property-level indexing is a schema-time decision, and it is the reason a filter can fail or behave unexpectedly on a property that looks perfectly ordinary in the returned objects. Data types matter too: filtering a numeric-looking value that was declared as text with `greater_than` is a type mismatch, not a lexicographic comparison. ## Cross-reference filters When a collection has cross-references, `Filter.by_ref(link_on="hasAuthor").by_property("name").equal("Ada")` filters on a linked object's property. This is powerful and genuinely relational-feeling, but it costs an extra traversal per candidate; heavy use on a hot query path is a design smell that usually means the value should be denormalised onto the object being searched.

  • A colleague writes filters=f1 and f2 and the query ignores one condition. Why?
    Python's and keyword is not overloadable for this purpose: it evaluates truthiness and returns one operand, so the query receives a single filter object. Weaviate sees a perfectly valid query with half the constraint and reports no error. Use the & operator, or Filter.all_of([f1, f2]), which also reads better when conditions are built in a loop.
  • You need every object where a filter matches, with no similarity ranking at all. What do you call?
    collection.query.fetch_objects(filters=..., limit=...) performs a purely structured retrieval against the inverted index with no vector search. Using a near_ query with an arbitrary vector to fake this wastes an index traversal and returns results ordered by an irrelevant similarity, which is misleading to anyone reading the code later.
  • How would you express "category is one of news, sport or opinion" without three chained OR conditions?
    For a single-valued text property, Filter.by_property("category").contains_any(["news", "sport", "opinion"]) expresses the set membership directly. Filter.any_of with a list comprehension is the alternative when the conditions differ in shape rather than just in value, for example mixing equality and range conditions.

saying these in an interview costs you the question

  • Using Python's and / or keywords to combine filters
  • Passing a raw dictionary as the filters argument
  • Filtering on a property never indexed for filtering
  • Assuming like patterns are regular expressions
  • Faking a filter-only query with a dummy vector search

context