skip to content

Metadata Filtering

Metadata filters restrict the candidate set of a similarity query — for tenant scoping, date windows, or document types. The interesting part is what filtering does to recall and latency when the filter is very selective.

on this pageshow

questions

4

In a Pinecone metadata filter, how do you combine a category match with a numeric range?

level: middleimportance: must knowfreq 70%

answer

  1. small MongoDB-style operator set
  2. top-level keys are conjoined
  3. two operators, one field, one range
  4. $or must always be explicit
  5. ranges need a numeric field

basics

~20 s

Write one clause per field and join them with $and — or list both keys at the top level, which Pinecone treats as an implicit AND. Use $in for the category and $gte plus $lte on a numeric timestamp field for the range.

solid answer

~50 s

Pinecone filters are a small MongoDB-style expression language: `$eq`, `$ne`, `$gt`, `$gte`, `$lt`, `$lte`, `$in`, `$nin`, `$exists`, plus `$and` and `$or` for composition. For "documents of type policy or faq, published in 2026", you write `{"$and": [{"doc_type": {"$in": ["policy", "faq"]}}, {"published_at": {"$gte": 1735689600, "$lte": 1767225600}}]}`. Two shorthands matter: multiple keys in the same object are ANDed implicitly, so the outer `$and` is optional here; and multiple operators on the *same* field are also ANDed, which is exactly how you express a closed range with one `published_at` clause. `$or` is never implicit — you must spell it out. The range operators only work because `published_at` is a **number**; the same window expressed against ISO date strings would not compare. Filters are evaluated within the namespace the query targets, so namespace scoping and metadata filtering compose rather than compete.

code

python · 16 lines
python
from pinecone import Pinecone

pc = Pinecone(api_key="pk-example")
index = pc.Index("docs")

results = index.query(
    vector=[0.02] * 1536,
    top_k=10,
    include_metadata=True,
    filter={
        "$and": [
            {"doc_type": {"$in": ["policy", "faq"]}},
            {"published_at": {"$gte": 1735689600, "$lte": 1767225600}},
        ]
    },
)

go deeper

for a junior

Know the operator names and that listing several keys means all of them must match. Be able to write a simple two-condition filter without hesitating.

for a middle

Explain the three places conjunction is implicit, why $or must be explicit, and why range filters require the field to be stored as a number.

for a senior

Show how you generate filters programmatically, log them with slow queries, and choose positive membership over negation so new data cannot silently widen a filter.

for a principal

Own the boundary between what is expressed as a filter and what is expressed structurally in the index layout — per-query constraints belong in filters, dataset-wide partitions usually do not.

## The operator vocabulary Pinecone's filter language is deliberately small and will look familiar to anyone who has written a MongoDB query: - **Equality**: `$eq`, `$ne` - **Ranges** (numbers only): `$gt`, `$gte`, `$lt`, `$lte` - **Set membership**: `$in`, `$nin` - **Presence**: `$exists` - **Composition**: `$and`, `$or` A bare value is shorthand for `$eq`, so `{"doc_type": "faq"}` and `{"doc_type": {"$eq": "faq"}}` mean the same thing. ## Three levels of implicit AND This is where most confusion lives, and it is worth being precise, because the same conjunction shows up in three places: 1. **Multiple keys in one object are ANDed.** `{"doc_type": "faq", "is_public": true}` requires both. 2. **Multiple operators on one field are ANDed.** `{"published_at": {"$gte": 1735689600, "$lte": 1767225600}}` is a single closed interval — the idiomatic way to write a window, rather than two separate clauses. 3. **`$and` makes the conjunction explicit** and is needed when you want to nest, for example when one branch of an `$or` is itself a conjunction. Disjunction is never implicit. If you want "tier is enterprise **or** the document is public", you must write `{"$or": [{"tier": "enterprise"}, {"is_public": true}]}`. ## Putting the example together "Policies and FAQs published in calendar 2026" becomes: - a set-membership clause on the categorical field: `{"doc_type": {"$in": ["policy", "faq"]}}` - a closed numeric interval on the timestamp: `{"published_at": {"$gte": 1735689600, "$lte": 1767225600}}` - joined with `$and` (or simply placed as two top-level keys) The filter travels with the similarity query and constrains which vectors are eligible to be returned as neighbours. ## Why the range field must be a number The range operators compare numbers. They do not lexicographically compare strings, so a `published_at` stored as `"2026-01-15"` cannot be windowed. The standard encoding is a Unix epoch timestamp stored as a number. Teams that want a readable date as well store it separately as a display-only string. Getting this wrong is one of the most common causes of "my filter silently returns nothing sensible". ## $exists and heterogeneous corpora In practice several ingestion pipelines write into the same index, and they rarely agree on every field. `$exists` is the escape hatch: `{"language": {"$exists": false}}` finds records a newer pipeline has not yet backfilled, and `{"language": {"$exists": true}}` restricts a query to the enriched subset. It is also the correct way to express "missing", because `null` is not a storable metadata value. ## $ne and $nin are not free Negation is expressive but works against you operationally: a filter that excludes a handful of values still qualifies most of the index, while a filter that excludes almost everything is extremely selective. Prefer stating what you want (`$in`) over what you do not want (`$nin`) when both are available — the positive form is easier to reason about and usually easier for the engine to satisfy. ## Composition with namespaces A query targets a single namespace, and the metadata filter is applied within that namespace. The two mechanisms stack: a namespace narrows the search space up front, and the filter expresses per-query business constraints inside it. Because of this, a dimension that partitions your whole dataset is usually not the best candidate for a metadata filter — filters are best at expressing the constraints that vary from query to query. ## Keeping filters maintainable As filters grow, build them programmatically rather than by hand-writing nested dictionaries at each call site: a small helper that takes the user's constraints and emits the filter object keeps the operator semantics in one place and makes the conjunction rules explicit. Log the generated filter alongside slow queries — when result quality degrades, the filter is very often the thing that changed.

  • When do you actually need $and rather than relying on the implicit conjunction?
    Whenever you cannot express the condition as distinct top-level keys — most often when nesting inside `$or`, or when two conditions target the same field in a way one object cannot hold. Many teams also write `$and` unconditionally in generated filters, because a uniform shape is easier to build and to log than one that sometimes collapses to bare keys.
  • What does $exists buy you when several pipelines write into the same index?
    It lets you tell enriched records from un-enriched ones without a sentinel value, which matters because Pinecone does not accept null. During a backfill you can serve queries with `{"field": {"$exists": true}}` so users only see fully-enriched results, then drop the clause once the backfill completes. It is also how you audit how far a migration has progressed.
  • Why prefer $in over $nin when both would express the same intent?
    The positive form states exactly which values qualify, which is easier to reason about and to test, and it does not silently start matching new values that appear later in the corpus. A `$nin` filter admits every category you have not thought to exclude, so a new document type added by another team quietly enters your results.

saying these in an interview costs you the question

  • Assumes multiple top-level keys are ORed rather than ANDed
  • Expects $or to apply implicitly without writing it
  • Uses range operators against ISO date strings
  • Writes two separate clauses instead of one field with $gte and $lte
  • Believes a filter escapes the namespace the query targets

context

open as a page

In Pinecone, which metadata value types can you store and filter on?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Pinecone metadata accepts strings, numbers, booleans, and lists of strings. Nested objects and null values are rejected, so flatten your structure and simply omit fields that have no value. The whole metadata payload is capped at 40 KB per vector.

open as a page

Why can a very selective Pinecone metadata filter make a query slower, not faster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Pinecone evaluates a metadata filter during the approximate search rather than after it. When only a tiny slice of the index qualifies, the engine must examine far more candidates before it accumulates top_k qualifying neighbours, so work and latency go up even though fewer records match.

open as a page

What does Pinecone's selective metadata indexing buy you on a pod-based index?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Selective metadata indexing builds filter indexes only for the fields you name in metadata_config when creating a pod-based index. That keeps high-cardinality fields out of pod memory so more vectors fit. Unlisted fields are still stored and returned — just not filterable.

open as a page