What does Pinecone's selective metadata indexing buy you on a pod-based index?
answer
- you choose which fields are filterable
- filter indexes share pod memory
- cost scales with field cardinality
- unindexed still stored and returned
- fixed at index creation time
basics
~20 sSelective metadata indexing builds filter indexes only for the fields you name in metadata_config when creating a pod-based index. That keeps high-cardinality fields out of pod memory so more vectors fit. Unlisted fields are still stored and returned — just not filterable.
solid answer
~50 sOn a pod-based index you can pass `metadata_config={"indexed": [...]}` to `create_index`, listing the fields that should be filterable. Only those fields get a filter index; everything else is still stored with the record and still returned when you request metadata, but a filter referencing it will not work. The reason to bother is memory economics: metadata indexes live in the same pod memory as your vectors, so a high-cardinality field — a per-chunk id, a source URL, a per-record hash — can consume a meaningful share of the pod and reduce how many vectors you fit before scaling out. Naming three low-cardinality fields instead of indexing everything is a direct capacity win. The cost is rigidity: the configuration is fixed at index creation, so adding a newly filterable field means creating a new index and re-upserting. Serverless indexes do not expose this knob — there you shape filterability through your data model instead.
code
python · 14 linesfrom pinecone import Pinecone, PodSpec
pc = Pinecone(api_key="pk-example")
pc.create_index(
name="docs",
dimension=1536,
metric="cosine",
spec=PodSpec(
environment="us-east-1-aws",
pod_type="p1.x1",
metadata_config={"indexed": ["doc_type", "tenant_id", "published_at"]},
),
)go deeper
Know that on pod-based indexes you can choose which metadata fields are filterable, and that unlisted fields are still stored and returned.
Explain the memory argument: filter indexes share pod memory with vectors, and cost scales with a field's cardinality, so indexing everything shrinks capacity.
Demonstrate the operational judgment — pick indexed fields from real query patterns, keep high-cardinality display fields out, and know that widening the list means a new index and a re-upsert.
Own the capacity model: tie the indexed-field decision to pod sizing and growth forecasts, and document the reasoning, because an immutable configuration made casually becomes a migration later.
## The mechanism When you create a pod-based Pinecone index you can pass a `metadata_config` object with an `indexed` list naming the metadata fields that should be filterable. Only the named fields get a filter index built for them. Everything else about the record is unchanged — the metadata is still stored, still travels back on queries that request it, and still costs storage. What it loses is the ability to appear in a filter expression. If you omit `metadata_config` entirely, all metadata fields are indexed. Passing an empty `indexed` list is the opposite extreme: nothing is filterable, which is a legitimate choice for an index used purely for display metadata. ## Why the knob exists: pod memory economics A pod has a fixed memory budget, and it holds both the vector index and the metadata filter indexes. Every additional indexed field takes a share of that budget, and the size of that share scales with **cardinality** — the number of distinct values the field takes. - `is_public` (2 distinct values) is nearly free. - `doc_type` (a dozen values) is cheap. - `tenant_id` (a few hundred values) is real but usually worth it. - `chunk_id`, unique per vector, is the pathological case: its index is as large as your record count and buys you nothing, because filtering by a unique id is what fetching by vector id is for. The practical effect of indexing a high-cardinality field is that fewer vectors fit per pod, so you scale to more pods sooner and pay for capacity that is holding an index nobody queries. Selective metadata indexing is how you avoid that. ## The tradeoff you are accepting The configuration is set when the index is created. You cannot decide six months later that `language` should have been filterable and flip a switch — you create a new index with the wider `indexed` list, re-upsert, and cut over. That makes this a decision worth taking deliberately rather than by default. The pragmatic middle ground: index the fields you filter on today plus the one or two you can clearly foresee, and leave genuinely display-only fields (titles, URLs, snippets, source paths) out. Those are exactly the fields that tend to be high-cardinality, so the memory win and the "never filtered anyway" property usually coincide. ## What happens when you filter an unindexed field This is the failure mode people hit: the field is visibly present in the returned metadata, so it looks like it should be filterable, but the filter does not behave as expected. Debugging it means going back to the index configuration rather than staring at the data. When someone reports "the filter isn't working but I can see the value right there in the response", check `metadata_config` first. ## Serverless is a different world Serverless indexes do not expose selective metadata indexing — you do not configure pods, so there is no pod memory budget for you to protect. That does not make the underlying discipline irrelevant: on serverless, a high-cardinality filter field still shapes query cost and latency, because a filter that matches a tiny, widely scattered slice of the index forces more of it to be read. The lever changes from "which fields are indexed" to "how the data is laid out and how selective the filters are". ## How to decide, field by field For each metadata field, ask three questions: 1. **Does any query filter on it?** If no, it does not need indexing — only storage. 2. **How many distinct values does it take?** Low cardinality is cheap; per-record uniqueness is a red flag. 3. **Is it a per-query constraint or a dataset partition?** Constraints that vary per query are what filters are for. A dimension that permanently splits your dataset into disjoint groups is usually better handled by the index layout than by paying for a filter index on every record. Write that decision down next to the schema. Because the configuration is immutable, the reasoning behind it is the thing future engineers most need and are least likely to reconstruct. ## Diagnosis in production Symptoms that point back here: pods filling up far faster than vector count alone predicts; a filter that silently fails on a field you can see in the response; a plan to add a new filterable dimension that turns out to require a full re-index. All three are cheaper to prevent at creation time than to fix later, which is why this comes up as a design-review question rather than a coding one.
- If a field is not in the indexed list, can you still see its value in query results?Yes. Selective indexing controls filterability only. The value is stored with the record and comes back when the query asks for metadata, so titles, URLs and snippets work exactly as before. This is also why the failure is confusing in practice: the data is visibly present, but a filter naming that field will not select on it.
- How would you add a newly filterable field to an index that was created without it?Create a new index with the widened `indexed` list, re-upsert the vectors into it, and cut traffic over — the configuration is fixed at creation. Because re-embedding is usually unnecessary (the vectors themselves are unchanged), the cost is mostly a bulk re-upsert and a dual-write or read-switch window rather than a full rebuild of your embeddings.
- Why is a per-record unique id a bad choice to index for filtering?Its filter index is as large as the record count while every lookup on it returns at most one vector — which is what fetching by vector id already does, more cheaply. You spend pod memory that would otherwise hold vectors, reducing how many records fit per pod and pushing you to scale out sooner for no query benefit.
saying these in an interview costs you the question
- Thinks unindexed metadata is dropped and not returned in results
- Assumes metadata_config also tunes serverless index behaviour
- Indexes a per-record unique id so it can be filtered on
- Believes the indexed field list can be changed on a live index
- Confuses the metadata filter index with the vector index itself