In Pinecone, which metadata value types can you store and filter on?
answer
- four value types only
- no nested objects allowed
- lists must hold strings
- omit the key instead of null
- 40 KB per vector cap
basics
~20 sPinecone metadata accepts strings, numbers, booleans, and lists of strings. Nested objects and null values are rejected, so flatten your structure and simply omit fields that have no value. The whole metadata payload is capped at 40 KB per vector.
solid answer
~50 sPinecone stores metadata as a **flat** map attached to each vector record, and only four value types are allowed: string, number, boolean, and list of strings. Anything else — a nested object, a `null`, a list of numbers — is a validation error at upsert time, so you flatten (`author.name` becomes `author_name`) and you omit a key rather than writing `null` for it. The type you choose determines which operators you can use: `$gt`/`$gte`/`$lt`/`$lte` only compare numbers, so dates must be stored as numeric epoch timestamps if you ever want range filters on them; strings and booleans support `$eq`/`$ne`, and `$in`/`$nin` do set membership. A list of strings behaves like a tag set — a match on any element matches the vector. Metadata is capped at 40 KB per vector, so keep raw document text out of it unless you genuinely need it returned.
go deeper
Be ready to name the four supported types — string, number, boolean, list of strings — and to say plainly that dates go in as numeric timestamps.
Explain why the type choice constrains the operators available, and show how you would flatten a nested payload and handle absent fields with $exists rather than null.
Show judgment about what belongs in metadata at all: filterable fields stay small and typed, display fields stay minimal, and bulk content lives in your own store keyed by vector id.
Own the schema contract across ingestion pipelines — a shared, versioned metadata schema, a rule for how new fields are backfilled, and a size budget per record so response payloads and storage stay predictable as the corpus grows.
## What metadata is in Pinecone A Pinecone record is three things: an `id` string, a `values` array (the embedding), and an optional `metadata` map. The metadata is what turns a pure nearest-neighbour store into something you can query with business constraints — tenant, document type, language, freshness, visibility. It is stored next to the vector, returned when you ask for it, and — crucially — usable as a filter on similarity queries. The important design constraint is that this map is **flat and typed**. It is not a general-purpose JSON document store. ## The four supported types - **String** — `{"doc_type": "faq"}`. The workhorse for enums and categorical fields. - **Number** — `{"published_at": 1735689600}` or `{"score": 0.83}`. Integers and floats are both numbers; there is no separate integer type. - **Boolean** — `{"is_public": true}`. - **List of strings** — `{"tags": ["billing", "eu"]}`. A tag set. That is the whole list. There is no date type, no list of numbers, no list of booleans, no nested object, and `null` is not a value you can store. ## What is rejected, and what to do instead A nested object such as `{"author": {"name": "Ada", "team": "search"}}` fails validation. Flatten it into `author_name` and `author_team`. If you truly need the structure and never filter on it, serialise it to a JSON **string** — but remember that string still counts against the size budget. A `null` value is likewise rejected. The idiom for "this record has no category" is to leave the key out entirely, and then use the `$exists` operator to distinguish records that have the field from those that do not. A list of numbers — `{"years": [2023, 2024]}` — is not supported. If you need set membership over numbers, either store them as strings (losing range comparisons) or model the dimension differently, for example as separate boolean or numeric fields. ## Type determines which operators you can use This is the part that actually bites in interviews and in production: - `$eq` / `$ne` work on strings, numbers, and booleans. - `$gt`, `$gte`, `$lt`, `$lte` compare **numbers only**. They do not do lexicographic comparison on strings. - `$in` / `$nin` test membership in a list of candidate values. - `$exists` tests whether the key is present on the record at all. The headline consequence: **store dates as numbers.** An ISO-8601 string like `"2026-01-15"` looks sortable to a human, but Pinecone will not range-compare it. Store a Unix epoch timestamp as a number for filtering; if you also want to display a human-readable date, store it as a second string field that you never filter on. ## Lists of strings behave like sets With `{"tags": ["billing", "eu"]}`, a filter that matches one of the tags matches the record — membership semantics rather than whole-value equality. This makes lists of strings the natural encoding for many-to-one attributes: labels, applicable regions, allowed roles. Keep the lists short; every element is stored per vector and counts toward the size cap. ## Size, cost, and what belongs in metadata Metadata is limited to 40 KB per vector. Even well under that ceiling, metadata is not free: it is stored per record, and it is transferred back on every query that asks for it. Two rules of thumb: 1. **Filterable fields**: small, low-cardinality where possible, typed for the operator you need. 2. **Display fields**: the minimum you need to render a result — a title, a source URL, a chunk offset. Full document text usually belongs in your own store, keyed by the vector id, rather than inside metadata. ## Common mistakes The recurring ones are storing dates as strings and then wondering why the range filter returns nothing useful; writing `null` for absent values and getting upsert errors; trying to push a nested JSON document in wholesale; and treating metadata as the system of record for content, which inflates every query response. Decide field by field: is this something I filter on, something I display, or something that belongs in the primary datastore?
- What happens if you upsert a vector whose metadata contains a nested object?The upsert is rejected with a validation error — Pinecone accepts only flat key/value pairs of the four supported types. The fix is to flatten the structure, so `author.name` becomes an `author_name` key, or to serialise the object into a single string you never filter on. A serialised blob still counts toward the 40 KB per-vector metadata limit, so keep it small.
- How would you store a date so that range filters work on it?As a number — a Unix epoch timestamp — because the range operators `$gt`, `$gte`, `$lt` and `$lte` compare numbers only. An ISO-8601 string supports equality and set membership but not ranges. A common pattern is to store both: a numeric timestamp for filtering and a formatted string for display, accepting the small extra bytes per record.
- How do you represent an absent value, given that null is not allowed?Leave the key out of the metadata map entirely. Then `{"category": {"$exists": true}}` selects records that carry the field and `{"$exists": false}` selects those that do not. This is also useful during backfills, when a newly added field exists on freshly written records but not on older ones.
saying these in an interview costs you the question
- Claims you can store nested JSON objects directly in metadata
- Uses null in metadata to represent a missing value
- Assumes range operators compare ISO date strings correctly
- Thinks lists of numbers or booleans are supported metadata values
- Treats metadata as unlimited document storage for full chunk text