skip to content

Pinecone

The managed vector database most teams reach for first: serverless indexes, namespaces for tenancy, metadata filtering, and hybrid sparse-dense search, with index internals deliberately hidden. Interviews focus on data modelling and cost rather than on operating the engine.

on this pageshow

explore

questions

20

In Pinecone, what do the dimension and metric arguments to create_index lock in?

level: juniorimportance: must knowfreq 78%

answer

  1. Set once at creation time
  2. Vector length must match exactly
  3. Mismatch is rejected, not padded
  4. Embedding model change means new index
  5. configure_index cannot alter geometry

basics

~20 s

create_index fixes an index's vector dimension and distance metric permanently. Every vector you upsert must have exactly that length or the write is rejected, and switching either setting means creating a new index and re-upserting the data.

solid answer

~50 s

`pc.create_index(name=..., dimension=1536, metric="cosine", spec=ServerlessSpec(...))` declares the shape of everything the index will ever hold. `dimension` is the fixed length of each vector, so it must match the output size of the embedding model you plan to use; upserting a vector of any other length fails with a dimension-mismatch error rather than being padded or truncated. `metric` is one of `cosine`, `dotproduct` or `euclidean`, and it is the distance function the index is built and queried with. Neither is mutable afterwards — `configure_index` can change things like deletion protection or pod scaling, but not dimension or metric. In practice that makes the index tightly coupled to one embedding model: changing models with a different output size forces a new index and a full re-embed, so teams usually create the new index, backfill it, then cut traffic over.

code

python · 14 lines
python
from pinecone import Pinecone, ServerlessSpec

pc = Pinecone(api_key="YOUR_API_KEY")

pc.create_index(
    name="docs",
    dimension=1536,
    metric="cosine",
    spec=ServerlessSpec(cloud="aws", region="us-east-1"),
    deletion_protection="enabled",
)

index = pc.Index("docs")
print(index.describe_index_stats())

go deeper

for a junior

Know that create_index takes a name, dimension, metric and spec, that dimension must equal your embedding model's output length, and that a wrong-length vector is rejected outright.

for a middle

Explain why the metric is baked into the index rather than chosen per query, and how score is read differently for euclidean than for cosine or dotproduct.

for a senior

Show the migration you would run for a model swap: parallel index, backfill from a store of record, relevance comparison, config-level cutover, rollback window before deleting the old index.

for a principal

Own the coupling this creates between model choice and infrastructure — budget for periodic re-embedding, keep source text in a system of record, and make index identity a configuration value so model upgrades never require a code release.

## What an index is In Pinecone an **index** is the top-level container you create, size and pay for. It holds records, each of which is an id, a dense vector of floating-point numbers, and optional metadata. Everything else — namespaces, queries, filters — happens inside an index. Creating one is a control-plane call on the client: The modern SDK is instantiated as `pc = Pinecone(api_key="...")` and the index is created with `pc.create_index(name="docs", dimension=1536, metric="cosine", spec=ServerlessSpec(cloud="aws", region="us-east-1"))`. Three of those arguments are structural and one is a label: the `name` identifies the index, `spec` chooses the deployment shape, and `dimension` plus `metric` fix the mathematics the index is built around. ## dimension An embedding model emits a vector of a fixed length — 384, 768, 1024, 1536, 3072 and so on depending on the model. `dimension` tells Pinecone that length, and it is enforced on every write. If the index says 1536 and you upsert a 768-float vector, the API rejects the request with a dimension-mismatch error. It does **not** pad with zeros, truncate, or project the vector; there is no coercion at all, because a vector of a different length is simply not a point in the same space and any distance computed against it would be meaningless. The practical consequence is that an index is bound to one embedding model family. This is the single most common beginner mistake in Pinecone: someone builds an index against one model, later swaps the model for a better one with a different output size, and discovers that the existing index cannot accept the new vectors at all. Even when two models happen to share a dimension the situation is not much better — the numbers are the same length but live in unrelated geometries, so mixing them in one index produces silently poor results rather than an error. The rule of thumb is: one index (or at least one namespace) per embedding model. ## metric `metric` selects the distance function: `"cosine"`, `"dotproduct"` or `"euclidean"`. Which of those suits a given embedding model is embedding theory, but two Pinecone-specific facts matter here. First, the metric is used both when the index structures data and when it scores a query, so it is not a per-query option — you cannot ask a cosine index for a Euclidean ranking. Second, the score in a query response is interpreted according to the metric: for `cosine` and `dotproduct` higher is more similar, for `euclidean` the value is a distance and lower is more similar. Code that sorts or thresholds on `score` without knowing the index's metric is a latent bug. ## Immutability and the migration pattern Neither `dimension` nor `metric` can be altered after creation. `pc.configure_index(...)` exists, but it changes operational settings — deletion protection, index tags, and for pod-based indexes the replica count or pod type — never the vector geometry. There is no ALTER INDEX equivalent. So a model change is a migration, and the standard shape is: 1. Create a second index with the new `dimension` and `metric`. 2. Re-embed the source documents with the new model and upsert into the new index. Pinecone does not store your original text unless you put it in metadata, which is one reason people keep the raw chunk text in metadata or in an external store of record. 3. Verify counts with `describe_index_stats()` and run offline relevance checks against the new index. 4. Flip the application's index name (usually a config value) and delete the old index once you are confident. Because the cutover is a config change, the usual approach is to keep the index name out of the code and in configuration or environment so the swap requires no redeploy of logic. ## Other creation-time choices Two more arguments deserve a mention. `spec` chooses `ServerlessSpec` (a cloud and region, with capacity managed for you) or `PodSpec` (an environment and explicit pod sizing) and cannot be converted in place either — moving between them is also a create-and-backfill migration. `deletion_protection="enabled"` can be set at creation or later via `configure_index`, and it blocks `delete_index` until it is turned off, which is a cheap guard for a production index that took hours to build. ## What interviewers are checking The question looks trivial but it separates people who have run Pinecone from people who have read the quickstart. The signal is: do you know the index is immutable in its two most important properties, do you know a dimension mismatch is a hard error rather than a silent coercion, and can you describe the re-embed-and-cut-over migration without hesitating.

  • You need to move from a 768-dimension model to a 1536-dimension one with no search downtime. What's your plan?
    Create a second index at dimension 1536, re-embed and backfill every document into it while the old index keeps serving, then compare relevance on a held-out query set. When the new index passes, flip the index name in configuration so traffic cuts over atomically, keep the old index for a rollback window, and delete it once you're confident. Storing chunk text in metadata or an external store of record is what makes the re-embed possible.
  • Two embedding models both output 1536 floats. Is it safe to keep their vectors in one index?
    No. The index will accept both because the dimension matches, so you get no error — but the two models place meaning in different directions, so distances between vectors from different models are noise. Queries will return plausible-looking but wrong neighbours. Keep one model per index, or at minimum per namespace, and never mix them in a searchable partition.
  • Which index properties can configure_index actually change?
    Operational ones only: deletion protection, index tags, and for pod-based indexes the replica count and pod type for vertical or horizontal scaling. It cannot change dimension, metric, or the spec type. Anything that changes the vector geometry or deployment shape requires creating a new index and backfilling it.

saying these in an interview costs you the question

  • Thinks Pinecone pads or truncates mismatched vectors
  • Believes dimension can be changed on an existing index
  • Says the metric can be chosen per query
  • Mixes two embedding models in one index because dimensions match
  • Confuses configure_index with altering vector geometry

context

open as a page

What does a Pinecone vector record contain, and how should you batch upserts?

level: juniorimportance: must knowfreq 82%

basics

~20 s

A Pinecone record is an id string, a values list whose length must equal the index dimension, and optional metadata. index.upsert() takes a list of records, so send batches of roughly 100 to stay under the request size cap.

open as a page

Why must a Pinecone index use the dotproduct metric for hybrid sparse-dense search?

level: middleimportance: must knowfreq 62%

basics

~20 s

Pinecone accepts sparse values only in indexes created with metric="dotproduct". Dot product is linear, so scaling the dense and sparse query vectors weights their contributions predictably; cosine and euclidean indexes reject sparse vectors, and the metric cannot be changed later.

open as a page

In Pinecone, how do you upsert and query a record holding both dense and sparse vectors?

level: middleimportance: must knowfreq 55%

basics

~20 s

A hybrid record carries the dense embedding in values and a sparse vector in sparse_values as {"indices": [...], "values": [...]}. The query passes vector= and sparse_vector= together, and Pinecone returns one ranked list with a single combined score per match.

open as a page

What is a Pinecone namespace, and how does it change query behaviour?

level: middleimportance: must knowfreq 70%

basics

~20 s

A namespace is a logical partition inside a Pinecone index. Each upsert and each query names one namespace, and a query searches only that partition — there is no cross-namespace search in a single call. Namespaces are created implicitly on first write.

open as a page

In a Pinecone metadata filter, how do you combine a category match with a numeric range?

level: middleimportance: must knowfreq 70%

basics

~20 s

Write one clause per field and join them with $and — or list both keys at the top level, which Pinecone treats as an implicit AND. Use $in for the category and $gte plus $lte on a numeric timestamp field for the range.

open as a page

In Pinecone, what does index.query() return and what caps apply to top_k?

level: middleimportance: must knowfreq 74%

basics

~20 s

index.query() returns a list of matches ordered best-first, each with an id and a score; metadata and the raw vector come back only if you ask via include_metadata or include_values. top_k tops out near 10,000, and near 1,000 once metadata or values are included.

open as a page

Why can a vector just upserted to Pinecone be missing from query results?

level: middleimportance: must knowfreq 63%

basics

~20 s

Pinecone is eventually consistent: upsert() returns once the write is accepted and durable, not once it is searchable, so there is a short window where a query or fetch will not see it. Deletes and updates have the same lag.

open as a page

In Pinecone, which metadata value types can you store and filter on?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Pinecone metadata accepts strings, numbers, booleans, and lists of strings. Nested objects and null values are rejected, so flatten your structure and simply omit fields that have no value. The whole metadata payload is capped at 40 KB per vector.

open as a page

What does Pinecone's describe_index_stats() return, and when do you call it?

level: middleimportance: should knowfreq 48%

basics

~20 s

describe_index_stats() reports the index's vector dimension, the total record count, an index_fullness figure, and a per-namespace map of vector counts. It is the standard check that a load landed in the right namespace and that writes have become visible.

open as a page

How do Pinecone's upsert() and update() differ when changing one metadata field?

level: middleimportance: should knowfreq 52%

basics

~20 s

upsert() replaces the whole record, so any metadata you do not resend is lost and the embedding must be sent again. update(id=..., set_metadata={...}) patches just the named fields in place, leaves the rest alone, and never creates a record that does not already exist.

open as a page

In a Pinecone hybrid query, how do you shift the weighting between dense and sparse?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Pinecone has no server-side alpha parameter. You scale the query vectors yourself before sending: multiply the dense query by alpha and the sparse query weights by (1 - alpha). Because the metric is dot product, the returned score becomes a convex blend of the two.

open as a page

In Pinecone hybrid search, why must index-time and query-time sparse vectors share one fitted encoder?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Pinecone treats sparse indices as opaque integers and never validates them. If the encoder's vocabulary or corpus statistics change between indexing and querying, the term ids no longer line up, so the lexical half scores near zero — silently, with no error and no failed query.

open as a page

When creating a Pinecone index, what do ServerlessSpec and PodSpec trade off?

level: seniorimportance: should knowfreq 55%

basics

~20 s

ServerlessSpec hands capacity management to Pinecone: you pick a cloud and region, storage grows with your data, and you pay for stored data plus read and write work. PodSpec makes you size and pay for fixed pods that you must scale yourself.

open as a page

Why can a very selective Pinecone metadata filter make a query slower, not faster?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Pinecone evaluates a metadata filter during the approximate search rather than after it. When only a tiny slice of the index qualifies, the engine must examine far more candidates before it accumulates top_k qualifying neighbours, so work and latency go up even though fewer records match.

open as a page

What does Pinecone's selective metadata indexing buy you on a pod-based index?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Selective metadata indexing builds filter indexes only for the fields you name in metadata_config when creating a pod-based index. That keeps high-cardinality fields out of pod memory so more vectors fit. Unlisted fields are still stored and returned — just not filterable.

open as a page

How do you load ten million vectors into a Pinecone index efficiently?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Batch roughly 100 records per upsert to stay under the request size cap, send batches concurrently using an Index configured with pool_threads and upsert(async_req=True), retry 429 responses with exponential backoff, and reconcile counts with describe_index_stats after writes stop.

open as a page

How do you delete every vector for one document from a Pinecone serverless index?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Serverless indexes do not support deleting by metadata filter, so you must delete by id. Give a document's chunks a shared id prefix at ingest time, page those ids with index.list(), and pass them to index.delete(ids=[...]) in batches within the right namespace.

open as a page

When do you add a Pinecone reranking stage rather than keep tuning hybrid alpha?

level: principalimportance: should knowfreq 30%

basics

~20 s

Tune alpha while the right documents are already in the top-k but ordered badly by a linear blend. Add reranking when ordering needs the query and document read together — it costs latency and money per query, and it can only reorder what retrieval already returned.

open as a page

In Pinecone, how do you isolate thousands of tenants — namespace or index per tenant?

level: principalimportance: should knowfreq 42%

basics

~20 s

Namespace per tenant is the default at that scale: namespaces are created implicitly on write, cost nothing to mint, and give a hard query boundary. Index per tenant is justified only when tenants need different dimensions, metrics, regions or billing separation.

open as a page