In Pinecone, what does index.query() return and what caps apply to top_k?
answer
- Two fields always, two on request
- Score direction is not universal
- Payload flags shrink the ceiling
- Roughly ten thousand, or roughly one thousand
- No offset, no cursor, ever
basics
~20 sindex.query() returns a list of matches ordered best-first, each with an id and a score; metadata and the raw vector come back only if you ask via include_metadata or include_values. top_k tops out near 10,000, and near 1,000 once metadata or values are included.
solid answer
~40 sYou call `index.query(vector=[...], top_k=10, namespace="...", filter={...}, include_metadata=True)`. The response carries `matches`, a list ordered from best to worst, where each match has an `id` and a `score`; `metadata` and `values` are present only when you explicitly request them with `include_metadata` / `include_values`. How to read `score` depends on the index's metric — for cosine and dot product higher is better, for euclidean it is a distance where lower is better — so never hard-code a threshold without knowing the metric. `top_k` is capped: around 10,000 for bare id-plus-score results, but roughly 1,000 once you include metadata or values, because payload size is the real constraint. There is no offset or cursor on query, so you cannot page deeper by asking for results 100–200; you widen top_k or narrow with a filter instead.
code
python · 10 linesres = index.query(
vector=[0.1] * 1536,
top_k=20,
namespace="tenant-a",
include_metadata=True,
include_values=False,
)
for m in res.matches:
print(m.id, m.score, m.metadata.get("doc"))go deeper
Know the call shape — vector, top_k, and the include flags — and that matches come back ordered best-first with an id and a score. Say that metadata only appears if you ask for it.
Explain that score meaning follows the index metric, that the include flags govern payload not search, and that top_k's ceiling drops from around 10,000 to around 1,000 once metadata or values are included.
Demonstrate request-path judgment: include_values off, lean metadata, a top_k sized for a reranker rather than for a UI, and thresholds tuned against labelled data instead of hard-coded per model.
Own the retrieval contract. Decide how relevance cutoffs are set and re-validated across embedding-model changes, and rule out designs that need deep pagination over an approximate ranking.
## The call A query is a single request that carries the query vector and the shape of the answer you want: ``` res = index.query( vector=[...], top_k=10, namespace="tenant-a", filter={...}, include_metadata=True, include_values=False, ) ``` You can also query by an id already in the index (`index.query(id="doc-1#chunk-0", top_k=10)`), which uses that stored vector as the query — handy for more-like-this features, and it saves you a round trip to the embedding model. One query targets one namespace. If you pass none, you query the default namespace, which is a frequent cause of "my query returns nothing" when the writes went to a named namespace. ## The response `res.matches` is a list ordered best-first. Each match always carries: - **`id`** — the record key you supplied at upsert time. - **`score`** — the similarity or distance under the index's metric. And carries, only when asked: - **`metadata`** — the stored dictionary, when `include_metadata=True`. - **`values`** — the full embedding, when `include_values=True`. The response also reports the namespace it searched, and on serverless indexes a usage figure for the read, which is what your bill is computed from. ## Reading the score correctly The score is not a universal 0-to-1 relevance number. Its meaning follows the metric the index was created with: for cosine and dot product a **larger** score means more similar; for euclidean the value is a **distance**, so smaller is better and sorting is ascending. A retrieval layer that assumes "higher is better" and then gets pointed at a euclidean index quietly returns the worst matches first while looking perfectly healthy. Any hard threshold you apply — "drop results below 0.75" — is metric-specific and, worse, embedding-model-specific: the same cutoff means different things after a model swap. Prefer relative cutoffs (a gap from the top score) or tune the threshold against a labelled set. ## What include_values and include_metadata really cost These flags do not change what is searched — they change what is shipped back. Setting `include_values=True` on a 1536-dimension index means every match drags a full embedding through the network and into your process memory. For a RAG pipeline you almost never need the vectors: you need ids and enough metadata to fetch the source text. Turning `include_values` off is one of the cheapest latency and cost wins available, and on serverless it directly reduces the read usage the query is billed for. `include_metadata=True` is the common case, because the whole point of metadata is to travel with the result. But it is also the reason to keep metadata lean: fat metadata multiplies by `top_k` on every single query. ## The top_k ceilings Two different limits, which is what makes this a good interview question: - Bare results (id and score only): `top_k` can go up to about **10,000**. - With `include_metadata=True` or `include_values=True`: `top_k` is capped near **1,000**. The reason is response size — the ceiling exists to bound the payload, so asking for the payload lowers the ceiling. If a request fails with a top_k error, the fix is to drop the include flags and re-hydrate the ids you actually need with `fetch`, or to accept a smaller `top_k`. ## No pagination Query has no `offset`, `skip` or cursor. You cannot ask for "the next 100 after the first 100". This surprises people coming from SQL or from a search engine with `from`/`size`. The available moves are: raise `top_k` and slice client-side; narrow the candidate set with a metadata filter or a namespace; or, for enumeration rather than search, list ids and `fetch` them instead of querying. Building a UI with deep result paging on top of a vector query is a design mistake, not a missing parameter. ## Practical defaults For RAG, a `top_k` of roughly 20–100 into a reranking step, `include_metadata=True`, `include_values=False`, and an explicit namespace is the shape most production callers converge on. Big `top_k` values are for recall analysis and offline evaluation, not for the request path.
- Why does including metadata lower the maximum top_k you can request?Because the ceiling exists to bound response size, not result count. Ids and scores are tiny, so ten thousand of them fit; add a metadata dictionary or a full 1536-float vector to each match and the same payload budget is exhausted an order of magnitude sooner. That is why the cap drops to roughly a thousand once include_metadata or include_values is set.
- How would you implement result paging for a search UI on top of Pinecone?Not with query itself — there is no offset or cursor. Fetch one generous top_k, hold the ordered ids for that query in your own layer (cache or session), and page over them client-side, hydrating each page with fetch. If the result set is genuinely large, narrow it with metadata filters or namespaces instead of trying to scroll deeper into an approximate ranking, where deep tail ordering is not stable anyway.
- When would you query by id instead of by vector?For more-like-this: you already have a stored record and want its neighbours, so passing id lets Pinecone use the stored vector and you skip a call to the embedding model. Note the source record itself normally comes back as the top match, so request one extra result and drop it. It also helps in debugging, letting you inspect a known record's neighbourhood without reproducing its embedding.
saying these in an interview costs you the question
- Assuming a higher score always means more similar
- Leaving include_values on in a RAG request path
- Expecting an offset parameter for paging results
- Treating 0.75 as a universal relevance threshold
- Forgetting the query targets a single namespace