What is a Pinecone namespace, and how does it change query behaviour?
answer
- Logical partition inside one index
- Every data-plane call takes it
- Query sees one partition only
- Default namespace is the empty string
- Created implicitly on first upsert
basics
~20 sA namespace is a logical partition inside a Pinecone index. Each upsert and each query names one namespace, and a query searches only that partition — there is no cross-namespace search in a single call. Namespaces are created implicitly on first write.
solid answer
~50 sA namespace slices one index into independent groups of records that share the index's dimension and metric. You pass `namespace="tenant-a"` on `upsert`, `query`, `fetch` and `delete`; a query only ever scans the namespace you name, so records in other namespaces cannot appear in the results and cannot dilute the top_k. Omitting the argument targets the default namespace, whose name is the empty string `""`. You never declare a namespace up front — it comes into existence on the first upsert that mentions it and disappears when its last record is deleted. `index.describe_index_stats()` reports a per-namespace vector count, which is the usual way to confirm data landed where you expected. Because a query is scoped to a single namespace, searching N tenants at once means N calls plus your own merge — that constraint is what drives most Pinecone data-modelling decisions.
code
python · 15 linesfrom pinecone import Pinecone
pc = Pinecone(api_key="YOUR_API_KEY")
index = pc.Index("docs")
index.upsert(
vectors=[("doc-1", [0.1] * 1536, {"title": "onboarding"})],
namespace="tenant-a",
)
# Only tenant-a records can be returned here.
res = index.query(vector=[0.1] * 1536, top_k=5, namespace="tenant-a")
# Wrong namespace (or none) -> zero matches, not an error.
empty = index.query(vector=[0.1] * 1536, top_k=5)go deeper
Recall that namespace is an argument on upsert, query, fetch and delete, that omitting it uses the default empty-string namespace, and that a query only sees the namespace you name.
Explain implicit creation and deletion, why a mistyped namespace yields empty results rather than an error, and how top_k behaves when a namespace holds fewer records than requested.
Show the diagnostic reflex — read describe_index_stats() namespace counts first — and describe the fan-out-plus-merge pattern, including why merging scores is valid within one index.
Own the modelling call: namespaces give a structural isolation boundary and cheap bulk deletion, but they push cross-group search into application fan-out. Decide based on how often queries must span groups.
## The partition model A Pinecone index has one vector dimension and one distance metric, and inside it records are grouped into **namespaces**. A namespace is just a string label attached to records. Two records in different namespaces are completely independent for search purposes even though they live in the same index and share its configuration. The API surface is deliberately small: `namespace` is an optional argument on the data-plane calls. `index.upsert(vectors=[...], namespace="tenant-a")` writes into that partition. `index.query(vector=..., top_k=5, namespace="tenant-a")` searches it. `fetch`, `update` and `delete` take the same argument. If you omit it, the operation targets the **default namespace**, which is literally the empty string `""` — this is why `describe_index_stats()` on a naively populated index shows a namespace whose key is an empty string. ## Implicit lifecycle There is no create-namespace step in the index creation call. A namespace begins to exist the moment the first record is upserted with that name, and it ceases to exist when its last record is removed. That makes namespaces extremely cheap to mint compared with indexes: adding a tenant is a write, not a control-plane operation with provisioning latency. The same property makes typos dangerous. Upserting to `"tenat-a"` does not fail — it quietly creates a new namespace containing exactly those records, and the subsequent query against `"tenant-a"` returns nothing at all rather than an error. Empty results from a correct-looking query are the classic namespace bug, and the first diagnostic is always to print `describe_index_stats()` and look at the namespace keys and their vector counts. ## Query scoping is the whole point The behavioural rule that matters is: **one query, one namespace**. Pinecone does not offer a cross-namespace search or a wildcard. That has three consequences worth stating in an interview. First, isolation is structural rather than filter-based. A tenant's query physically cannot reach another tenant's vectors, because the other vectors are not in the searched partition. There is no filter to get wrong and no risk that a mis-specified condition leaks a neighbour's document — an attractive property when tenants are separate customers. Second, `top_k` is per namespace. If tenant A's namespace holds ten records and you ask for `top_k=20`, you get ten. Nothing is borrowed from elsewhere in the index to fill the gap. Conversely, a large noisy namespace can no longer crowd out a small one, because they are never scored against each other. Third, fan-out is your problem. A feature that must search fifty tenants at once needs fifty `query` calls issued concurrently plus a client-side merge of the scored results — and merging is only meaningful because all namespaces in an index share the same metric, so the scores are comparable. If cross-group search is the common case rather than the exception, namespaces are the wrong tool and a metadata attribute on the records is a better model. ## Namespaces versus metadata Both partition data, but they behave differently. A metadata attribute keeps all records in one searchable pool and narrows the candidate set at query time; that keeps cross-group search to one call, at the cost of the search touching a shared structure and of correctness depending on every query supplying the right condition. A namespace removes the records from each other's search space entirely and makes bulk removal trivial — deleting everything in a namespace is a single call, `index.delete(delete_all=True, namespace="tenant-a")`, whereas deleting by attribute across a large shared pool is a heavier operation. The two are also composable: it is normal to use a namespace for the hard tenant boundary and metadata for softer within-tenant slices such as document type or date. ## Operational habits Derive the namespace string from a stable identifier — an internal tenant id rather than a display name that can change — and centralise it in one function so no call site can forget it or spell it differently. Treat "query returned zero results" as a namespace question before it is a relevance question. And remember that writes are visible only in the namespace they were written to, so a smoke test that upserts into the default namespace and queries a named one will look broken for reasons that have nothing to do with embeddings. ## What good answers include Strong candidates state that the namespace is a hard query boundary, that it is created implicitly, that the default is the empty string, and that cross-namespace search requires application-level fan-out. Weaker answers describe namespaces as "like a folder" without knowing that a query cannot span them, which is exactly the fact every design decision hangs on.
- A query returns zero matches even though you just upserted 1,000 vectors. How do you diagnose it?Call describe_index_stats() and inspect the namespaces map. Most often the writes landed in the default namespace (empty string) because the upsert omitted the argument while the query names one, or a typo minted a near-miss namespace that now holds all the data. Namespaces are created implicitly, so neither mistake raises an error — the counts in the stats response show immediately which partition actually received the records.
- You need results across 40 tenant namespaces in one user-facing search. How do you implement it?Issue 40 concurrent query calls, one per namespace, each with a top_k at least as large as the final count you need, then merge and re-sort client-side. Scores are comparable because every namespace shares the index's metric. If this pattern is the norm rather than the exception, drop the namespaces and model the tenant as a metadata attribute so a single query can span them.
- How do you remove one tenant's data entirely?Delete the whole namespace's contents in one call — index.delete(delete_all=True, namespace="tenant-a") — rather than enumerating ids. The namespace stops existing once it holds no records, and it will silently reappear if anything writes to that name again. That single-call bulk removal is a real operational argument for namespaces when you have deletion or offboarding obligations.
saying these in an interview costs you the question
- Thinks one query can search all namespaces
- Believes namespaces must be declared at index creation
- Assumes top_k is filled from other namespaces
- Says the default namespace is called "default"
- Treats a typo'd namespace as an error rather than a new one