skip to content

How do you delete every vector for one document from a Pinecone serverless index?

level: seniorimportance: should knowfreq 46%

answer

  1. Three delete shapes, one unavailable
  2. Serverless cannot delete by filter
  3. Plan the erasure unit at ingest
  4. Prefix in the id, then page it
  5. Shrinking documents leave orphans

basics

~20 s

Serverless indexes do not support deleting by metadata filter, so you must delete by id. Give a document's chunks a shared id prefix at ingest time, page those ids with index.list(), and pass them to index.delete(ids=[...]) in batches within the right namespace.

solid answer

~40 s

`index.delete()` accepts three shapes: a list of ids, `delete_all=True` to clear a namespace, and a metadata `filter`. The catch is that **filter-based delete is not available on serverless indexes** — that form is a pod-based capability — so on serverless, "delete everything for document 4711" has to become "delete these ids". The design move is to make that cheap at ingest time: id chunks as `doc-4711#chunk-0`, `doc-4711#chunk-1`, and you can then enumerate them with `index.list(prefix="doc-4711#", namespace="tenant-a")`, which pages ids by prefix, and delete in batches (keep each call to a few hundred ids). If a whole tenant or corpus is going away, `index.delete(delete_all=True, namespace="tenant-a")` is far cheaper than enumerating. And remember deletes are eventually consistent, so a removed vector may still be returned for a short window.

code

python · 6 lines
python
# Purge one document's chunks from a serverless index
for id_page in index.list(prefix="doc-4711#", namespace="tenant-a"):
    index.delete(ids=list(id_page), namespace="tenant-a")

# Whole-tenant erasure is a single call
index.delete(delete_all=True, namespace="tenant-b")

go deeper

for a junior

Know that delete takes a list of ids, that delete_all clears a namespace, and that every delete is scoped to the namespace you name.

for a middle

Explain the serverless restriction on filter-based delete and the id-prefix workaround: prefix chunk ids by document, page them with index.list(), and delete in batches.

for a senior

Show that the erasure unit is designed at ingest time, not improvised during an incident, and name the traps: eventual consistency on deletes, and orphaned chunks when a re-ingested document shrinks.

for a principal

Own the data-lifecycle model — which unit is deletable, whether tenants get their own namespace, how erasure obligations are met at the serving boundary, and how re-indexing cuts over without leaving stale content behind.

## The three delete forms ``` index.delete(ids=["a", "b"], namespace="tenant-a") # by id index.delete(delete_all=True, namespace="tenant-a") # clear a namespace index.delete(filter={...}, namespace="tenant-a") # by metadata (pod-based only) ``` All three are scoped to a single namespace — deleting is never global across namespaces, which is both a safety property and a thing to remember when a cleanup "did nothing" because it ran against the default namespace. ## The serverless restriction that drives the design Delete-by-metadata-filter is **not supported on serverless indexes**. This is the fact the question is really about, and it is not a small operational detail: it means the obvious way to implement "forget this document" or "purge this user's data" is unavailable, and you have to plan for it at ingest time rather than discover it during an incident. The workaround has two moves: **1. Encode the deletion key in the id.** Since delete-by-id always works, make the ids carry the grouping you will need to delete by. The prevailing convention is a prefix and a separator: `doc-4711#chunk-0`, `user-99/doc-3/para-12`. Choose the prefix to match your actual erasure unit — usually the document, sometimes the source file, sometimes the tenant. **2. Enumerate by prefix, then delete.** `index.list(prefix="doc-4711#", namespace="tenant-a")` pages through the ids sharing that prefix, and you feed those pages into `index.delete(ids=[...])`. Keep each delete call to a few hundred ids rather than thousands, for the same request-size reasons that govern upsert batching. **Alternative: namespace as the erasure unit.** If the thing you delete is always a whole tenant or a whole corpus version, put it in its own namespace and clear the namespace in one call. This is dramatically cheaper than id enumeration and is the right shape for hard multi-tenant isolation and for blue/green re-indexing (build the new version in a fresh namespace, cut reads over, then drop the old one). ## When you did not plan ahead If ids are random UUIDs and there is no prefix to page by, you are left with poor options: query with a metadata filter at a large `top_k` to collect ids and delete them (approximate, so you may not get them all in one pass, and you must repeat until empty), or rebuild the namespace from your source of truth without the unwanted records. Both are slow and awkward, which is exactly why the id scheme is an architectural decision made on day one — an interviewer asking this question is usually probing whether you know that. ## Deletes are eventually consistent A delete is a write. It is acknowledged when accepted, and the record stops appearing in queries shortly afterwards, not instantly. Cleanup code that deletes and immediately asserts an empty result set will be intermittently wrong. If your requirement is that the record must never be served after the delete call returns — a data-erasure obligation, say — you enforce that in your own layer with a tombstone or a post-query filter, and treat the Pinecone delete as the eventual physical removal. ## Cost and scale considerations - Deleting a large corpus id-by-id is many requests; a namespace drop is one. Model the erasure unit before you model the query. - Enumeration itself costs requests, so for very large documents prefer coarse prefixes that let one listing pass cover the whole unit. - Re-ingestion is not deletion. If a document shrank from 40 chunks to 30, upserting 30 leaves chunks 30–39 behind as orphans that queries will still return. Deterministic ids make those orphans *findable* by prefix, but you still have to delete them explicitly. Getting this wrong produces stale content resurfacing in RAG answers long after the source was corrected — a genuinely damaging bug and a good thing to volunteer in the answer. ## Answer shape for an interview Name the three delete forms; state the serverless filter restriction plainly; give the id-prefix plus `list` plus batched `delete` recipe; offer namespace-as-erasure-unit as the cheaper design when it fits; and close on the two operational traps — eventual consistency and orphaned chunks after re-ingestion.

  • Why is a namespace often a better erasure unit than a metadata field?
    Because clearing a namespace is one call with delete_all=True, whereas erasing by attribute on serverless means enumerating and deleting ids. If tenants, corpus versions or ingestion batches are the things you delete wholesale, making each one a namespace turns an expensive scan-and-delete into a constant-time operation, and it gives you blue/green re-indexing for free: build the new namespace, switch reads, drop the old one.
  • A re-ingested document went from 40 chunks to 30. What is left behind, and why does it matter?
    Chunks 30 to 39 from the previous run. Upsert only replaces the ids you send, so the surplus ids survive and keep being returned by queries — stale text resurfacing in answers long after the source was fixed. The fix is to make re-ingestion delete-then-write for the document's prefix, or to list the prefix afterwards and delete any id beyond the new chunk count.
  • How would you satisfy a hard requirement that a deleted record is never served again, given eventual consistency?
    Do not rely on the delete alone. Record a tombstone in your own store at the moment of the request and filter deleted ids out of results in your retrieval layer, then let the Pinecone delete complete physical removal asynchronously. That gives you an immediate, verifiable guarantee at the serving boundary while the index catches up within its normal window.

saying these in an interview costs you the question

  • Assuming delete by metadata filter works on serverless
  • Relying on random UUID ids and no deletion strategy
  • Expecting a delete to take effect instantly
  • Thinking re-upserting a document removes its old chunks
  • Forgetting delete is scoped to one namespace

context