How do Pinecone's upsert() and update() differ when changing one metadata field?
answer
- PUT versus PATCH, in vector form
- Omitted fields are not merged
- One of the two never inserts
- No unset for a metadata key
- Last writer wins, no version check
basics
~20 supsert() replaces the whole record, so any metadata you do not resend is lost and the embedding must be sent again. update(id=..., set_metadata={...}) patches just the named fields in place, leaves the rest alone, and never creates a record that does not already exist.
solid answer
~40 s`upsert()` is a whole-record write keyed by id: whatever you send becomes the record. Send values without metadata and the previous metadata is gone; send metadata you had to re-read first and you have paid for a read plus a full embedding round trip. `index.update(id="doc-1#chunk-0", set_metadata={"status": "archived"}, namespace="...")` is the partial write — it sets the named keys, leaves every other metadata key and the stored vector untouched, and sends only a few bytes. Two behaviours worth stating: `update` never inserts, so patching an id that is not present simply does nothing rather than creating a record; and `set_metadata` can add or overwrite keys but has no delete-a-key form, so removing a field means re-upserting the full record without it. `update` can also replace the vector via `values=` when the embedding itself changed.
code
python · 13 lines# Patch one metadata field; vector and other keys untouched
index.update(
id="doc-1#chunk-0",
set_metadata={"status": "archived"},
namespace="tenant-a",
)
# Full replace: omitting metadata here would erase it
index.upsert(
vectors=[{"id": "doc-1#chunk-0", "values": [0.1] * 1536,
"metadata": {"doc": "doc-1", "page": 4, "status": "archived"}}],
namespace="tenant-a",
)go deeper
Be able to say that upsert writes a whole record while update changes named fields, and that resending a record without its metadata wipes that metadata.
Explain the cost difference — update sends a few bytes and keeps the stored vector, upsert requires the full embedding — plus the two edges: update never inserts and cannot delete a key.
Show that you design around the quiet failure: metadata your filters depend on being erased by a partial upsert, and last-writer-wins with no compare-and-set when two pipelines touch the same id.
Own where the system of record lives. Decide whether Pinecone metadata is authoritative or a projection, and set the ownership rules that keep concurrent mutation and periodic full re-ingestion from fighting each other.
## Two write paths, different semantics Pinecone gives you two ways to change an existing record, and they differ in exactly the way PUT differs from PATCH in an HTTP API. **`upsert()` — full replace.** The record you send *is* the record afterwards. Fields you omit are not merged with what was stored; they are simply absent from the new record. That is a feature for ingestion (re-running a job over a source produces a clean, deterministic state) and a hazard for edits (a partial payload silently drops data). **`update()` — partial patch.** `index.update(id=..., set_metadata={...}, values=[...], namespace=...)` touches only what you name. `set_metadata` writes the given keys and leaves every other metadata key intact. Omit `values` and the stored embedding is untouched. ## Why the difference matters in practice Consider a document chunk whose metadata is `{"doc": "doc-1", "page": 4, "tenant": "acme", "status": "active"}` and you want `status` to become `archived`. With `update` it is one small request naming one key. The embedding never leaves your process, the other three keys survive, and there is nothing to re-read first. With `upsert` you must reconstruct the entire record. That means either keeping the original vector around (a few kilobytes per record you now have to carry through your code path) or fetching it back with `include_values=True`, and merging the metadata yourself. If you skip that merge and send only `{"status": "archived"}`, you have just destroyed `doc`, `page` and `tenant` — and the loss is silent. Queries keep working; the filters that depended on `tenant` simply stop matching. This is the classic bug in this area, and it shows up as "our tenant filter suddenly returns nothing for old records." ## Two sharp edges on update **It does not insert.** `update` is an operation on an existing record. Point it at an id that is not in the namespace and no record appears — you do not get a new record and you should not rely on an error to tell you. That makes it wrong for any "write if missing" path: for create-or-modify semantics you want `upsert`, and if you must know whether the record existed, `fetch` first. **It cannot remove a metadata key.** `set_metadata` sets keys. There is no unset form, so a field you want gone stays gone only if you rewrite the whole record with `upsert`, omitting it. Teams that expect to prune metadata should design for that: either use a sentinel value they can filter on, or accept a periodic full re-upsert from the system of record. ## Choosing between them A workable rule: - **Ingestion and re-ingestion → `upsert`.** You have the full record in hand, deterministic ids make it idempotent, and full replacement is exactly the semantics you want when the source changed. - **Small mutations to existing records → `update`.** Flags, lifecycle status, counters, ownership changes, anything where the embedding is unchanged. - **Re-embedding after a model change → `upsert`** (you are rewriting values and probably metadata anyway) — unless the change is purely the vector for a subset, where `update(values=...)` is leaner. ## Consistency and ordering Both are writes, so both are subject to the same visibility delay before queries and fetches reflect them; do not read back immediately and assert. And neither gives you compare-and-set: two concurrent writers patching the same record produce last-writer-wins, with no version check and no conflict error. If two pipelines can touch the same id, either partition ownership so only one does, or serialise the mutation upstream. There is no transaction spanning multiple records either — a batch upsert is not atomic across the batch, so partial visibility during a large write is normal and your readers must tolerate it. ## What an interviewer is checking That you know a vector database record is not a row you can freely `SET column = value` on, that you can name the operation that patches versus replaces, and that you have internalised the failure mode: the destructive one is quiet, and it destroys the metadata your filters depend on.
- What happens if you call update() with an id that does not exist in that namespace?Nothing is created. update patches an existing record and has no insert path, so the call leaves the namespace unchanged. That makes it unsuitable for create-or-modify flows — use upsert there. It also means you cannot infer existence from the call succeeding; if the distinction matters, fetch the id first, and remember that a very recent write may not be visible yet.
- How do you actually remove a metadata key from a Pinecone record?By rewriting the record with upsert, sending the metadata dictionary without that key. set_metadata only sets or overwrites keys; there is no unset operation, so a partial update can never shrink the metadata. This needs the vector in hand, so either keep it or fetch it back with include_values=True first — one reason to keep the system of record outside Pinecone.
- Two services patch the same record's metadata at the same time. What are the semantics?Last writer wins. There is no version, etag or compare-and-set on the write path, so neither call fails and one set of values simply survives. If concurrent mutation is realistic, give a single service ownership of each id, or serialise the change upstream — for example by making the source-of-record system emit the change and having one consumer apply it.
saying these in an interview costs you the question
- Thinking upsert merges the metadata you omit
- Expecting update() to create a missing record
- Believing set_metadata can delete a key
- Assuming concurrent writes are conflict-checked
- Re-upserting full records for a one-flag change