What does a Pinecone vector record contain, and how should you batch upserts?
answer
- Three parts to every record
- One of them is length-constrained
- Insert-or-replace, keyed by your string
- Request size cap drives batch size
- About 100 records, under ~2 MB
basics
~20 sA Pinecone record is an id string, a values list whose length must equal the index dimension, and optional metadata. index.upsert() takes a list of records, so send batches of roughly 100 to stay under the request size cap.
solid answer
~50 sEvery record you write to Pinecone has three parts: a string `id` you choose, a `values` list of floats whose length must match the index's configured dimension exactly, and an optional `metadata` dictionary of strings, numbers, booleans or string lists. You write them with `index.upsert(vectors=[...], namespace="...")`, passing either dicts (`{"id": ..., "values": [...], "metadata": {...}}`) or `(id, values, metadata)` tuples. Upsert is insert-or-replace keyed by id, so re-sending an id overwrites the whole record — that makes re-ingestion idempotent as long as your ids are deterministic. Batch: one record per call wastes round trips, and one giant call hits the per-request size ceiling (about 2 MB), so ~100 records per call is the usual working batch, smaller if metadata is fat. A single upsert call targets one namespace, so group your batches by namespace before sending.
code
python · 12 linesfrom pinecone import Pinecone
pc = Pinecone(api_key="...")
index = pc.Index("docs")
records = [
{"id": "doc-1#chunk-0", "values": [0.1] * 1536, "metadata": {"doc": "doc-1", "page": 1}},
{"id": "doc-1#chunk-1", "values": [0.2] * 1536, "metadata": {"doc": "doc-1", "page": 2}},
]
for i in range(0, len(records), 100):
index.upsert(vectors=records[i:i + 100], namespace="tenant-a")go deeper
Be able to name the three parts of a record — id, values, metadata — say that values must be the index dimension, and show a loop that upserts in batches instead of one at a time.
Explain why upsert is insert-or-replace by id and what that buys you: deterministic ids make ingestion idempotent and re-runnable. Know the request size cap is what sets batch size.
Show ingestion judgment: aligning chunk, embed and upsert batch boundaries so retries are cheap, keeping metadata lean because it rides on every request and response, and partitioning batches by namespace.
Own the data model. Decide the id scheme that lets you re-ingest and delete by source, decide what belongs in metadata versus your system of record, and set the contract that keeps embedding model changes from silently invalidating an index.
## What a record is Pinecone stores *records*, not rows. A record has: - **`id`** — a string you supply. Pinecone never generates one for you. Ids are the primary key: they are what `fetch`, `update` and `delete` address, and what makes a re-run of your ingestion idempotent. Practical ids are deterministic and derived from the source, e.g. `doc-4711#chunk-12`, rather than a fresh UUID per run (a fresh UUID per run duplicates your whole corpus on the second run). - **`values`** — the dense embedding, a list of floats. Its length **must equal the dimension the index was created with**. A mismatch is a hard request error, not a silent truncation, and it is the single most common first-day failure: you swapped embedding models and the new one emits a different width. - **`metadata`** — an optional flat dictionary. Values may be strings, numbers, booleans, or lists of strings. Nested objects and nulls are not stored; if a field is absent, simply omit the key. Metadata has its own per-record size ceiling (about 40 KB), which matters because people are tempted to stuff the whole source chunk text in there. Dense `values` are the normal case; an index can also carry sparse values alongside them for hybrid retrieval, which is a separate subject. ## The write call ``` index.upsert(vectors=[...], namespace="tenant-a") ``` Two accepted record spellings — a dict with `id`/`values`/`metadata`, or a positional tuple `(id, values, metadata)`. The dict form is worth preferring because it is self-describing and survives someone adding a field later. The operation is **upsert**, one word for insert-or-replace: if the id is new it is created, if it exists the stored record is replaced. There is no duplicate-key error and no second copy under the same id — a fact worth stating explicitly in an interview, because it is what lets you re-run a failed ingest job over the same input without cleaning up first. ## Why batching matters Each `upsert` call is one HTTPS round trip to a remote service. Writing 50,000 vectors one at a time means 50,000 round trips and, on serverless, 50,000 chances to hit a write rate limit. So you batch. The constraint from the other side is the **request size limit — roughly 2 MB per upsert call**. Work out what that means for your data: a 1536-dimension float embedding is a few kilobytes of JSON on its own, so a couple of hundred records with small metadata is already near the ceiling, and records carrying long text metadata get there much faster. That is why ~100 records per batch is the conventional default rather than a magic number: it comfortably fits under the size cap for typical embedding widths, keeps individual retries cheap when one call fails, and amortises the round trip well. If you get a request-too-large error, the fix is a smaller batch, not a bigger timeout. A related habit: chunk your input *before* embedding, so batch boundaries line up between your embedding calls and your upsert calls, and a failure retries a single aligned unit. ## Namespaces and the call boundary One `upsert` call writes into one namespace (the default namespace if you pass none). Records for different tenants therefore cannot share a batch — partition your work by namespace first, then batch within each partition. ## Failure modes to name - **Dimension mismatch** — rejected outright; check the index's configured dimension against your model's output width. - **Oversized request** — batch too large, or metadata too heavy; shrink the batch. - **Non-deterministic ids** — re-running ingestion silently doubles your data instead of replacing it. - **Fat metadata** — storing whole documents in metadata inflates every request and every response that includes metadata; store a reference and keep the text in your own store or a small field. - **Assuming the write is immediately queryable** — Pinecone acknowledges the write before it is visible to search, which is a separate consistency concern.
- What happens if you upsert a record whose values list is shorter than the index dimension?The request is rejected with an error — Pinecone does not pad, truncate, or silently accept it. Dimension is fixed when the index is created, so every record must match it. In practice this error means the embedding model changed: either re-embed the whole corpus with the model the index was sized for, or create a new index at the new width and cut over.
- Why prefer deterministic ids over generated UUIDs for chunked documents?Because upsert replaces by id. With deterministic ids like `doc-4711#chunk-12`, re-running ingestion over the same source overwrites the previous version in place, so the job is idempotent and safe to retry. With a fresh UUID per run, the second run inserts a full duplicate set that queries will happily return alongside the originals, and you have no cheap way to identify the stale copies.
- What kinds of values can metadata hold, and what should stay out of it?Strings, numbers, booleans, and lists of strings — a flat structure, no nested objects, no nulls. Keep it small: metadata counts toward the request size and toward a per-record ceiling of roughly 40 KB, and it is returned on every query that asks for it. Store filterable attributes and a pointer (document id, offset, URL); keep the full chunk text in your own store.
saying these in an interview costs you the question
- Thinking Pinecone generates the record id for you
- Believing upserting an existing id creates a duplicate
- Assuming metadata can hold nested JSON objects
- Upserting one vector per call for a large corpus
- Expecting dimension mismatch to be auto-padded