skip to content

How do you load ten million vectors into a Pinecone index efficiently?

level: seniorimportance: should knowfreq 40%

answer

  1. Round trips are the enemy
  2. Two dials: batch size and concurrency
  3. Rejections mean slow down, not retry harder
  4. Collect every future, or lose batches
  5. Reconcile counts after writes stop

basics

~20 s

Batch roughly 100 records per upsert to stay under the request size cap, send batches concurrently using an Index configured with pool_threads and upsert(async_req=True), retry 429 responses with exponential backoff, and reconcile counts with describe_index_stats after writes stop.

solid answer

~50 s

Sequential single-record writes are the failure mode here: ten million round trips is hours of pure latency. Three levers fix it. **Batch** — group about 100 records per `upsert` call, sized to stay under the roughly 2 MB request limit given your dimension and metadata. **Parallelise** — create the index client with `pc.Index(name, pool_threads=30)` and issue `index.upsert(vectors=chunk, async_req=True)`, then collect the returned futures with `.get()` so failures surface instead of being swallowed. **Respect backpressure** — a serverless index will rate-limit you; treat 429s as a signal to back off exponentially and reduce concurrency, not to retry harder. Around it: use deterministic ids so any batch can be safely re-sent, keep metadata lean because it inflates every request, and verify at the end with `describe_index_stats()` rather than during, since counts lag. For very large initial loads Pinecone also offers an import path from object storage, which avoids driving the whole corpus through client requests.

code

python · 15 lines
python
from pinecone import Pinecone

pc = Pinecone(api_key="...")
index = pc.Index("docs", pool_threads=30)

def chunks(seq, n=100):
    for i in range(0, len(seq), n):
        yield seq[i:i + n]

futures = [
    index.upsert(vectors=batch, namespace="corpus-v2", async_req=True)
    for batch in chunks(records, 100)
]
for f in futures:
    f.get()  # surfaces failures instead of dropping batches

go deeper

for a junior

Know that you batch records into each upsert call rather than sending one at a time, and that the request size limit is what caps the batch.

for a middle

Explain both dials — batch size from the roughly 2 MB request cap, and concurrency via pool_threads with async_req=True — and why every future must be collected so failures are not lost.

for a senior

Show operational judgment: adaptive concurrency under 429 backpressure, deterministic ids plus checkpoints for restartability, reconciliation after writes settle, and the fact that embedding usually dominates the runtime.

for a principal

Own the ingestion architecture — whether to push from clients or import from object storage, loading into a fresh namespace for a clean cutover and rollback, and the cost model across embedding, writes and the metadata you chose to store.

## Where the time actually goes At ten million vectors, three costs compete and only one of them is Pinecone: 1. **Embedding** — usually the real bottleneck. Ten million texts through an embedding model is typically far more time and money than the writes. Any credible answer mentions that the ingest pipeline is embedding-bound, and that you checkpoint embeddings so a failed load never re-embeds. 2. **Network round trips** — one request per record is ten million sequential HTTPS calls. At even 30 ms each that is 83 hours. This is the number worth quoting because it makes the case for batching and concurrency instantly. 3. **Write throughput at the service** — bounded, and it pushes back with rate limiting when you exceed it. ## Lever one: batch Group records into `upsert` calls of about 100. The constraint is the request size ceiling of roughly 2 MB — a 1536-dimension float vector serialises to several kilobytes, so a couple of hundred records with modest metadata is near the limit and fat metadata gets there sooner. Batching turns ten million round trips into a hundred thousand. Size the batch by measurement, not folklore: serialise one representative record, divide the budget, take a safety margin. If you see request-too-large errors, halve the batch. ## Lever two: concurrency Batching alone still leaves you serial. The Python client supports in-flight parallelism: ``` index = pc.Index("docs", pool_threads=30) futures = [index.upsert(vectors=chunk, async_req=True) for chunk in chunks] results = [f.get() for f in futures] ``` `pool_threads` sizes the connection pool; `async_req=True` makes each `upsert` return a future instead of blocking. The critical discipline is **calling `.get()` on every future**: without it, failed batches vanish silently and you discover the hole weeks later as missing search results. Bound how many futures are outstanding at once (a sliding window rather than materialising ten million calls) so memory stays flat. If you drive ingestion from multiple worker processes, partition by id so two workers never write the same record, and remember there is no cross-record transaction — partial visibility during the load is normal. ## Lever three: backpressure Serverless indexes rate-limit writes. Getting 429s is not a bug, it is the system telling you the ceiling. The correct response is exponential backoff with jitter and a reduction in concurrency; the incorrect one is an immediate tight retry, which converts a slowdown into a thundering herd. A good loader adapts: raise concurrency while responses are clean, drop it on rejection. Also make failures loud — count them per batch, log the ids, and let the job exit non-zero rather than reporting success over a partial load. ## Idempotency and restartability A multi-hour load *will* be interrupted. Two properties make that survivable: - **Deterministic ids** (derived from the source, not generated fresh per run), so re-sending a batch overwrites rather than duplicates. - **A checkpoint** of which input shards have been acknowledged, so a restart resumes rather than re-runs. Combined, resuming is safe even if the checkpoint is slightly stale — the overlap is harmless overwrites. ## Verifying the load `describe_index_stats()` reports vector counts per namespace, but those counts lag while writes are in flight, so it is a **reconciliation** tool, not a progress bar. Track progress from your own acknowledged-batch counter. When writing has stopped and settled, compare your count against the index stats; a persistent shortfall means real dropped batches and warrants investigation, whereas a gap mid-load means nothing. Also remember: freshly written vectors are not instantly searchable. Do not kick off recall evaluation the moment the loader exits — let the index settle first. ## Alternatives worth naming - **Import from object storage.** For very large initial loads Pinecone provides a bulk import path that reads prepared files from object storage rather than having your client push every record. Where it fits, it removes client throughput from the equation entirely and is usually cheaper than a client-side load. It is an initial-load tool, not a substitute for streaming upserts of ongoing changes. - **Load into a fresh namespace and cut over.** Building a large corpus into a new namespace, then switching reads and dropping the old one, gives a clean rollback and avoids serving a half-built corpus. It also means a failed load costs you nothing but the namespace. - **Keep metadata small.** Every unnecessary metadata byte is multiplied by ten million on the way in and by `top_k` on every query afterwards. ## The judgment being tested An interviewer wants to hear that you sized batches from the request limit rather than guessing, that you parallelised with backpressure rather than open-loop, that failures are surfaced and the job is restartable, and that you know the write path is probably not the slowest part of the pipeline.

  • Why must you call .get() on the futures returned by upsert(async_req=True)?
    Because that is where errors surface. Without it a rejected or failed batch is never observed, the loader reports success, and you are left with a silently incomplete index that only shows up as missing search results later. Collecting results also gives you natural backpressure — bound the number of outstanding futures so you do not queue the entire corpus in memory.
  • Your loader starts receiving 429 responses halfway through. What do you change?
    Back off exponentially with jitter and lower concurrency; a 429 means you are above the write ceiling, so retrying immediately just amplifies the pressure. A well-behaved loader adapts in both directions — increasing parallelism while responses stay clean, shedding it on rejection — and never treats rate limiting as a fatal error, since the batches are safely retryable given deterministic ids.
  • How do you make a multi-hour load restartable?
    Deterministic ids plus a checkpoint. Ids derived from the source mean re-sending a batch overwrites rather than duplicates, so overlap after a restart is harmless. Checkpoint which input shards were acknowledged, and resume from slightly before the last one. Cache embeddings separately, since re-embedding is usually the expensive half of the pipeline, not the writing.
  • Why is describe_index_stats a poor progress indicator during the load?
    Because its counts are eventually consistent and lag behind accepted writes, so the number trails and then jumps, which looks like stalls and bursts that are not real. Track progress from your own count of acknowledged batches, and use index stats once writing has stopped and settled as a reconciliation check — a persistent shortfall then is a genuine signal of dropped batches.

saying these in an interview costs you the question

  • Upserting one vector per request for a large corpus
  • Firing unbounded concurrent requests with no backpressure
  • Ignoring futures so failed batches disappear silently
  • Retrying 429 responses immediately without backoff
  • Using index stats as a live progress bar

context