skip to content

How do you embed a large document corpus with Gemini's embedding API?

level: seniorimportance: should knowfreq 40%

answer

  1. Model input limit is small, chunk first
  2. One request, many texts
  3. Quota, not compute, sets the pace
  4. Backoff with jitter, not tighter loops
  5. Resume by chunk ID after interruption

basics

~20 s

Chunk documents to fit the model's input token limit, send many chunks per request via the batch embedding endpoint, run bounded concurrency with backoff on rate-limit errors, and make the job resumable by chunk ID so an interruption does not restart it.

solid answer

~50 s

Gemini's embedding models cap input at a fixed token count per text, so the first step is chunking — the model will not summarise a long document for you, and an oversized input is a per-request problem you must solve before the call. Then batch: the API's `batchEmbedContents` endpoint (a list passed to `contents` in the GenAI SDK) embeds many texts in one round trip, and each entry carries its own config, so a batch can even mix task types. Throughput is then governed by your project's rate limits, not by compute, so run bounded concurrency across batches and retry rate-limit responses with exponential backoff and jitter rather than tightening the loop. Finally, make the job resumable and idempotent per chunk ID: a corpus-sized backfill will be interrupted, and you want to restart where it stopped. Pair each returned vector with its source chunk immediately, since results are positional.

code

python · 18 lines
python
from google import genai
from google.genai import types

client = genai.Client(api_key="YOUR_KEY")

def embed_batch(chunks):
    resp = client.models.embed_content(
        model="gemini-embedding-001",
        contents=[c["text"] for c in chunks],
        config=types.EmbedContentConfig(
            task_type="RETRIEVAL_DOCUMENT",
            output_dimensionality=768,
        ),
    )
    return [
        {"id": c["id"], "vector": e.values}
        for c, e in zip(chunks, resp.embeddings)
    ]

go deeper

for a junior

Know that embedding models take limited-length inputs, so documents must be chunked, and that you can send a list of texts in one call instead of looping one request at a time.

for a middle

Explain batching versus per-item requests, that results are positional and must be zipped to chunk IDs, and that rate limits rather than model speed set the throughput ceiling.

for a senior

Show the production pipeline: token-aware structural chunking, bounded concurrency with jittered backoff, idempotent upserts by chunk ID, dead-lettering bad inputs, and a counted completeness assertion before cutover.

for a principal

Own the cost and reliability tradeoffs: choose synchronous versus offline batch processing against latency needs and price, size pacing against quota so the job does not oscillate, and mandate that model, task type and dimension come from one governed configuration.

## The four constraints Embedding a corpus is not a scaled-up version of embedding one string. Four things bind, in this order: input length, request batching, rate limits, and job restartability. ## 1. Input length forces chunking Gemini's embedding models accept a bounded number of input tokens per text — a couple of thousand, far less than a generation model's context window. Anything longer must be split by you. This is the most common surprise for teams coming from generation, where a long document is simply a long prompt. Chunking is a retrieval-quality decision, not just a plumbing one: - Split on structure (headings, paragraphs) before splitting on length, so chunks are coherent units rather than sentences cut mid-thought. - Overlap adjacent chunks modestly so a fact spanning a boundary is retrievable from either side. - Keep chunks well under the token limit rather than exactly at it, so a tokenisation surprise does not fail a request. - Record a stable chunk ID and its source document; you need it for resumability, for dedup, and for citations at query time. Measure token counts rather than characters — a limit expressed in tokens is not a character count, and code, non-Latin scripts and dense punctuation all tokenise worse than English prose. ## 2. Batch to amortise round trips One HTTP request per chunk is the naive shape and it is slow: you pay connection and request overhead thousands of times. The API provides `:batchEmbedContents`, and in the google-genai SDK you reach it simply by passing a list to `contents`. Key properties: - Each entry in the batch carries its own configuration, so task type can differ per item — useful when one job ingests documents and pre-computes query vectors for an evaluation set. - Results come back **in input order**, one per entry. Zip them against your chunk IDs immediately; anything that reorders or filters between call and zip risks pairing the wrong vector with the wrong document, which produces plausible nonsense at search time. - There is a cap on how many items one request may carry, and the batch is also bounded by total tokens, so "send the whole corpus" is not an option — you are choosing a batch size, not eliminating batching. ## 3. Rate limits govern wall-clock time At corpus scale the bottleneck is your project's request and token throughput quota, not the model. The pattern that works: - **Bounded concurrency** — a fixed worker pool over a queue of batches, sized empirically against your quota. Unbounded concurrency just converts throughput into rate-limit errors. - **Exponential backoff with jitter** on rate-limit responses, honouring any retry hint the response carries. Jitter matters because a synchronised fleet retrying in lockstep re-creates the spike that caused the throttle. - **Retry only what is retryable.** Rate limits and transient server errors, yes; an invalid request or an over-length input, no — those need fixing, and retrying them burns quota that successful work needs. - **Steady pacing beats bursts.** A job that runs at 70% of quota for four hours finishes sooner than one that saturates, throttles, backs off and oscillates. If the provider offers an asynchronous batch/offline mode for embeddings, a corpus backfill is exactly the workload it exists for: you trade latency you do not need for throughput and often a lower price. Check the current pricing page before assuming either. ## 4. Make the job resumable A multi-hour backfill will be interrupted — a deploy, a network blip, a quota exhaustion, an OOM. Design for it: - Persist progress per chunk ID, not per document and certainly not per run. - Make writes idempotent — upsert by chunk ID so a replayed batch overwrites rather than duplicates. - Log failures with their chunk IDs into a dead-letter list to be inspected and re-run, instead of aborting the whole job on one bad input. - Emit a counted completeness check at the end: chunks expected versus vectors stored. Never gate a cutover on "the job exited"; gate it on that count matching. ## Consistency with the index Every chunk in one index must be embedded with the same model, the same task type and the same output dimensionality. Read all three from one configuration source in the ingestion job, write them as collection metadata, and assert on them before inserting. The failure mode of a drifted setting is not an exception — it is an index that returns worse results for reasons nobody can reproduce. ## Rough shape of a good pipeline Documents → structural chunker with token-aware limits → stable chunk IDs → queue of fixed-size batches → worker pool with backoff → upsert vectors keyed by chunk ID → completeness assertion → offline recall check against an evaluation set.

  • Why is pairing results with chunk IDs by zipping immediately so important?
    Batch results are positional: the Nth embedding belongs to the Nth input and nothing in the payload identifies it. Any code that filters, reorders or partially retries between the call and the pairing can silently associate a vector with the wrong document. The index then returns confident, wrong neighbours, and the bug is nearly untraceable from search logs. Zip at the call site and carry the ID onward.
  • Your backfill starts getting rate-limit responses. What do you change?
    Reduce concurrency and add exponential backoff with jitter, honouring any retry hint on the response. Do not retry immediately or in lockstep across workers — that recreates the spike. Aim for steady pacing below quota rather than saturation, and only retry retryable failures; invalid or over-length inputs go to a dead-letter list for fixing.
  • Can a single batch embedding request mix different task types?
    Yes — the batch endpoint takes a list of individual embedding requests, each with its own configuration, so one call can embed documents with the document task type and evaluation queries with the query task type. It is rarely needed in ingestion, where every chunk shares one setting, but it is legitimate and worth knowing.
  • How do you decide chunk size beyond just fitting the token limit?
    Fit is the floor, not the goal. Split on structure so each chunk is a coherent unit, add modest overlap so facts on a boundary stay retrievable, and keep a margin below the limit against tokenisation surprises. Then validate the choice with recall@k on an evaluation set — chunk size affects retrieval quality at least as much as the model does.

saying these in an interview costs you the question

  • Sending whole documents and expecting silent truncation
  • One HTTP request per chunk at corpus scale
  • Retrying rate-limit errors immediately without backoff
  • Restarting an interrupted backfill from the beginning
  • Gating cutover on the job exiting rather than a chunk count

context