skip to content

How would you embed a large document corpus through OpenAI's Embeddings API?

level: seniorimportance: should knowfreq 46%

answer

  1. Chunk first, pack second
  2. 8192 tokens per input element
  3. Up to 2048 array elements per request
  4. Join results on the index field
  5. Bulk backfill belongs on the Batch API

basics

~20 s

Chunk documents under the model's 8192-token input limit, pack many chunks per request as an array (up to 2048 entries), reassemble results by the index field, and for very large corpora submit the work through the asynchronous Batch API instead of live calls.

solid answer

~60 s

Three mechanics do most of the work. First, **chunking**: each element of `input` must fit the model's 8192-token maximum, and oversized input returns a 400 rather than being truncated, so you split documents yourself and count tokens before sending. Second, **packing**: `input` accepts an array, and one request with 500 chunks costs one round trip instead of 500. The documented cap is 2048 array elements, and you also want to keep each request's total tokens bounded so a single failure does not cost you much. Results come back with an `index` per element — join on that, never on arrival order. Third, **offline execution**: for a corpus of any real size, write the requests to a JSONL file and submit them through `/v1/batches` targeting the embeddings endpoint. It completes within a 24-hour window at a reduced rate and against separate limits, which keeps a backfill from starving your live query traffic. Around that: retry 429 and 5xx with exponential backoff and jitter, checkpoint completed chunks so a restart is incremental, and hash chunk text so unchanged content is never re-embedded.

code

python · 14 lines
python
from openai import OpenAI

client = OpenAI()
MODEL = "text-embedding-3-small"
MAX_PER_REQUEST = 256

def embed_chunks(chunks):
    vectors = [None] * len(chunks)
    for start in range(0, len(chunks), MAX_PER_REQUEST):
        window = chunks[start:start + MAX_PER_REQUEST]
        resp = client.embeddings.create(model=MODEL, input=window)
        for item in resp.data:
            vectors[start + item.index] = item.embedding
    return vectors

go deeper

for a junior

Know that input accepts an array so many texts go in one call, and that each document must be split into chunks that fit the model's token limit.

for a middle

Explain the concrete limits — 8192 tokens per element, 2048 elements per array — and that results must be matched back to inputs using the index field on each data entry.

for a senior

Demonstrate the operational pipeline: jittered backoff on 429 and 5xx, quarantine rather than retry on 400, checkpointing, content hashing for resumability, and routing bulk backfill through the Batch API so it cannot starve live queries.

for a principal

Own the capacity story — estimate the token bill before launching, decide how backfill and live traffic share the rate-limit budget, and set the re-embedding policy that keeps a corpus migration a scheduled job rather than an incident.

## The shape of the problem Embedding a corpus is a bulk ETL job that happens to call an HTTP API. The failure modes are ETL failure modes — partial progress lost, duplicated work, one poison record killing a batch — not model failure modes. Interviewers ask this because it separates people who have embedded ten documents in a notebook from people who have backfilled millions. ## Chunking, and the hard limit behind it The `text-embedding-3-*` models accept at most 8192 tokens per input element. Exceeding it returns a 400 error; the API does not silently truncate, which is the right behaviour but means the responsibility is yours. In practice you chunk far below the ceiling anyway. A single embedding is one point in space, so a chunk covering several unrelated topics produces a blurred average that matches nothing well — the "one idea per vector" principle. Typical chunk sizes land in the hundreds of tokens with some overlap between neighbours so a sentence straddling a boundary is not lost. Count tokens with a real tokeniser rather than estimating from characters, and reject or split anything still over the limit before the request goes out. Also filter empty and whitespace-only chunks: an empty string in the array fails the entire request, taking every good element with it. ## Packing requests The `input` field takes an array. This is the single biggest throughput lever, because per-request overhead — TLS, HTTP, queueing — dominates for short chunks. One request carrying 500 chunks does the work of 500 requests. Two caps to respect: - **Array length**: the documented maximum is 2048 elements per request. - **Total tokens per request**: keep it bounded by policy, not just by the array cap. A request that fails after carrying 2048 large chunks wastes far more work on retry than one carrying 200. Batch size is a blast-radius decision as much as a throughput one. The response's `data` array carries an `index` on each element pointing back at the input position. Join on it. If you fan requests across a worker pool and write results as they land, index is the only thing tying a vector back to its chunk. ## Synchronous pipeline versus the Batch API For a live pipeline you run a bounded worker pool issuing packed requests concurrently. Concurrency is tuned against your account's rate limits, and every worker must handle: - **429 rate limited** — back off exponentially with jitter and retry. Uniform backoff across workers produces a thundering herd that re-triggers the limit. - **5xx** — retry, same policy. - **400** — do *not* retry. It is a malformed request: oversized input, empty string, bad model name. Retrying burns time forever. Quarantine the batch, bisect it to find the offending element, and log it. For a genuine backfill, the better tool is the asynchronous **Batch API**. You upload a JSONL file where each line is a request targeting `/v1/embeddings`, create a batch job via `/v1/batches`, poll its status, and download an output file when it completes. It runs within a 24-hour completion window, at a reduced rate, and against separate queues — which is the operationally important part: a multi-million-chunk backfill submitted this way cannot exhaust the rate-limit budget your production query path depends on. Live embedding of user queries and incremental new documents stays on the synchronous endpoint; bulk history goes to batch. ## Idempotency and resumability Assume the job will be interrupted. Two habits make that cheap: **Content hashing.** Key each chunk by a hash of its normalised text plus the model name and dimension width. Before embedding, check whether that key already has a vector. This makes re-runs nearly free, handles documents that are re-ingested unchanged, and means a partially failed job resumes rather than restarts. **Checkpointing.** Write vectors to durable storage in batches with a progress marker, rather than accumulating in memory and writing at the end. A five-hour job that loses everything on an unhandled exception at hour four is an avoidable outage. ## Cost hygiene Embedding is billed on input tokens, and there is no output side, so the total is knowable in advance: tokenise the corpus, multiply by the rate, and you have the bill before you spend it. Do that estimate before launching a backfill. The controllable levers are the model choice, how much text you send (deduplication, boilerplate stripping, chunk overlap — overlap of 20% means you pay for 20% more tokens), and never re-embedding unchanged content. The `dimensions` parameter is not a lever here; it shrinks the vectors you get back, not the tokens you are charged for. ## The summary answer "Chunk under 8192 tokens and well below it for retrieval quality, pack a few hundred chunks per request and join results on index, run bulk work through the Batch API so it cannot starve live traffic, retry 429 and 5xx with jittered backoff while quarantining 400s, and hash chunks so re-runs and restarts are incremental."

  • One chunk in a 500-element request exceeds the token limit. What happens, and how do you recover?
    The whole request fails with a 400 and no vectors are returned for any element. Retrying is pointless — it is deterministic. Recover by validating token counts before sending, and if a 400 still occurs, bisect the batch to isolate the offending element, quarantine it for inspection, and re-submit the rest. Never blind-retry a 400.
  • Why route a backfill through the Batch API rather than just running more workers?
    Because the synchronous endpoint shares a rate-limit budget with your production query path. A large backfill running at full concurrency causes 429s for real users. The Batch API takes a JSONL file of embedding requests, runs within a 24-hour window at a reduced rate against separate queues, and isolates the bulk job from live traffic.
  • How do you make a re-run of the pipeline cheap?
    Key each chunk by a hash of its normalised text plus the model name and dimension width, and skip anything already present. Unchanged documents cost nothing on re-ingest, an interrupted job resumes from where it stopped, and a model migration becomes an explicit key change rather than an accidental mixed index.
  • Does packing more chunks per request reduce the token bill?
    No. Billing counts input tokens regardless of how they are distributed across requests. Packing reduces round-trip overhead, connection count and wall-clock time, and it uses your request-per-minute budget more efficiently — but the token total is identical. The real cost levers are deduplication, less chunk overlap, and not re-embedding unchanged text.

saying these in an interview costs you the question

  • Sends one HTTP request per chunk for the whole corpus
  • Assumes oversized input is silently truncated to 8192 tokens
  • Retries 400 errors with backoff like a rate limit
  • Reassembles results by arrival order instead of index
  • Runs a full backfill on the live endpoint at max concurrency

context