skip to content

How do you embed a large document corpus with Cohere's Embed API without hitting limits?

level: seniorimportance: should knowfreq 42%

answer

  1. ninety-six texts per call, not unlimited
  2. over-long inputs: fail or lose the tail
  3. 429 is the signal, backoff is the answer
  4. checkpoint so a crash is not a re-bill
  5. one model and input type for the whole run

basics

~20 s

Chunk the corpus, send batches within the documented per-call text limit of 96, keep input_type constant, set truncate deliberately, and drive concurrency with exponential backoff on 429. For very large corpora, use the asynchronous Embed Jobs API over an uploaded dataset instead.

solid answer

~50 s

The synchronous endpoint is batch-shaped but bounded: a single `/v2/embed` call accepts at most 96 texts, and each text is capped by the model's context length — Embed v3 models cut off around 512 tokens, later generations accept much longer inputs. The `truncate` parameter decides what happens past that: `END` and `START` silently drop the overflow, `NONE` makes the call fail instead, which is what you want during ingest so you learn your chunker is producing over-long chunks rather than quietly embedding half a document. Beyond that it is ordinary throughput engineering: fixed-size batches, bounded concurrency, exponential backoff with jitter on 429, and idempotent checkpointing so a crashed job resumes rather than re-embeds and re-bills. Keep `model`, `input_type` and `embedding_types` identical across the whole run, and rely on response ordering to zip vectors back to chunks. For very large corpora, Cohere's asynchronous Embed Jobs API embeds an uploaded dataset without you managing the loop at all.

go deeper

for a junior

Know that embed calls are batched, that there is a maximum number of texts per call, and that you send many texts per request rather than one call per document.

for a middle

Explain the two caps — texts per call and tokens per text — and what each truncate value does to an over-long input.

for a senior

Demonstrate ingest engineering: backoff on 429, bounded concurrency, batch-level checkpointing so a crash is not a re-bill, and truncate NONE to surface chunker bugs.

for a principal

Treat the run's configuration as an index contract — one model, one input type, one set of embedding types — and decide up front what a future re-embed would cost before choosing what to capture.

## The hard limits Two caps shape any ingest job: 1. **Texts per call.** A single embed request takes at most **96** texts. Sending more is a 400, not an auto-split. Your batcher must chunk the work; the client will not do it for you. 2. **Tokens per text.** Each input is bounded by the model's context length. Embed v3 models cap around 512 tokens per input — short enough that ordinary paragraphs sometimes exceed it. Later Embed generations accept far longer inputs, which changes chunking strategy substantially, so pin the model and check its documented limit rather than assuming. ## truncate: choose it deliberately `truncate` takes `NONE`, `START` or `END`. `END` (the common default behaviour) keeps the beginning of an over-long text and discards the tail; `START` does the reverse; `NONE` returns an error instead of embedding a partial text. The senior move is to run **ingest with `NONE`**. Silent truncation is exactly the kind of defect that survives to production: your chunker emits a 900-token chunk, the API keeps the first 512, the tail is never searchable, and nothing anywhere reports it. Failing the call surfaces the bug at ingest time, where you can fix the chunker. Once the pipeline is proven, a permissive setting is a reasonable safety valve for pathological inputs — but only as a conscious choice with a metric counting how often it fires. ## Throughput Within the caps this is standard concurrent-client work: - **Fixed batches.** Group chunks into batches at or below the cap and keep the batch list around — response order is the only way to rejoin vectors to chunks, so the exact list you sent must survive until the response is consumed. - **Bounded concurrency.** A worker pool of a handful of in-flight requests, not an unbounded fan-out of every batch at once. Unbounded fan-out is how you convert a rate limit into a thundering herd of retries. - **429 handling.** Rate limits differ sharply between trial and production keys, so do not hardcode a rate — treat 429 as the signal, back off exponentially with jitter, and reduce concurrency adaptively. Retry 429 and 5xx; do not retry 400, which means your request is wrong and will stay wrong. - **Timeouts and idempotency.** A timed-out call may have been billed. Checkpoint by batch: record which chunk ranges are durably stored, and on restart resume from the checkpoint rather than from zero. Re-embedding a 10M-chunk corpus because a job died at 90% is a real and avoidable cost. ## Cost tracking `meta.billed_units.input_tokens` on every response gives exact billing. Sum it as you go and log it per batch; you get a live cost meter for the run and a sanity check on your token estimates before committing to the full corpus. Always pilot on a 1% sample first and extrapolate — both cost and wall-clock. ## Consistency across the run Every vector in one index must be produced with the same `model`, the same `input_type` (`search_document` for corpus text) and the same `embedding_types`. A run that changes any of them halfway — because someone bumped a model id mid-job, or a retry path used different defaults — leaves an index with two incompatible vector populations and no way to tell them apart afterwards. Pin these in one config object read once at job start, and record them as index metadata. Since a single call can return several `embedding_types` at one token cost, decide up front whether to capture a compressed form alongside float. Capturing it during ingest is nearly free; adding it later is a full re-embed. ## When to stop writing the loop For genuinely large corpora, Cohere offers an **asynchronous Embed Jobs** path: you upload the data as a dataset and submit a job that embeds it, rather than driving batches from your own process. That trades interactive control for not owning the retry, concurrency and checkpoint logic, and it is the right default when the corpus is a bulk backfill rather than a streaming ingest. Incremental, near-real-time ingest stays on the synchronous endpoint. ## What a strong answer sounds like Name the 96-text cap and the per-input token cap; make a deliberate `truncate` choice and justify it; describe backoff, bounded concurrency and checkpointing; mention consistency of model/input_type/embedding_types as an index invariant; and finish by noting the async job path exists for bulk work.

  • Why would you deliberately set truncate to NONE for an ingest job?
    So over-long chunks fail loudly instead of being half-embedded. With END or START the API keeps part of the text and returns a perfectly normal vector, so a broken chunker produces an index where document tails are simply unsearchable and nothing reports it. NONE turns that silent data-loss bug into an error at ingest time, when you can still fix the chunker and re-run.
  • How do you make an interrupted embedding job safe to restart?
    Checkpoint at batch granularity: assign each batch a deterministic id from the chunk range, and only mark it complete after the vectors are durably written to the store. On restart, skip completed ranges. Without that, a job that dies at 90% either re-embeds and re-bills the whole corpus or, worse, leaves duplicate vectors if the writer is not idempotent.
  • How do you estimate the cost of embedding a corpus before running it?
    Pilot on a random 1% sample and read `meta.billed_units.input_tokens` from the responses, then extrapolate by corpus size. That measures your real chunking and your real text rather than a token-per-word heuristic, and it also gives you a wall-clock rate under your chosen concurrency, so you can predict both the bill and the duration of the full run.
  • Should you request compressed embedding types during the initial ingest?
    Usually yes. Because embed is billed on input tokens and a single call can return several representations, adding a compressed type alongside float costs essentially nothing at ingest time. Adding it six months later means re-reading and re-billing the entire corpus. Capturing both forms in one pass is cheap insurance against a future storage or latency change.

saying these in an interview costs you the question

  • Assuming the client auto-splits batches over the text cap
  • Leaving truncate at a permissive value during ingest
  • Retrying 429s immediately with no backoff
  • Firing every batch concurrently with no bound
  • Changing model or input_type partway through a run

context