skip to content

When would you run Mistral embeddings through the batch jobs API instead of /v1/embeddings?

level: seniorimportance: should knowfreq 34%

answer

  1. Is anyone waiting for this?
  2. JSONL in, output file out
  3. Your id travels with each line
  4. Discounted, not instant
  5. Off the interactive rate-limit budget

basics

~20 s

Use batch jobs for large offline work with no latency requirement — a corpus backfill or re-embed. You upload a JSONL file of requests, create a job naming the target endpoint, and collect an output file later, at roughly half the synchronous price and without hammering per-minute rate limits.

solid answer

~50 s

The synchronous route is for anything a user is waiting on; the batch API is for volume you can wait hours for. The flow is three steps: upload a JSONL file to the Files API with a batch purpose, where each line carries a `custom_id` and a `body` holding exactly the payload you would have POSTed; create a job with `input_files`, the `model`, and an `endpoint` field naming the API those requests target (`/v1/embeddings`, `/v1/chat/completions` and so on); then poll the job and download its output file when it reports success. The economics are the point — batch runs at a substantial discount (around half) and is processed asynchronously within a completion window, typically 24 hours, so it does not compete with your live traffic for per-minute limits. The `custom_id` matters: results are not guaranteed in input order, so it is your join key back to your own chunk ids.

code

json · 2 lines
json
{"custom_id": "chunk-0001", "body": {"model": "mistral-embed", "input": ["first chunk of text"]}}
{"custom_id": "chunk-0002", "body": {"model": "mistral-embed", "input": ["second chunk of text"]}}

go deeper

for a junior

Know that there are two ways to send the same request — immediately, or as an offline job built from a JSONL file — and that the offline route is cheaper but not instant.

for a middle

Walk the three steps: upload the JSONL to the Files API, create a job with input_files, model and endpoint, then poll and download the output. Explain what custom_id is for.

for a senior

Show the operational judgment: batch keeps a backfill off the interactive rate-limit budget, partial failures are normal and need an error-file path, and ingest must be idempotent on custom_id so reruns are safe.

for a principal

Own the trade explicitly — roughly half the cost and no contention with live traffic, bought with a completion window you cannot compress. Decide where the deadline-sensitive boundary sits and make sure no launch depends on a job you cannot reprioritise.

## The decision Two questions decide it. *Is anyone waiting?* and *how much of it is there?* A user typing a search query needs a query embedding in milliseconds — that is always the synchronous `/v1/embeddings` call. Re-embedding four million document chunks after a chunking-strategy change is nobody's interactive request; it is a job. Running that through the synchronous endpoint means saturating your per-minute limits for hours, building your own concurrency and retry machinery, and paying full price while starving live traffic. The batch API exists precisely to take that class of work off the hot path. ## The three-step flow **1. Build and upload a JSONL file.** One request per line. Each line has: - `custom_id` — your identifier for this request. It is echoed on the result. - `body` — the request payload exactly as the target endpoint expects it. ``` {"custom_id": "chunk-0001", "body": {"model": "mistral-embed", "input": ["first chunk text"]}} {"custom_id": "chunk-0002", "body": {"model": "mistral-embed", "input": ["second chunk text"]}} ``` Upload it through the Files API with the batch purpose; you get back a file id. **2. Create the job.** POST to the batch jobs route with `input_files` (a list of uploaded file ids), `model`, and `endpoint` — the string naming which API the bodies target, such as `/v1/embeddings`. `metadata` lets you tag the job with your own run identifier, which is worth doing because you will be looking at a list of jobs later trying to work out which was which. A timeout window governs how long the job may take; the default is on the order of a day. **3. Poll and collect.** The job object exposes a status that walks through queued and running to a terminal state — success, failure, cancellation, or timeout exceeded — along with counters for total, succeeded and failed requests. On completion you download the output file (and, where present, an error file for the lines that failed). Poll on a sane interval; this is a job that runs for hours, so polling every few seconds buys nothing. ## Why custom_id is load-bearing Batch results are not promised in input order, and a partially failed job returns fewer results than you sent. So you never zip results against inputs positionally. `custom_id` is the join key, and it should be *your* stable identifier — the chunk id or document id in your own store — not a running counter you would have to regenerate to interpret the file. Doing that also makes the job restartable: rerun only the `custom_id`s missing from the output. ## Economics and limits Batch pricing is discounted relative to synchronous calls — around half — which for a multi-million-chunk backfill is the difference between a line item you approve and one you argue about. Just as important, batch throughput does not consume your interactive per-minute allowance the way a self-built concurrent loop does, so your production search stays fast while the backfill runs. The trade you accept is latency and control: results arrive within the completion window, not at a time you choose, and you cannot reprioritise a line mid-run. Practical ceilings apply to file size and to how many requests one batch may carry, so very large corpora get split into several files or several jobs — which is fine, and gives you natural checkpoints. ## Failure handling Treat a batch job as a pipeline stage with three outcomes, not two: - **Success with all requests succeeded** — ingest the output file. - **Success with some failures** — ingest what landed, read the error file, and requeue the failed `custom_id`s. Common causes are inputs over the model's sequence limit, which is a chunking bug on your side. - **Terminal failure or timeout exceeded** — nothing usable; investigate before blindly resubmitting, because resubmitting an oversized file will just time out again. Make the ingest idempotent on `custom_id` so a partial rerun cannot double-write vectors. ## What it does not change Batch is a delivery mechanism, not a different model. The vectors are the same ones the synchronous endpoint produces for the same model, so you can mix batch-produced document vectors with synchronously produced query vectors in one index without concern. What you must not mix is vectors from *different models* — that stays true whichever route produced them.

  • Why does each line of the batch input file need a custom_id?
    Because results are not guaranteed to come back in input order, and a partly failed job returns fewer results than you sent — so positional zipping is unsafe. The `custom_id` is echoed on each result and is your join key back to your own chunk or document id. Using a stable id from your own store also makes reruns trivial: resubmit only the ids missing from the output.
  • A batch job finishes with a status of success but 3% of requests failed. What do you do?
    Ingest the results that succeeded, then read the error file to see why the rest failed. In an embedding backfill the usual cause is inputs over the model's sequence limit — a chunking bug on your side. Fix the chunking for those ids and resubmit just them as a smaller follow-up batch, with an ingest keyed on `custom_id` so nothing is double-written.
  • Which endpoints can a batch job target, and how does the job know?
    The job creation request carries an `endpoint` field naming the API the request bodies target — `/v1/embeddings`, `/v1/chat/completions` and similar — and every line in the input file must be a valid body for that endpoint. One job targets one endpoint, so a run that mixes embedding and chat work is two jobs, not one file with mixed bodies.
  • When is batch the wrong choice even for a large volume of work?
    When the result has a deadline shorter than the completion window, or when downstream work is blocked on it. Batch guarantees processing within a window, not by a specific minute, and you cannot reprioritise mid-run. If a launch depends on the embeddings being ready at a fixed hour, run them synchronously with your own throttling, or start the batch far enough ahead that the window cannot hurt you.

saying these in an interview costs you the question

  • Zipping batch results against inputs by position
  • Thinking batch uses a different or weaker model
  • Expecting batch results within seconds
  • Putting bodies for two different endpoints in one job
  • Treating a completed job as proof every request succeeded

context