skip to content

When does an LLM provider's batch API beat synchronous calls for bulk work?

level: middleimportance: should knowfreq 50%

answer

  1. job deadline, not item deadline
  2. half price for patience
  3. the window is a ceiling, not a promise
  4. results reconcile against submitted IDs
  5. budget lead time for the tail

basics

~20 s

Use it when the job has a deadline but no single item needs a fast answer. Batch endpoints trade a completion window of up to a day for roughly half price and much higher throughput off the provider's spare capacity.

solid answer

~50 s

Providers expose an asynchronous batch endpoint alongside the synchronous one: you submit a file of many requests, the provider schedules them against spare capacity, and you collect results when the job finishes — typically within a window of up to 24 hours, at roughly half the per-token price, and usually against a separate quota so the batch does not starve your interactive traffic. The trade is explicit: you give up per-item latency and gain throughput and cost. That is the right trade for offline work — overnight classification, backfills, bulk enrichment — where the *job* has a deadline but no *item* does. It is the wrong trade for anything a user is waiting on. Two operational facts shape the design: the window is a ceiling, not a promise, and items that do not finish inside it come back unfinished, so you need a synchronous re-drive path and enough lead time before your real cutoff.

go deeper

for a junior

Know that providers offer an asynchronous batch endpoint that is cheaper but slower per item, and be able to say that it suits offline bulk jobs while interactive requests stay on the synchronous path.

for a middle

Explain the trade in both directions: throughput and roughly half price against per-item latency and a completion window of up to a day. Name concrete workloads on each side of the line.

for a senior

Show you have operated one. Talk about partial completion, reconciling results against submitted IDs, sizing the synchronous re-drive so it does not starve interactive traffic, and working the submission time backwards from the business cutoff.

for a principal

Own the policy question: which workload classes are batch-by-default, how batch and interactive quota are kept separate, and what the organisation does when a nightly job's tail regularly threatens a downstream cutoff.

## Two endpoints, one model Most providers offer the same models through two front doors. The synchronous endpoint answers one request now. The batch endpoint accepts a large collection of requests, queues them, and runs them when the provider has capacity to spare — usually within a window of up to 24 hours, at roughly a 50% discount per token, and typically metered against a separate allowance so bulk work does not consume the same per-minute ceilings your interactive traffic depends on. This is the cleanest cost lever in the toolkit and the most under-used, because teams reach for the synchronous endpoint by habit and then spend their effort fighting rate limits they did not need to touch. ## What you are actually trading Throughput for per-item latency. On the synchronous path, an item's latency is roughly the model's response time, and total throughput is capped by how many calls you can keep in flight inside your rate limits. On the batch path, an item's latency is effectively the whole job's completion time — results are collected when the batch finishes, not as each item lands — while total throughput is set by the provider's scheduler rather than your concurrency and quota. The economics follow the same shape: half price, plus a large saving in engineering effort, because you are no longer building and tuning a fan-out that has to stay inside a token ceiling, back off on throttling, and survive partial failure. ## When batch is the right call The test is whether any *item* has a deadline, or only the *job*. Consider an insurer that must classify 380,000 first-notice-of-loss claim narratives every night so the adjuster queues are correctly prioritised when they open at 3am. No individual narrative needs an answer within seconds; the collection needs to be done by a fixed hour. That is a textbook batch workload: submit in the evening, collect before the cutoff, pay half, and leave the synchronous quota entirely free for adjusters working live claims. The same shape covers backfills over historical records, periodic re-scoring, bulk metadata enrichment, offline evaluation runs, and anything driven by a scheduler rather than a user. ## When it is the wrong call Anything interactive. Also anything whose *next* step depends on the answer within the same window — a multi-step agent loop cannot batch its steps, because step two's prompt is not known until step one returns. And anything whose input goes stale inside a day: pricing against a live market, fraud decisions on an in-flight transaction, triage of an active incident. If the answer would be wrong or worthless by the time the window closes, the discount buys nothing. ## Operating a batch job: partial completion is normal The critical operational fact is that the window is an upper bound on how long the provider will keep trying, not a guarantee that everything finishes. A submission of 380,000 items can come back reporting, say, 94% complete, with 22,000 items expired or failed. Individual items can also fail on their own merits — a malformed request, an oversized prompt, a content refusal — independently of the batch's overall state. So the pipeline needs three things the synchronous path never made you build: 1. **A reconciliation step.** Match returned results back to submitted item IDs and compute the set difference. Never assume the result file is complete because the job reported "done". 2. **A re-drive path.** The remaining items go through the synchronous endpoint at bounded concurrency, or into a second batch, depending on how much time is left before the real cutoff. Size this path for the tail you actually observe, not for zero. 3. **Lead time.** If the business cutoff is 3am and the window is 24 hours, submitting at 9pm is gambling. Submit early enough that a slow batch plus a full re-drive still lands before the cutoff, and alarm on completion percentage rather than on job status alone. ## Sizing the re-drive The re-drive is where batch workloads meet rate limits. Twenty-two thousand items pushed through the synchronous endpoint at unbounded fan-out is exactly the shape that trips throttling and, worse, competes with the interactive traffic you were protecting. Give the re-drive its own bounded concurrency, run it at a lower priority than user-facing work, and let it take as long as the remaining time allows rather than as fast as the client can go. ## The judgement to show Interviewers are looking for the explicit trade rather than the feature knowledge. The strong answer names the axis — per-item latency sacrificed for throughput and cost — identifies which workloads sit on which side of it, and then, unprompted, raises partial completion and lead time, because that is the part that turns a working prototype into a pipeline that meets a cutoff every night.

  • Your overnight batch returns 94% complete with 22,000 items unfinished an hour before the cutoff. What now?
    Reconcile the returned results against the submitted IDs to get the exact missing set, then re-drive those through the synchronous endpoint at bounded concurrency, prioritised so the highest-value items land first. Keep the re-drive below the quota your interactive traffic needs. Longer term, submit earlier so the tail has room, alarm on completion percentage rather than job status, and treat partial completion as the expected case rather than an incident.
  • Why can't an agent loop use the batch endpoint even though it makes many model calls?
    Because the calls are sequential by construction: each step's prompt depends on the previous step's output and tool results, so there is no set of independent requests to submit together. Batching only pays when the requests are known up front and independent. Agent loops can still batch genuinely parallel work — for example scoring many candidate documents in one step — but not the loop itself.
  • How would you decide the submission time for a nightly batch with a fixed morning cutoff?
    Work backwards from the cutoff: allow the full advertised completion window, plus the observed p95 duration of the re-drive for the unfinished tail, plus downstream processing, plus a margin for a failed submission that has to be resent. Measure the actual distribution over a few weeks rather than trusting the advertised window, and alarm when the projected finish crosses the cutoff, not when it misses it.

saying these in an interview costs you the question

  • Assumes the batch window is a guarantee that everything completes
  • Uses the batch endpoint for anything a user is waiting on
  • Skips reconciliation and treats the result file as complete
  • Re-drives the unfinished tail at unbounded concurrency
  • Thinks batching changes the model or the output quality

context