skip to content

Your LangSmith experiment over 2,000 examples keeps hitting provider 429s — how do you tune the run?

level: seniorimportance: should knowfreq 46%

answer

  1. cap examples in flight, not calls
  2. multiplier: target plus each judge
  3. judges share the target's quota
  4. async target needs the async entry point
  5. retry in the client, not the platform

basics

~20 s

Lower max_concurrency on evaluate(), which caps how many examples run in parallel. Count calls per example — one target call plus one per LLM evaluator — since that multiplier, not the example count, sets your request rate. Use aevaluate for async targets and shrink the run to a split while iterating.

solid answer

~50 s

The knob the tool gives you is `max_concurrency`, which caps how many examples `evaluate` (or `aevaluate`) executes in parallel; the default fans out aggressively, which is exactly what saturates a provider quota. But before turning it down, do the arithmetic: each example is one target call **plus one call per LLM-based evaluator**, so an experiment with two judges at concurrency 8 has up to 24 requests in flight, not 8. If the target and the judges hit the same provider account, they compete for one quota — putting the judge on a different key or a different model separates them. If the target is `async def`, use `aevaluate` and await it, which gets you throughput from I/O overlap rather than from raw parallelism. Beyond that: keep provider-side retry/backoff enabled in your model client so transient 429s don't turn into errored rows, and iterate against a small split so you only pay the full 2,000-example run when you actually mean to.

code

python · 21 lines
python
import asyncio

from langsmith import aevaluate


async def my_async_app(question: str) -> str:
    return "stub answer"


async def target(inputs: dict) -> dict:
    return {"answer": await my_async_app(inputs["question"])}


asyncio.run(
    aevaluate(
        target,
        data="support-qa",
        max_concurrency=4,
        experiment_prefix="throttled-rerun",
    )
)

go deeper

for a junior

Know that evaluate() runs examples in parallel and that max_concurrency is the argument that caps it; turning it down is the first response to throttling.

for a middle

Explain the arithmetic — requests in flight are concurrency times calls per example, counting each LLM evaluator — and that aevaluate is the entry point for an async target.

for a senior

Show the production instincts: separate credentials and judge model so evaluation does not eat production quota, client-side retry so throttles do not become errored rows, and verifying the error count before trusting the score.

for a principal

Own the budget and scheduling policy — which suites may run at what size and cadence, what evaluation is allowed to cost per run, and how quota is partitioned so a sweep can never degrade live traffic.

## What max_concurrency actually caps `evaluate` and `aevaluate` both take `max_concurrency`, and it bounds **examples in flight**, not requests in flight. That distinction is the whole answer to why people set it to 4 and still get throttled. Each in-flight example may issue: - one or more calls from your target (a chain with a rewrite step and a generation step is already two), - one call per LLM-based evaluator attached to the run. So the request rate is roughly `max_concurrency × calls_per_example`, and with two judges and a two-step target that multiplier is five. Compute the multiplier first, then choose the concurrency that keeps you under the quota you actually have. ## Whose quota is being consumed A subtlety worth raising in an interview: the target and the judges are usually the same provider account. The evaluation is therefore competing with itself, and worse, if the same account serves production, a big experiment can throttle live traffic. The clean separations are (a) a separate API key or project for evaluation, (b) a different — usually cheaper and higher-quota — model for judging than for the system under test, and (c) running the big sweep off-peak. This tree owns the concurrency knob; the provider's own quota model and how limits are enforced belong to the provider's documentation, and the honest interview answer names which side each lever sits on. ## Sync versus async `evaluate` runs a synchronous target with a thread pool. `aevaluate` is the async twin: an `async def` target awaited inside an event loop, with the same `data`, `evaluators`, `experiment_prefix`, `max_concurrency` and `num_repetitions` arguments. For an I/O-bound target — which almost every LLM application is — the async path gets more useful overlap per unit of resource, and it is mandatory if your application code is already async, because handing a coroutine function to `evaluate` records coroutine objects as outputs rather than results. What async does **not** do is exempt you from the provider's rate limit. It changes how efficiently you wait, not how many requests per minute you are allowed. ## Handling the 429s that still happen Even correctly sized, a long run will occasionally get throttled — other traffic exists. The defence is in the model client, not in LangSmith: most provider SDKs retry with backoff by default and let you raise the retry count. Get that right, or transient throttling shows up as **errored rows** in the experiment, and an average computed over the rows that happened to succeed is a biased number. Always read the error count on a finished experiment before reading its score. ## Shrink the loop The cheapest fix for a 2,000-example run is not to run 2,000 examples. While iterating on a prompt, run a small split of a few dozen rows for the fast signal, and save the full sweep for the change you are ready to commit to. This is a cost and wall-clock decision as much as a rate-limit one: 2,000 examples with two judges is 6,000 model calls, and that bill is paid every time somebody clicks run. ## Wall clock and the long tail With concurrency capped low, the experiment's duration is dominated by the slowest examples. If a handful of rows involve long retrieval or very long outputs, they will still be finishing when everything else is done. Two practical responses: set a sane timeout inside the target so one hung call cannot stall the run, and expect that halving concurrency roughly doubles wall clock — which is fine for a nightly sweep and painful for an interactive loop, another reason the fast tier exists. ## What to say when asked A strong answer moves in this order: measure the calls-per-example multiplier; cap `max_concurrency` from that number rather than by guessing; separate evaluation credentials and judge model from production; ensure client-side retry so throttles don't become errored rows; verify the error count afterwards; and keep the everyday loop on a small split so the full run is a deliberate act. A weak answer is "turn concurrency down until it stops failing", which works but leaves you with an unexplained magic number and no idea whether the run is now needlessly slow.

  • You set max_concurrency to 4 and are still throttled. What did you miss?
    That the cap counts examples, not requests. With a two-step target and two LLM evaluators, four examples in flight can be twenty requests in flight. Work out calls per example, multiply, and size the concurrency from that. If the judges and the target share one provider account, the run is also competing with itself for the same quota.
  • Why do throttled runs make an experiment's average untrustworthy?
    Because a request that exhausts its retries raises, and the row lands as an error rather than a low score. The aggregate is then computed over whichever rows happened to get through, which is not a random sample — the slow, long, expensive examples are the likeliest to fail. Check the error count before quoting any number.
  • When is aevaluate genuinely better than evaluate, and when does it change nothing?
    It is better when your target is already async or clearly I/O-bound: awaiting overlapping provider calls in one event loop is cheaper than holding a thread each. It changes nothing about provider quota — the same requests per minute are still the same requests per minute — and it is required rather than optional when the target is defined with async def.
  • How would you keep a large experiment from throttling production traffic?
    Give evaluation its own credentials or project so the two draw on separate quota, pick a cheaper high-throughput model for the judges instead of the production model, and schedule the full sweep off-peak. Iterating against a small split during the day keeps the heavy run to the moments you actually intend it.

saying these in an interview costs you the question

  • Thinking max_concurrency limits requests rather than examples
  • Ignoring evaluator calls in the rate arithmetic
  • Assuming async removes the provider rate limit
  • Reading an average without checking errored rows
  • Running the full dataset for every prompt tweak

context