In LangSmith, how do you pick an online evaluator's sampling rate at high volume?
answer
- cost is linear in traffic
- three multipliers, not one
- narrow who is eligible first
- work back from runs per day
- cheap and wide plus strict and narrow
basics
~20 sWork backwards from the judge bill: sampled runs equal eligible runs times the rate, and each sampled run costs at least one judge model call. Narrow the rule's filter first, then set a rate that yields a few hundred to a few thousand scored runs a day.
solid answer
~60 sThe arithmetic is the whole conversation. Judge calls per day equal eligible runs times sampling rate times calls per run, and a multi-criteria judge or a long-context prompt multiplies that further. So the first lever is not the rate at all — it is the filter, because the rate is a fraction of whatever the filter admits. Restrict to root runs of the specific chain you care about, or to runs already carrying an error or a low user score, and the same rate suddenly costs a tenth as much while telling you more. Then choose the rate by what you need statistically: you want enough scored runs per window to see a distribution move, which in practice is hundreds to low thousands a day, not a percentage that sounds impressive. Split the budget rather than averaging it — a cheap model at a low rate across everything for a trend line, plus a stricter judge at a high rate on a narrow, high-stakes slice. And treat the rate as something you tune down after launch: the temptation is to leave it where you set it while debugging.
go deeper
Understand that every sampled run costs a real model call, so scoring all production traffic is expensive, and that sampling is how you keep it affordable.
Be able to do the arithmetic out loud — eligible runs times rate times calls per run — and explain why filtering to root runs of one chain is the first thing you change.
Show you size the rate from the number of scored runs you need to read a chart, run several rules with different rates instead of one compromise, and know that judge-side rate limits and failed judge calls bite at volume.
Own the standing budget: what fraction of LLM spend goes to evaluation, which teams may attach an evaluator to a high-volume project, and the review cadence that catches a rule left at full sampling after a debugging session.
## Why this is the question interviewers reach for An online evaluator is the only part of a tracing setup that spends money proportionally to your traffic *and* produces no user-facing value. It is a pure observability cost. So the interviewer wants to know whether you have ever sized one, or whether you would attach a judge to a firehose and find out on the invoice. ## The arithmetic Judge calls per day equal eligible runs per day, times the sampling rate, times the judge calls made per sampled run. Judge cost per day is that call count times the tokens each call consumes, times the model's price. Three multipliers, and people usually only think about the middle one. **Eligible runs** is what the filter admits, not what the project receives. This is the biggest lever and the cheapest to pull. **Calls per sampled run** is often more than one. A judge that scores three criteria may issue three calls. A judge that reasons before scoring produces far more completion tokens than a classifier. And the *prompt* side is not small: an evaluator over a retrieval-augmented answer has to be shown the retrieved context, which can be thousands of tokens per call. It is entirely normal for a judge call to cost more than the production call it is judging, because the judge sees the input, the output, and the context all at once. **Rate** then scales the whole thing linearly. ## Filter before rate A project receiving every nested run in a chain has maybe five to fifteen runs per user interaction. Filtering to root runs alone can be a ten-fold reduction before you have chosen a rate. Beyond that: scope to the one chain whose quality you are actually tracking; exclude health-check and internal traffic; and consider filtering to runs that already look interesting — an error status, a thumbs-down from the user, an unusually long latency. A filter that admits 2% of traffic at a 0.5 rate is often a better instrument than one that admits everything at 0.01, because it concentrates the judge on the runs where a bad score is informative. ## Choosing a rate statistically, not aesthetically "We sample 1%" is a number about your traffic, not about your confidence. What you need is enough scored runs in the window you intend to read to notice a real move. If you are looking at a daily chart and your metric is a pass rate around 0.9, a few hundred scored runs a day gives you a usable daily point; a few dozen gives you a chart that jitters and that people learn to ignore. Work back from that: pick the runs-per-day you want, divide by eligible runs per day, and that is your rate. On a small project this can legitimately be 1.0. On a large one it might be 0.002. A consequence people miss: if you want to compare two prompt versions from online scores, each version only receives its share of the sample, so your effective N per version is halved or worse. Online sampling is a poor instrument for a fast A/B verdict — that is what an offline experiment on a curated dataset is for. Online scoring is for detecting that something changed in production, not for adjudicating a change you are about to ship. ## Stratify instead of averaging One global rate forces a single compromise between cost and resolution. Several rules do not. A realistic setup: a small, fast judge at a low rate over all root runs, giving a cheap continuous trend; a second rule at a much higher rate scoped to the newly deployed prompt version for its first week; and a third at a high rate on the regulated or high-value flow where being wrong is expensive. Each has its own cost, and each is justifiable on its own terms. ## Judge choice is part of the rate decision Rate and judge model trade against each other for a fixed budget. A cheaper judge at ten times the rate may detect a regression sooner than an expensive judge at a tenth of the volume — as long as the cheap judge actually correlates with the outcome you care about, which is something you establish with human-labelled runs rather than assume. If it does not correlate, a bigger sample of noise is still noise. ## Operational hazards at volume Raising the rate during an incident is the instinct and it is usually wrong: traffic is already elevated, so cost scales twice over, judge-side rate limits start throttling, and evaluation falls behind ingest so the chart you are staring at lags reality. The better incident move is to keep the rate and tighten the filter onto the failing route, or to route the failing runs into an annotation queue and look at them directly. Also remember the judge is itself a model call with its own failure modes — timeouts, refusals, malformed scores. At high volume a nonzero fraction of judge calls fail, and a metric whose denominator is silently shrinking looks like a quality improvement. ## What a strong answer sounds like Name the multipliers, say that the filter comes before the rate, size the rate from the runs-per-day you need to read a chart rather than from a nice-looking percentage, and mention that you run several rules with different rates rather than one compromise. Then add the honest limit: online sampling detects, offline experiments decide.
- During an incident, is it a good idea to push the sampling rate to 1.0?Usually not. Traffic is already elevated, so you scale cost twice over, you start hitting the judge provider's rate limits, and evaluation queues up behind ingest so the chart lags the incident you are trying to watch. Tightening the filter onto the failing route, or routing those runs into an annotation queue and reading them directly, gets you an answer faster and cheaper.
- You sample 5% and want to compare two prompt versions from the online scores. What is the catch?Each version only gets its share of that 5%, so your effective sample per version is halved and the daily numbers are noisy for days. Grouping the chart by the metadata carrying the version works, but it is a slow instrument. If you need a verdict before shipping, run an offline experiment on a curated dataset; use the online scores afterwards to confirm the change held in production.
- Why can a judge call cost more than the production call it is judging?Because the judge is shown more input. A production call sends the prompt; the judge sees the input, the generated output, and, for a grounding metric, the whole retrieved context — often several thousand prompt tokens. If it reasons before scoring, or scores several criteria in separate calls, the completion side multiplies too.
saying these in an interview costs you the question
- Choosing a percentage that sounds good rather than sizing from runs per day
- Ignoring that the filter, not the rate, sets the base
- Assuming one judge call per sampled run
- Raising sampling during an incident when traffic is already spiking
- Treating online sampled scores as a pre-deploy A/B verdict