skip to content

In LangSmith, how do you split coverage between an online judge and human review?

level: principalimportance: should knowfreq 36%

answer

  1. two currencies, not one budget
  2. breadth is cheap, truth is not
  3. wide and shallow versus narrow and deep
  4. overlap so you can check agreement
  5. detect online, gate offline

basics

~20 s

Two different budgets: dollars for the judge, reviewer-hours for the queue. Send the judge wide and cheap across all traffic for a trend line, and send humans narrow and deep into the slices where being wrong is expensive and where their labels calibrate the judge.

solid answer

~60 s

They are not interchangeable resources, so do not think of one number. The judge buys breadth: a low sampling rate over all root runs, a cheap model, a continuous series that tells you *something moved*. Human review buys truth: a much higher rate on a deliberately narrow filter — the newly deployed prompt version, the regulated flow, runs the user thumbed down, runs where a judge and a user disagreed — producing labels no model can produce. Implement it as separate automation rules rather than one compromise rate, because the filters differ as much as the rates do. The two connect in one direction that matters: human labels are the only evidence that the judge's series is tracking reality, so keep some overlap where both score the same runs. Size the human side from actual reviewer hours, not aspiration — a queue nobody drains is worse than none because it looks like coverage. And be clear what each is for: online scores detect, human review confirms, and an offline experiment on a curated dataset is what you gate a deploy on.

go deeper

for a junior

Know that automated judges and human reviewers do different jobs — the judge covers a lot of traffic cheaply, humans produce the labels you can actually trust.

for a middle

Be able to describe the split concretely: a low-rate judge rule over all root runs, plus a higher-rate rule feeding an annotation queue on a narrow, high-value filter.

for a senior

Show that you keep an overlap so judge and human scores are comparable, that you size queue intake from real reviewer throughput, and that reviewed failures end up as dataset examples rather than as opinions.

for a principal

Own the whole allocation: which products get which coverage in proportion to blast radius, who reviews and at what cost, what the evaluation budget is as a share of LLM spend, and the standing rule that deploys are gated offline while online scores only detect.

## Two currencies, not one budget The judge costs money and scales instantly. Human review costs attention from people who have other jobs, scales badly, and cannot be bought at the last minute. Treating them as a single "quality budget" leads to the classic mistake of trading reviewer hours for sampling rate as if they were fungible. They buy different things: the judge buys *coverage*, humans buy *credibility*. ## Assign each to what it is good at **The judge goes wide and shallow.** A low sampling rate across all root runs of the application, a cheap fast model, reference-free criteria. Its output is a continuous line you can look at daily. Its job is detection: something changed, or nothing did. **Humans go narrow and deep.** A short list of slices where a wrong answer is expensive or where you need a label you can trust: - the first week of any newly deployed prompt or model version - flows with regulatory, financial or safety consequences - runs already carrying a low user score, an error, or a low judge score - runs where the judge and the user disagreed — the most information-dense sample in the whole system - a small unbiased random slice, deliberately kept, so you notice failure modes your filters were not designed to catch That last one is easy to drop and expensive to have dropped, because every other filter only finds problems you already anticipated. ## Implement as several rules One global rule at one rate is a compromise nobody chose. In practice you have three or four: a judge rule at a low rate over everything; a judge rule at a much higher rate scoped to the new version; an annotation-queue rule filtered on bad feedback; and a small random-sample queue rule. Each has its own cost line and its own justification, and each can be turned off independently when it stops earning its keep. ## The one dependency between them Human labels are the calibration set for the judge. If the judge writes a score on the same runs a reviewer grades, you can compare the two on that overlap. When they agree, the judge's wide, cheap series is meaningful and you can act on it. When they do not, the series is decoration, and every decision made from it is unfounded. So deliberately keep an overlap — do not construct the rules so that judged runs and reviewed runs are disjoint sets, which is what happens by accident if the queue is fed only by runs the judge scored low. This check is not one-off. It has to be repeated whenever the judge's prompt or model changes, whenever the application changes materially, and on some periodic cadence, because a judge that tracked reality in March can quietly stop doing so. ## Size the human side from real hours Do the arithmetic honestly. Reviewers get through some tens of substantive runs an hour. Multiply by the hours the organisation has genuinely committed per week. That is your queue's intake, and the feeding rule's sampling rate follows from it — not the other way round. The characteristic failure is a rule set at an aspirational rate, a queue that accumulates thousands of pending runs, and a team that stops opening it while continuing to describe human review as part of their process. Who reviews matters as much as how many. Labels from a domain expert are worth calibrating against; labels from whoever was free are noise with a timestamp. If the rubric needs domain knowledge, engineers are the wrong reviewers and the queue instructions must be written for the people who are right. ## What each layer is allowed to decide Be explicit, because this is where teams get burned: - **Offline experiments over a curated dataset gate a deploy.** They are repeatable, they have references, and they run before anything reaches a user. - **Online judge scores detect**, over hours and days. Traffic mix shifts underneath them, sample sizes are modest, and they cannot distinguish "the model got worse" from "the questions got harder". Never gate on them alone. - **Human review confirms**, and produces the corrections that become tomorrow's dataset examples. A regression discovered by the online judge should end its life as a dataset example added via the annotation queue, so the same failure is caught before deploy next time. If it does not, you will keep rediscovering it in production, which is a process defect rather than a tooling one. ## Making the spend defensible Expect to justify this. The useful framing is per-incident: what did an escaped quality regression cost last time in support load, customer churn or remediation, and what fraction of that is the evaluation budget. Coverage should also be proportional to blast radius — a high-volume, low-stakes summarisation feature does not deserve the same review intensity as a low-volume flow that can commit the company to something. Finally, review the whole arrangement periodically. Rules accumulate. A judge rule created for a launch six months ago is still sampling and still billing, and a queue that no longer matches how the product works is producing labels against an obsolete rubric.

  • How often should you re-check whether the judge agrees with your human reviewers?
    Whenever the judge's prompt or model changes, whenever the application changes materially, and on a standing cadence in between. Agreement is not a property you establish once; a judge that tracked reality in March can drift out of it silently. Keep a slice of runs that both score, so the comparison is always available rather than requiring a special exercise.
  • What do you actually gate a deployment on?
    An offline experiment over a curated dataset, because it is repeatable, has references, and runs before any user is affected. Online scores are a detector: sample sizes are modest, traffic mix moves underneath them, and they cannot separate a worse model from harder questions. Gating a release on a sampled online metric produces both false confidence and random blocked deploys.
  • Why deliberately keep a small random slice in human review when you could review only suspicious runs?
    Because every targeted filter only finds problems you already thought of. Failure modes nobody anticipated — a new question type, a subtle tone problem, a policy the model was never told about — show up in ordinary-looking traffic that no filter selects. A small unbiased slice is your only sight line into the unknown-unknowns, and it is the first thing teams cut.

saying these in an interview costs you the question

  • Treating judge spend and reviewer hours as one interchangeable budget
  • Setting queue intake from ambition rather than committed reviewer hours
  • Never comparing judge scores against human labels on the same runs
  • Gating a deploy on a sampled online score
  • Reviewing only flagged runs and losing sight of unanticipated failures

context