How do you keep judge cost sane when adding LLM evaluators to a LangSmith experiment?
answer
- rows times repetitions times evaluators
- free checks carry the broad layer
- one call, several sub-scores
- score again without running again
- repetitions are the priciest knob
basics
~20 sCount first: judge calls equal rows times repetitions times judged evaluators. Then cut the multipliers — put free deterministic checks first, keep one judged metric rather than four, emit several sub-scores from a single judge call, run the expensive judge on a sample or on a nightly schedule.
solid answer
~50 sEvery judged evaluator is a model call per row, so an experiment's evaluation bill is `examples x num_repetitions x judged evaluators x calls per evaluator` — and that is on top of the target's own calls. Four judged metrics over 500 examples with three repetitions is 6,000 judge calls before your application has spent anything. The levers, in the order I would pull them: 1. **Move what you can to deterministic evaluators.** Schema validity, required citations, length, forbidden strings and exact match cost nothing and catch the crudest regressions. 2. **Collapse judged metrics.** One judge call returning `{"results": [...]}` with three sub-scores costs a third of three separate judges. 3. **Right-size the judge model.** A small model is often adequate for a binary check; reserve the expensive one for the metric that decides things. 4. **Shrink the row count.** A small per-PR split plus the full dataset nightly. 5. **Re-score, don't re-run.** `evaluate_existing` applies new evaluators to recorded runs without paying for the application again.
go deeper
Know that every LLM-as-judge evaluator costs a model call for each dataset row, so adding judges multiplies the bill, and that deterministic checks are free.
State the cost formula — examples times repetitions times judged evaluators times calls each — and name the concrete reductions: fewer judged metrics, sub-scores from one call, a smaller judge model.
Show the operating pattern: a broad free assertion layer, a small curated per-PR split with the full dataset nightly, and re-scoring recorded runs instead of re-executing the application.
Own the budget as policy — what fraction of token spend evaluation may take, which runs are allowed repetitions, and the rule that spend is tied to a decision someone acts on rather than to a habit.
## Do the multiplication before you write the evaluator The cost model of a LangSmith experiment has two independent halves. The **target** runs once per example per repetition and spends whatever your application spends. The **evaluators** then run over those runs: deterministic ones are free, and each judged one is at least one model call per row. So the judge bill is `examples x num_repetitions x judged_evaluators x calls_per_evaluator`. Every one of those factors is something you chose, and the fourth is the one people forget: a judge that makes a planning call and then a scoring call, or that re-asks on a parse failure, doubles quietly. Write the number down before the first run. Teams are routinely surprised that the evaluation of an experiment costs several times the experiment. ## Layer the evaluators: free checks first Most real regressions are not subtle. The output stopped being valid JSON. The required citation field is empty. The answer is 4,000 characters when the contract says 300. It emitted the internal system prompt. Every one of those is a two-line deterministic evaluator that costs nothing, runs instantly, and never flakes. Build that layer first and keep it broad. The judged layer then exists only for the genuinely fuzzy question — is this answer actually correct, is it grounded in the retrieved context, is the tone right — which is where a model is the only instrument you have. ## Collapse judged metrics into one call If three of your judged metrics look at the same output with the same context, they can usually be one prompt. An evaluator that returns `{"results": [{"key": "grounded", ...}, {"key": "complete", ...}, {"key": "on_tone", ...}]}` produces three columns from one call. The saving is exactly proportional: three judged evaluators become one, so the second factor in the cost formula divides by three. The cost of collapsing is coupling — a change to one criterion means touching the shared prompt, and one parse failure loses all three scores. It is still usually the right trade at scale. ## Right-size the judge and the sample Binary, well-specified checks often do not need a frontier model; a small fast model with a tight rubric will agree with the big one most of the time. Reserve the expensive judge for the metric that gates a decision. On row count, split the dataset by purpose. A per-PR run over a small, hand-picked, high-signal split gives fast feedback at bounded cost. The full dataset with the expensive judge and repetitions runs nightly or before release, where a longer wall clock and a bigger bill are acceptable. This is a scheduling decision, not a quality compromise, provided the small split is curated rather than randomly sampled. ## Do not re-pay for the target The most wasteful pattern is re-running the whole experiment because you thought of a new evaluator. The application's outputs have not changed; only the scoring has. `evaluate_existing` (and `aevaluate_existing`) applies evaluators to an experiment that already exists, so you pay for judge calls only. Making that a habit changes how you work: run the app once, score it as many times as your understanding of quality improves. ## Repetitions are a deliberate purchase `num_repetitions` multiplies both target and judge cost linearly. It buys variance estimates on a stochastic pipeline, which is a real thing to want — but it is the most expensive knob on the page, so turn it up only for the runs whose conclusion you intend to act on, not for every iteration. ## Instrument the spend The judge calls are themselves LLM calls, so trace them and attribute them like any other traffic. Knowing that the eval suite consumed a third of the month's token budget is what turns "we should optimise this someday" into a scheduled change. Tag the judge traffic distinctly so it does not silently inflate what looks like production usage. ## What not to cut Do not economise by removing the judged metric that actually reflects the user-visible failure and keeping only the cheap proxies. A suite of free assertions that all pass while the answers are subtly wrong is worse than expensive — it is misleading. Cut multipliers, not coverage of the thing you care about.
- Your eval suite costs more than the feature it evaluates. Is that automatically wrong?No. Before launch, or on a suite that gates a release, paying more to score than to serve can be entirely rational because the run happens a handful of times against thousands of production requests. It becomes wrong when the same expensive suite fires on every commit and nobody reads the result. Tie the spend to a decision someone acts on.
- How do repetitions interact with the cost formula?num_repetitions multiplies both halves: the target runs again per repetition and every judged evaluator runs again over each new run. Three repetitions triple the whole bill, not just the evaluation. That is worth paying when you need to know whether a difference exceeds the pipeline's own variance, and wasteful on every exploratory iteration.
- When does collapsing several judged metrics into one call backfire?When the criteria are unrelated, so the shared prompt becomes long and the model's attention is split, or when one metric changes often and every edit risks the others. A single parse failure also loses all the sub-scores at once. Collapse metrics that look at the same evidence; keep an independently evolving criterion separate.
saying these in an interview costs you the question
- Never multiplies out the number of judge calls before running
- Uses the largest judge model for every metric by default
- Re-runs the whole experiment to add one new evaluator
- Turns up num_repetitions without noticing it multiplies target cost too
- Replaces the meaningful judged metric with cheap proxies that always pass