skip to content

When does an LLM judge over the full trace beat reference-trajectory matching?

level: seniorimportance: should knowfreq 38%

answer

  1. some qualities have no string form
  2. one vague score versus five narrow criteria
  3. cheap deterministic checks come first
  4. the judge is itself an instrument
  5. agreement with humans, measured not assumed

basics

~20 s

Use a judge when the quality you care about cannot be expressed as string comparison — was the plan sensible, did the agent recover well, did it stop at the right time — or when the task has too many valid paths to enumerate references for.

solid answer

~50 s

Reference matching is exact, cheap and deterministic, but it only measures conformance to a path someone wrote down. A judge model reading the whole trace can score properties that have no string form: whether the decomposition was sensible, whether the agent recovered gracefully from a tool error, whether it stopped as soon as it had enough evidence, whether it respected a stated policy. It also scales to open-ended tasks where enumerating reference trajectories is impractical. The cost is that the judge is itself a stochastic, fallible model, so treat it as an instrument to be calibrated: define a rubric of a handful of narrow criteria rather than one vague quality score, require a per-criterion verdict, and measure agreement against a set of human-labelled traces before you trust it. The standard arrangement is layered — deterministic checks for everything expressible in code, the judge only for the residue.

go deeper

for a junior

Know that an LLM judge can read an agent's whole trace and score it against written criteria, and that this is used where a simple string comparison cannot capture what good looks like.

for a middle

Be ready to say what a judge adds over reference matching — plan quality, recovery, stopping behaviour, policy adherence — and why a rubric of several narrow criteria beats a single quality rating.

for a senior

Show the engineering discipline: deterministic checks first, judge only for the residue, per-criterion binary verdicts with justifications, agreement measured against hand-labelled traces, judge model and rubric pinned and versioned.

for a principal

Own the cost and trust model. Decide how much of the eval budget a judge may consume, when judge-graded criteria are allowed to block a release rather than inform it, and how you keep scores comparable across judge upgrades.

## What a judge buys you Reference-trajectory comparison answers one question well: did the agent follow an acceptable path someone wrote down. Whole classes of trajectory quality escape it, because they are judgements about *reasoning* rather than about token equality: - **Plan quality** — was the decomposition sensible, or did it work in an order that made later steps harder? - **Error recovery** — after a tool returned an error, did the agent adapt, or retry the identical call? - **Evidence sufficiency** — did it stop when it had enough, or keep gathering after the answer was determined? - **Policy adherence** — did it follow stated operating rules, such as confirming identity before disclosing account details? - **Faithfulness of the narration** — does the agent's stated reasoning match what its tool calls actually did? A judge model that reads the full trace — user goal, every call, every observation, the final answer — can score all of these. It also scales where references do not: for open-ended tasks with a combinatorial space of valid approaches, authoring and maintaining reference trajectories costs more than it returns, while a rubric written once applies to every task in the suite. ## Rubric design is where the quality comes from The single biggest determinant of judge usefulness is the rubric. A prompt asking "rate this trajectory 1-10" produces a number that is neither reproducible nor actionable. What works is a small set of narrow, independently checkable criteria — five is a common size — each with explicit pass conditions and an example of failure, scored separately, with a short justification and a citation to the step in the trace that drove the verdict. Per-criterion scores are diagnosable in a way a scalar is not: when a release regresses, you can see that recovery-quality fell while plan quality held. Two refinements matter. **Per-step (process) rubrics** score each turn as it happened rather than the episode as a whole, which localizes failure and produces a much richer signal on long trajectories; whole-trace rubrics are cheaper but blur where things went wrong. And **binary criteria beat graded ones** — "did the agent verify the record existed before updating it: yes/no" is far more stable across runs than "rate the verification thoroughness 1-5". ## Calibration is not optional A judge is a measuring instrument, and an uncalibrated instrument produces confident numbers about nothing. Before letting judge scores influence a decision, label a sample of traces by hand — sixty is a realistic working size — spanning passes, failures and ambiguous cases, and measure how often the judge agrees with the humans, per criterion. Criteria with poor agreement are usually badly specified rather than badly judged, and rewriting them is the fix. Recalibrate whenever you change the judge model, the rubric or the agent's output format, because all three shift the distribution the judge sees. Nail down what you can: fix the judge's sampling to be as deterministic as the provider allows, pin the model version, and keep the rubric under version control alongside the eval suite, since a silent rubric edit invalidates every historical comparison. ## Where a judge is the wrong tool Anything checkable in code should be checked in code. If the criterion is "the agent called the retrieval tool before answering", a five-line assertion is cheaper, faster, perfectly reproducible and never drifts. Reaching for a judge there adds cost, latency and noise for nothing. A judge is also poorly suited to fine numeric distinctions — it cannot reliably tell a 12-step trajectory from a 14-step one, which is what a counter is for. The layered arrangement most teams converge on: deterministic outcome check first, then programmatic trajectory constraints (required calls, forbidden calls, ordering, budgets), then the judge rubric over whatever residue is genuinely a matter of judgement. Each layer is cheaper and more reliable than the one after it, so you want as much weight as possible on the early layers. ## Reading judge output honestly Treat judge scores as a noisy estimate with a confidence interval, not a measurement. Small movements between runs are usually noise; act on trends across a suite, not single-task deltas. Keep the judge's per-criterion justifications, because they are the artifact a human actually reviews during triage — a score with no reasoning is unauditable, and an audit trail is how you notice that the judge has started rewarding verbosity rather than correctness.

  • How do you decide which criteria belong in the rubric versus a deterministic check?
    Ask whether the property can be decided from the trace by code without interpretation. Required calls, forbidden calls, ordering, step budgets and schema validity all can, so they go in code — cheaper, faster and exactly reproducible. The rubric takes only what needs reading comprehension: plan quality, recovery behaviour, evidence sufficiency, policy adherence. If a criterion keeps getting poor human agreement, it is often a code check in disguise, badly specified.
  • Why prefer per-step rubrics over one score for the whole trajectory?
    A whole-trace score tells you an episode was mediocre; a per-step score tells you which turn broke it, which is what you fix. Per-step grading also degrades less on long trajectories, where a single scalar gets dominated by whatever the judge read last. The tradeoff is cost — you are running the judge once per turn — so many teams grade every step on a sampled subset and whole-trace on the rest.
  • What breaks when you swap the judge model for a newer one?
    Comparability with every historical score. Different models weight rubric criteria differently, so a suite can appear to improve or regress with no change to the agent. Treat the judge model version as part of the eval configuration: pin it, re-run the human-agreement check after any change, and rescore a baseline window with the new judge before comparing across the switch.

saying these in an interview costs you the question

  • Uses a judge for checks that a simple assertion would decide
  • Asks the judge for one overall quality score out of ten
  • Trusts judge scores without ever measuring human agreement
  • Uses the same model and prompt as the agent under test
  • Treats a small judge-score delta between runs as a real regression

context