skip to content

When does an agent eval task need a judge model instead of a programmatic verifier?

level: seniorimportance: must knowfreq 55%

answer

  1. write the verifier before the task
  2. free, stable, path-agnostic
  3. assert what changed and what must not
  4. the grader is a model too
  5. gate in code, judge the residue

basics

~20 s

Use a programmatic verifier whenever success has a checkable end state — asserted database rows, a passing test suite, a valid schema. Reach for a judge model only for the residue that has no machine-checkable form, such as whether a drafted customer message was accurate and appropriately toned.

solid answer

~50 s

Design tasks so the verifier can be code. If the task is "refund order 4471 and notify the customer", assert the refund row exists with the right amount and status and that a message was queued — that check is deterministic, near-free, stable across repeated rollouts, and it survives the agent taking a different path to the same outcome. Repository-level coding suites use the same trick by running the project's tests as the verifier. A judge model earns its place only where the deliverable is open-ended prose or a policy judgement no assertion captures, and it costs you twice: model spend on every rollout, multiplied by k, and a second nondeterministic component whose scores you must first validate against human labels before you trust them. The practical pattern is hybrid — code gates the hard requirements, a judge scores only the free-text remainder.

go deeper

for a junior

Know the two kinds of verifier — code that checks the result, and a model that grades it — and that code is preferred whenever success can be checked by an assertion or a test run.

for a middle

Explain why programmatic checks scale: near-zero cost, stable scores across reruns, and indifference to which path the agent took. Name the cases they cannot cover, such as open-ended prose or tone.

for a senior

Show the hybrid design — assert hard requirements as a gate, judge only the free-text residue — and account for the judge's costs: spend multiplied by repeats, its own sampling variance, and the need to validate it against human labels before trusting it.

for a principal

Own the upstream consequence: what the suite can verify determines which product behaviours you can safely ship autonomously. Be ready to argue for specifying tasks so success is checkable, and for the budget split between cheap deterministic breadth and expensive judged depth.

## The verifier is the task A harness answers one question per rollout: did the agent succeed? Everything else — fixtures, sandboxes, tool fidelity — exists to make that question answerable repeatedly. So the verifier is not a detail bolted on afterwards; it constrains which tasks belong in the suite at all. A useful discipline is to write the verifier first. If you cannot state success as a check, either the task is under-specified or it does not belong on the deterministic tier. ## Programmatic verifiers A programmatic verifier is code that inspects the end state or the produced artifact. - **End-state assertions.** After the rollout, query the fixture: does the refund row exist, with the right amount, status and audit entry, and were no other orders touched? The last clause matters — a good verifier checks both the intended change and the absence of collateral damage. - **Executable checks.** Run the repository's test suite, compile the patch, validate the emitted JSON against a schema, execute the generated SQL. These are strong because the check is independent of how the agent phrased anything. - **Deterministic string or structural matching.** Weakest form, and brittle for agent output; prefer parsing to matching. *Why they are the default.* They cost essentially nothing, so they scale with k. They are stable: the same end state scores the same way on every rerun, which means a change in the score is a change in the agent, not in the grader. And they are path-agnostic — the agent may reach the correct end state by a different route than your reference, and a state assertion still passes where a path comparison would not. *Their limits.* They only see what they were told to look at. An agent can satisfy the assertions and still have done something unacceptable that no assertion covers, so state checks need to include negative conditions. And many real deliverables genuinely have no canonical form: a drafted apology email, a triage summary, an explanation of a policy decision. ## Judge models A judge model reads the agent's output against a rubric and scores it. It is the right tool when the artifact is open-ended and quality is the thing being measured. The costs are concrete and often understated: - **Spend and time.** Every rollout invokes the judge; k repeats multiply that. On a suite of 120 tasks at k=5, that is 600 judge calls per run, on top of the agent's own tokens. This is frequently what pushes a suite past its CI budget, and it is the main reason the pre-merge tier tends to be programmatic-only. - **A second variance source.** The judge samples too. Two identical agent outputs can receive different scores, which muddies exactly the reliability signal pass^k exists to measure. Judges are usually run at low temperature with a tight rubric for this reason. - **Unvalidated scores are not evidence.** A judge is a measuring instrument and must be calibrated against human labels on a sample before its numbers are used to make decisions, then re-checked when you change the judge model or the rubric. A judge nobody has checked can be confidently wrong for months. ## The hybrid pattern In practice most tasks decompose. "Handle this return request" has hard requirements (a refund of exactly this amount was issued, the return label was created, no other order was modified) and soft ones (the customer message was accurate, apologetic and did not promise anything outside policy). Assert the hard requirements in code as a gate; if the gate fails, the task fails and you never pay for a judge call. Score the soft residue with a judge only on the tasks that passed the gate. This keeps the expensive, noisy grader off the majority of rollouts and keeps the pass/fail decision anchored on something deterministic. A second useful split is by tier: programmatic-only on the fast pre-merge suite, judge-assisted on the nightly run where the budget is larger and latency does not block a developer. ## Designing tasks to be checkable Much of the skill is upstream. Prefer tasks whose goal terminates in a state change over tasks whose goal is a conversation. Specify the target precisely enough that a check can exist ("refund the full amount and set status to REFUNDED", not "make the customer happy"). Where a free-text artifact is unavoidable, extract the checkable parts — did the message contain the correct order number and refund amount — and judge only what remains. ## How to answer Lead with the default: code, because it is free, stable and path-agnostic, and it is why test-based coding suites work. Name what code cannot see, and admit the judge's three costs — spend times k, its own variance, and the need for validation against human labels. Then give the hybrid gate as the design you would actually ship. The weak answer defaults to a judge for everything because it is easy to write, and never mentions that the grader itself is an unvalidated model.

  • Why is asserting the end state usually better than comparing the agent's actions to a reference sequence?
    Because there is rarely one correct path. An agent may search before fetching, batch two lookups, or use a different but equivalent tool, and still land on exactly the right end state. End-state assertions accept all of those and reject genuinely wrong outcomes; sequence comparison penalizes harmless variation and makes the suite brittle to prompt changes. Path information is still worth capturing for diagnosis — it is just a poor pass/fail criterion.
  • A team's eval suite blows its nightly cost budget. How does verifier choice help?
    Judge calls scale with tasks times k and often rival the agent's own spend. Gate each task with programmatic assertions first and only invoke the judge on rollouts that pass the gate, so failures cost nothing extra. Move judge-scored tasks off the per-PR tier entirely, and consider judging a sampled subset of rollouts rather than all k when the judged dimension is a quality score rather than the pass/fail decision.
  • What is the risk of a programmatic verifier that only checks the intended change?
    It rewards agents that achieve the goal destructively. An agent that issues the correct refund but also cancels an unrelated order, deletes rows, or refunds twice passes a check that looks only at the target row. Verifiers should assert invariants as well as outcomes — nothing outside the expected diff changed, totals still reconcile, no duplicate transactions — which also converts a whole class of silent production incidents into eval failures.

saying these in an interview costs you the question

  • Defaulting to a judge model because writing assertions is more work
  • Trusting judge scores that were never checked against human labels
  • Verifying only the intended change, ignoring collateral damage
  • Ignoring that judge cost multiplies by tasks times repeated rollouts
  • Treating exact-match on free text as a rigorous verifier

context