skip to content

questions

4

When does delegating a test's pass/fail verdict to a model beat an explicit assertion, and when is it strictly worse?

level: middleimportance: must knowfreq 46%

answer

  1. Something has to decide pass or fail
  2. Can you list every acceptable answer?
  3. Exact value present means no delegation
  4. Split: facts asserted, remainder judged

basics

~20 s

Delegate the verdict only when the acceptable result is a family you cannot enumerate: free-form text, wording shown to a user, perceived quality. Where an exact value or a structural rule exists, an explicit comparison is strictly better.

solid answer

~50 s

A judged verdict hands the produced output plus a written standard to a model and takes its answer as the outcome, instead of comparing against a value the author wrote down. It earns its place where the acceptable set is open — a generated summary, a support reply, *"does this error tell the user what to do next"* — because enumerating that set as assertions is impossible, so the honest alternative is no check at all. It is strictly worse wherever an exact value exists (amounts, identifiers, statuses), wherever a schema or invariant states the rule once, and wherever the outcome must be identical on every run, such as an access decision. The mature shape is a split: assert the hard facts explicitly, delegate only the remainder, and keep the judged question narrow and written down.

code

pseudocode · 16 lines
pseudocode
# one case, two decision layers
result = submit_refund(order_id = "ord-7741")

# hard facts: explicit, exact, reproducible
expect result.status equals "REFUNDED"
expect result.amount equals 42.50
expect result.currency equals order.currency
expect result.confirmation_id is not empty

# the open-ended remainder: one narrow judged question
verdict = judge(
    output   = result.customer_message,
    standard = "states the refunded amount and the reason,"
               " and names no customer other than this one"
)
record(verdict)   # reported alongside the case, not the thing that fails it

go deeper

for a junior

Recall the two ways a check can decide: comparing against a value the author wrote down, or asking a model to judge the output against a written standard. Be ready to say that a known amount or identifier is always compared, never judged.

for a middle

Explain the mechanics of the split: which parts of a response are hard facts under explicit assertion, which remainder is open-ended enough to delegate, and why a narrow judged question is more stable than a broad one.

for a senior

Show production judgement about where delegation is genuinely the only option versus where it quietly weakens a check. Be ready to argue that an open acceptable set, not effort saved, is the test for delegating a verdict.

for a principal

Own the policy: which categories of check may ever be judged, which are permanently off limits such as access outcomes and money, and how the team keeps judging standards reviewed like any other test asset.

## The two ways a check can decide Every automated check needs something that decides pass or fail. The usual arrangement is an **explicit assertion**: the author writes the expected value, or the expected shape, into the case and the runner compares. That comparison is total, repeatable and effectively free, and when it fails it points at the exact difference between what was produced and what was expected. A **judged verdict** replaces that comparison with a question put to a model. The check hands over the produced output together with a written standard — *"does this refund explanation state the amount, the reason and when the money arrives?"* — and takes the model's answer as the outcome of the check. Nothing is compared against a stored value; a judgement is made while the suite runs. Both decide, so both are oracles. What differs is where the decision comes from, how stable it is, and what you can do with the failure afterwards. ## Where delegating the verdict earns its place Delegation pays exactly where the acceptable answer is a **family you cannot enumerate**. - **Free-form language.** A generated summary, a support reply, a validation message written for a human: dozens of wordings are all correct and no equality comparison covers them. Asserting on an exact string here produces a case that fails on every harmless copy edit. - **Human-facing quality attributes.** *"Does this error tell the user what to do next"* and *"does this page name the customer whose order it is, and nobody else's"* are real acceptance criteria that no comparison against a stored value expresses. - **Sweeps nobody would otherwise write.** Asking one narrow question across four hundred rendered pages — *"does anything here read as unfinished: placeholder text, a truncated label, an untranslated phrase"* — is work that would never be written as four hundred assertions, so the realistic alternative is not a cheaper check, it is no check. - **Triage of a wide surface.** A judged pass over a broad area is weak evidence, but a judged fail is a useful pointer to where a human should look. The common thread: an explicit assertion would require you to enumerate the acceptable set, and the acceptable set is open. ## Where it is strictly worse Strictly worse, not merely more expensive, whenever an exact expected value or an expressible rule already exists. 1. **Known values.** Money, counts, identifiers, dates, statuses. If you can write `42.50`, write it. 2. **Structural rules.** Required field present, correct type, value within range, parts summing to a total. A rule stated once as a schema or an invariant is cheaper and stronger than a paragraph of instructions that re-describes it. 3. **Authorization outcomes.** Whether a request was allowed or refused must be decided the same way every time. A probabilistic answer to *"was this user permitted?"* is not an acceptable check. 4. **Anything a build gate depends on.** A verdict you cannot reproduce cannot be bisected, and a failure whose text differs each run cannot be deduplicated or tracked. 5. **Very hot paths.** A judged step costs latency and money on every execution; a comparison costs neither, and a suite running thousands of times a day feels the difference. | What the check decides | Explicit assertion | Judged verdict | | --- | --- | --- | | An exact known value | Right choice: exact, stable, free | Wrong: adds variance, buys nothing | | A structural rule | Right choice: stated once, enforced always | Wrong: an unstable restatement of a rule | | Open-ended prose | Cannot enumerate the acceptable set | Right choice | | Perceived quality for a human | Not expressible at all | Right choice, with a human reviewing findings | | An access decision | Required: must be identical every run | Never appropriate | ## The working rule: split the check, do not swap it The mature shape is not *assertions or judgement* but both inside one case. Assert every hard fact explicitly — status, amount, currency, identifier, required fields — and delegate only the remainder that genuinely resists enumeration. Two things follow. First, the judged question gets narrower, and a narrow question is far more stable than *"is this page correct?"*. Second, when the case fails you usually learn which layer failed, which is most of triage. Keep each judged question single-purpose and written down as an artefact of the suite, the way an expected value is. *"Is this output good?"* is not a check. *"Does this paragraph state a refunded amount and a reason, and mention no other customer's name?"* is one, and a reviewer can argue with it. The moment you find yourself listing the acceptable answers inside the instruction, you have discovered that an explicit assertion was available after all — write that instead.

  • You find yourself listing the acceptable phrasings inside the judging instruction. What does that tell you?
    That the acceptable set turned out to be enumerable, so an explicit comparison was available all along. Once you can list the accepted answers, write them as a membership check: it runs faster, costs nothing per execution, gives the same result every time, and its failure names the exact mismatch. A judged verdict is only justified while the list refuses to close.
  • Is a judged pass over a broad area worth anything at all?
    Very little on its own. A broad judged pass means only that nothing crossed one model's reading of one loosely worded standard, so it should never be recorded as coverage. A judged fail is worth much more: it is a cheap pointer telling a human where to look on a surface nobody had checks for. Treat the two directions asymmetrically.
  • Where does a judged verdict sit for an access decision?
    Nowhere. Whether a request was permitted or refused has an exact expected outcome and must be decided identically on every run, so it belongs to an explicit assertion. A model can still be useful next to that check — for instance judging whether the refusal message leaks internal detail — but the allow-or-refuse outcome itself is never delegated.

An explicit assertion is a lock and key: one shape fits. A judged verdict is a doorkeeper reading a written policy — useful where the guest list is open-ended, pointless where you already hold the key.

saying these in an interview costs you the question

  • Delegating a verdict when the exact expected value is known
  • Treating a broad judged pass as coverage
  • Asking one instruction to judge a whole page at once
  • Claiming judged checks can replace an assertion suite
  • Keeping the judging standard out of version control
open as a page

Why can a check that asks a model to judge a page's error text pass one run and fail the next on unchanged output?

level: juniorimportance: should knowfreq 38%

basics

~20 s

A judged step generates its answer by sampling, so identical input can produce different verdicts. Borderline outputs flip most, wording changes in the standard move verdicts, and the model doing the judging can be updated without a commit anywhere.

open as a page

Why should a model-judged verdict never be the only thing that fails a build, and what deterministic check sits behind it?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A build gate must be reproducible, reviewable and stable in its wording; a judged verdict is none of these. Let explicit assertions on statuses, amounts, identifiers, required fields and forbidden content fail the build, and report judged findings as signals.

open as a page

Before a suite trusts a model-judged check, how do you calibrate it against human-labelled examples?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Run the judged check over outputs people have already labelled against the written standard, then read every disagreement. Trust it only once it catches known past defects and the subtly-wrong near-misses, not just obviously good and obviously broken examples.

open as a page