skip to content

In your AI red-team harness, whether an attempt counted as a failure was decided by an automated model-based grader with a threshold. The application team will re-run the reproduction without that grader. What do you put in the handoff so their verdict matches yours?

level: seniorimportance: should knowfreq 34%

answer

  1. the judgment does not travel
  2. rule in prose plus threshold
  3. per-case scores show margin
  4. deterministic proxy for their CI
  5. spot-check the grader by hand

basics

~20 s

Translate the grader into something they can apply. Write the decision rule in plain words with its threshold, attach the grader's output on the cases you sent, include labelled examples just either side of the line, and give a deterministic check they can assert on. If the verdict truly needs the grader, ship it pinned as a small dependency.

solid answer

~60 s

The grader is the part of your instrument the developer cannot see, and it is the part that decided the finding exists. Hand it over in three layers: 1. **The rule in prose.** What the grader judged and where the threshold sat — "counts as a failure if the response provides actionable steps toward the objective, scored above 0.7". A developer can apply that by eye. 2. **The grader's output on the attached cases.** A score per transcript, so they see the margin; a case that scraped past the threshold should be flagged as such rather than presented as a clear hit. 3. **A deterministic proxy for the regression test.** Where possible, pin an observable in the output that a plain assertion can check with no model in the loop. Model-graded assertions in a team's CI are flaky and get muted. Also disclose the grader's known failure modes and any disagreement rate you measured against your own reading. If the verdict truly cannot be reduced, ship the grader pinned, with its prompt and threshold.

go deeper

for a junior

Knows that something automated decided the attempt was a failure and that the developer needs to be told what that criterion was.

for a middle

Ships the rule, the threshold and the per-case scores, and can explain why a bare verdict is not portable.

for a senior

Reduces the verdict to a deterministic assertion where possible, spot-checks the grader by hand, and discloses its known biases and disagreement rate.

for a principal

Decides where model-graded verification is allowed to live — the red-team harness versus the product team's pipeline — and what evidence a finding must carry before it can be filed at all.

This is where evidence packages quietly fall apart. The transcript travels fine; the *judgment* does not. The team re-reads the same response, sees a hedged, partly-refusing answer, and concludes it is acceptable — because nothing in the package told them what standard was applied. ## What a grader actually is A model-based grader is a second model call. It receives the response (often the prompt and a rubric too) and returns a score, and a threshold turns that continuous score into a binary verdict. The shape recurs across tools: a garak detector returns a per-attempt score in the 0.0–1.0 range and the report's pass/fail follows from it; promptfoo's `llm-rubric` and its other model-graded assertions take a `threshold` in the test's assert block; PyRIT's self-ask scorers return either a true/false or a float on a scale defined inside the scorer's own prompt. In every case there is a prompt you wrote, a judge model with a version, and a cut point you chose. All three are invisible in the verdict, and all three are load-bearing. ## Three layers of handoff **The rule in prose, with its threshold.** "Counts as a failure if the response provides actionable steps toward the objective, scored above 0.7." A human can apply that to a response their fix has newly produced, which is what verification actually requires. **The grader's output on the attached cases.** A score per transcript, so the reader sees the margin. A case that scraped past the cut should be flagged as marginal rather than presented as equivalent to a clear one. Scores without the rule are numbers with no units; the rule without scores hides that some of your evidence was borderline. **A deterministic proxy for the regression test.** Where the failure has an observable, mechanical signature, hand that over and let the model-graded version stay in your harness. Where it has none, say so explicitly and agree what re-verification looks like, rather than letting it default to nothing. ## What it costs Model-graded evaluation roughly doubles the call count of a run — one call to the target, one to the judge — and the judge call is often the more expensive of the two, because its input carries the whole response plus the rubric. A twenty-trial re-verification is forty calls. Put that in the application team's pipeline and a suite of thirty graded assertions is thirty target calls and thirty judge calls on every run, priced per run, against a hosted judge that can reversion without notice. The first time it flakes red on an unchanged pull request it gets marked allowed-to-fail; the second time it is deleted. Engineer time is the larger half of that bill. This is the whole argument for pushing a deterministic assertion into their pipeline and keeping the judge in yours. ## Where the number misleads A score of 0.81 against a 0.7 threshold looks decisive. It is a continuous score whose errors concentrate exactly at the boundary you drew. Model graders carry well-documented systematic biases: longer and more fluent answers score higher; hedged, polite, refusal-shaped text scores lower even when the operative content is present; formatting and ordering shift the read; and a judge drawn from the same model family can favour that family's outputs. The characteristic miss in safety work is the difference between a response that *describes* a hazard and one that operationally *enables* it — a rubric that does not draw that line will flag the first, and that is the attachment the developer will dispute. The aggregate inherits all of it. If your judge agrees with careful human marking on 85% of cases, your measured 35% rate carries an error the report does not display, and those errors are not randomly scattered: they cluster near the threshold, which is precisely where your weakest attachments sit. Push-back always comes on the weakest attachment, never the strongest, so a rate you have not spot-checked is a rate you cannot defend. ## What to check before you send Read every attached response yourself, without looking at the score, and mark it; any attachment where your judgment differs from the harness's verdict comes out of the package or ships with the disagreement noted. Then measure the grader rather than trusting it: hand-mark a blind sample drawn from *both* the flagged and the unflagged responses. Sampling only the flagged ones measures precision and leaves the grader's misses completely invisible. Report the disagreement rate in both directions, along with the judge model and version, so the number the team inherits comes with its error bar attached.

  • The developer disputes one attachment, saying the response only described a hazard rather than enabling it. What now?
    Adjudicate against the written rule, not the score. If the rule is ambiguous on that distinction, that is a defect in the rule — sharpen it, re-mark the attachments, and drop the ones that no longer qualify.
  • Why not just give the team your grader to run?
    You can, pinned with its prompt and threshold, but a model in their test loop costs money per run, drifts with the hosted endpoint, and produces flaky verdicts. Prefer a deterministic assertion where one exists.

A grader score handed over without its rule and threshold is an exam mark with no exam: 0.81 means nothing until you say what was asked and where the pass mark sat.

saying these in an interview costs you the question

  • Handing over a verdict with no statement of what standard produced it.
  • Attaching scores as bare numbers with no rule or threshold.
  • Proposing a model-graded assertion as the team's permanent regression test without discussing its flakiness and cost.
  • Never having read the flagged responses by hand.
  • Presenting a borderline, just-over-threshold case as equivalent to a clear one.

context