skip to content

You need a PyRIT scorer for a success condition none of the shipped scorers covers. When do you write it as a deterministic rule over the response, and when do you back it with a model call?

level: middleimportance: should knowfreq 45%

answer

  1. checkable surface vs judgment of meaning
  2. canary, schema, tool call = rule
  3. third metered call per turn
  4. non-determinism breaks re-scoring
  5. cheap pre-filter, model on the middle

basics

~20 s

Write a deterministic rule when the success condition has a checkable surface: a specific string, a schema, a tool call, a leaked identifier you planted. Back it with a model only when success is a judgment about meaning. The rule is cheap, reproducible and re-runnable over stored transcripts; the model call adds cost, latency and run-to-run variation.

solid answer

~1 min

Start from the condition, not from the technology. Many bespoke red-team objectives are mechanically checkable: did the response contain the canary you planted in the document, did it emit a tool call with a particular argument, did it return valid JSON matching a schema, did it echo a value from a system prompt you control. All of those are rules, and a rule is the better scorer — deterministic, free, fast, and replayable over stored transcripts with identical results. A model-backed scorer earns its cost when the condition is semantic: did the response provide meaningful uplift toward the objective rather than a refusal wrapped in polite language, did it stay in a persona, did it describe a process at operational detail. No regex captures that. The tradeoff to state out loud: a model-backed scorer is a third metered call per turn on top of the attacker model and the target, it adds latency to a loop that already runs serially, and it is non-deterministic, so re-scoring the same transcript twice can produce different verdicts. That last point is the one that bites at report time, because your hit count stops being a fixed property of the transcripts. A common resolution is a cheap rule as a pre-filter with the model call only on responses the rule cannot settle.

go deeper

for a junior

Says a rule works when you can look for something specific in the response, and a model is needed when you have to judge meaning.

for a middle

Names the concrete cost of the model-backed choice: a metered call per turn, added latency in a serial loop, and non-deterministic verdicts on re-score.

for a senior

Designs the hybrid pre-filter, pins the scoring prompt as a versioned artefact, defines the error path so a failed scoring call is an unscored turn rather than a non-hit.

for a principal

Sets the default for the team as rule-first with planted canaries designed into the test setup, and requires an agreement measurement before any model-backed verdict is quoted in a deliverable.

### The decision rule Ask one question about the success condition: could a careful colleague settle this verdict with a text editor and a written spec, without interpreting anything? If yes, write a deterministic rule. If settling it needs a reading for intent, meaning or degree, you need a model. Many bespoke red-team objectives are mechanically checkable and people miss it. Did the response contain a canary token you planted in a document the target can read? Did it emit a tool call with a particular argument? Did it return JSON matching a schema? Did it echo a value you placed in a system prompt you control? Each of those is a substring test or a parse, and the rule is not a compromise — it is strictly the better scorer. ### What a deterministic scorer buys inside a run **Reproducible verdicts.** Re-score a stored transcript next year and get the same number. That is what makes a hit count a property of the evidence rather than of the afternoon you ran it. **No per-turn cost or latency.** The scorer runs on every turn of every conversation, so whatever it costs is multiplied by turns × conversations × objectives. **No new dependency in the loop.** A model-backed scorer that rate-limits, times out or errors mid-run turns into wrong verdicts unless you wrote the error path deliberately. **What it costs you instead:** rules are brittle exactly where a target's output varies — a paraphrase, another language, a code fence, a partial compliance, an answer delivered as a summary. Each of those is a false negative, and on an engagement a false negative is a finding you never counted. ### The cost of going model-backed, in numbers A 50-conversation run with an 8-turn budget is up to 400 scored turns. Rule-based: 400 string operations, microseconds, zero calls. Model-backed: 400 additional calls, each carrying the transcript so far — three metered calls per turn (attacker, target, scorer) instead of two. The money is rarely the constraint; at small-model rates a scoring call is a fraction of a cent. Two other costs bind: - **Wall clock.** The loop is serial within a conversation. One or two seconds of scoring latency per turn is minutes to tens of minutes added across a run, spent inside an authorised testing window. - **Variance.** Two scoring passes over identical stored transcripts can disagree, so the hit count stops being a fixed property of the evidence. ### The hybrid shape ```text response -> cheap deterministic checks planted canary present? -> hit, no model call known hard-refusal shape? -> non-hit, no model call anything else -> model-backed scorer decides ``` This keeps the metered call off the easy majority of turns and spends it on the ambiguous middle, which was the only place it bought anything. ### Where the number misleads **Scoring-side refusals.** A model-backed scorer is itself a safety-trained model being asked to read harmful content. When it declines, a scorer that maps "not a clear yes" onto false records a non-hit — an infrastructure failure laundered into a silent false negative, and precisely on the transcripts most likely to be real findings. Require a structured verdict and treat an unparsable scoring response as an *unscored turn* you can count, never as a non-hit. **Unmeasured variance.** A count produced by a non-deterministic scorer, quoted without a re-score check, implies a precision it does not have. Two runs of the same scorer over the same corpus giving 12 and 15 is not a target that changed. ### What to check - Score one fixed set of stored transcripts five times with the same model-backed scorer and count the verdicts that flip. That flip rate is your variance floor, and it belongs next to the count. - Pin the scoring prompt as a versioned artefact beside the code, request the lowest-variability sampling the endpoint offers, and require a structured verdict plus a rationale rather than free prose. - Count unscored turns explicitly and report them. A run with 30 unscored turns and 12 hits is a different result from 12 hits across 400 scored turns. - Measure agreement against hand-labelled transcripts from the run itself before quoting anything. For a model-backed scorer, agreement is the only evidence that the verdict means what you assume it means.

  • Your model-backed scorer's scoring call errors on some turns. What is the wrong way to handle it?
    Catching the error and returning a non-hit. That converts an infrastructure failure into a silent false negative. Surface it as an unscored turn so the count of unscored turns can be reported.
  • How does a planted canary change what the scorer needs to be?
    It moves the condition from semantic to mechanical: if you control a unique token in the document or system prompt, the scorer is a substring check, which is deterministic, free, and replayable.

saying these in an interview costs you the question

  • Defaulting to a model-backed scorer for a condition a planted canary would settle exactly.
  • Not mentioning that a model-backed scorer runs on every turn and multiplies by the turn budget.
  • Ignoring run-to-run variance, so the same stored transcripts produce a different hit count on re-score.
  • Leaving the scorer's error path undefined, so a rate-limited scoring call silently becomes a non-hit.

context