skip to content

Your team runs PyRIT campaigns against several internal endpoints on a recurring schedule. What should the scorer be allowed to decide on its own, and what has to reach a human before it counts?

level: principalimportance: should knowfreq 40%

answer

  1. autonomous inside the run, reviewed outside it
  2. ranking beats verdict accuracy
  3. frozen labelled set = the scorer's regression test
  4. pick the error direction per campaign type
  5. never automate 'no findings'

basics

~20 s

Let the scorer decide only inside the run: when to stop a conversation and how to rank what gets read first. Anything leaving the team as a finding needs a human who read the transcript. Keep a fixed labelled transcript set as the scorer's own regression test, re-run whenever the scorer changes.

solid answer

~60 s

Draw the line at the boundary of the run. **Inside the run** the scorer has to be autonomous - it is the loop condition, and no one is going to adjudicate every turn. It also earns the right to rank: ordering conversations by score decides what a human reads first, and getting that ordering roughly right is worth far more than perfect verdicts. **Leaving the run**, its verdict is a candidate, not a finding. A hit that goes into a report, a ticket or a risk register should carry the transcript and a name of someone who read it. That is not ceremony - the scorer's failure mode is precisely the response that looks like compliance, and the cost of one wrong finding landing on another team is measured in trust. Around both, the scorer needs its own test: a frozen set of hand-labelled transcripts, re-scored whenever the scorer, its backing model or its rubric changes, before the new configuration is allowed to end runs. Without that, a silent change to a component quietly redefines what the program calls a success.

go deeper

for a junior

Should say a scorer verdict is not by itself a finding and that someone should read the transcript.

for a middle

Should place the boundary at the run - autonomous as a stop condition, reviewed once it leaves - and mention keeping the transcript with each hit.

for a senior

Should add the frozen labelled regression set for the scorer and the split between a cheap in-loop stop rule and a richer offline verdict.

for a principal

Should own the tradeoff explicitly: error direction chosen per campaign type, a named owner for the scorer configuration, ranking as the realistic ask, and no automated null results.

## The authority boundary Automate the decisions whose errors are **cheap and self-correcting**; keep a human on the decisions whose errors **travel**. - **Stopping a conversation is cheap** - if the scorer got it wrong you re-run that objective for the price of one conversation. - **Declaring a finding is not cheap**: it moves to another team, becomes a remediation ask, consumes their sprint, and is remembered for a long time when it turns out to have been a scorer artefact. So inside the run the scorer must be autonomous, because nobody is going to adjudicate every turn of every conversation. Crossing the run boundary, its verdict is a *candidate*, and what leaves the team carries the transcript and the name of a person who read it. ## Ranking is the underused middle A recurring program produces more transcripts than anyone will read. Work the arithmetic honestly: 300 objectives a week, several turns each, is thousands of responses; a realistic human budget is an hour or two, which is perhaps forty to sixty transcripts. The realistic ask of a scorer is therefore not "be right on every verdict" but "put the forty conversations most worth reading at the top". That reframing lowers the stakes on individual verdicts and raises the value of persisting **raw scores** rather than only booleans, because a boolean cannot sort. PyRIT's **human-in-the-loop scorer** fits here too: route only the ambiguous band to a person instead of everything or nothing. ## Separate the stopping decision from the recorded one They need not be the same component. A cheap, deliberately conservative rule can drive the loop while a richer scorer, replayed offline over persisted conversations, produces the verdicts and the reading order the report is built from. That keeps expensive judgment off the hot path of every turn and gives you a second, independent opinion on every candidate hit - for the price of scorer calls only, since no target turns are repeated. **Auxiliary scorers** can annotate in-flight when you want the label recorded but not acting. ## Choose the error direction per campaign type - A **broad coverage sweep** across many endpoints can tolerate false negatives: you are deciding where to dig, and a missed hit costs a bill and a later re-run. - A **deep engagement** on one high-value target cannot tolerate false positives, because every premature stop abandons a line of attack you had budget for. Set the stop rule accordingly, and write down which way you biased it and why, because that choice - not the target - is what makes two campaigns' numbers non-comparable. ## Give the scorer an owner and a test Name who owns the scorer configuration, and hold a **frozen, hand-labelled transcript set** drawn from your own runs as its regression suite. Re-run it before any change to the scorer, its rubric or its backing model is allowed into campaigns. The cost is one scorer call per item plus the labelling you already paid for once; the cost of skipping it is that the definition of "success" in your program drifts with a component nobody is watching, and week-over-week movement becomes unattributable. This is the single most common way a red-team metric silently stops meaning anything: the judge model gets upgraded and the trend line is read as a change in the systems under test. ## Where the numbers mislead at program scale Three readings to refuse. 1. First, a trend line across a scorer or rubric change - it measures the scorer. 2. Second, a league table of endpoints scored by different scorers or objectives - the denominators are not the same. 3. Third, and most dangerous, "no findings this week": a blind scorer and a genuinely resistant target produce identical runs, because both exhaust their budgets. A null result must clear the same evidence bar as a positive one. ## What you would check before a report leaves the team - That the scorer still recognises known-compliant responses, using the frozen set - that is the smoke test that makes a null result publishable. - That every hit going outward has a transcript and a reader. - That early-turn successes were reviewed. - And that nothing in the scorer's configuration changed since the last number you published, or that the change is stated next to the delta. ## What I would not automate - Turning a hit into an external report; - comparing endpoints on scorer output alone; - and any claim of "no findings".

  • Why is 'no findings this week' the most dangerous automated output in this setup?
    Because a blind scorer and a genuinely resistant target produce identical runs - both exhaust their budgets. A null result should require the same evidence as a positive one, starting with proof the scorer still recognises known-compliant responses.
  • Does the component that stops the run have to be the one whose verdict is reported?
    No, and separating them is usually better. A cheap conservative rule drives the loop; a richer scorer replayed over persisted transcripts produces the recorded verdicts and the reading order, giving a second independent opinion off the hot path.
  • How would you bias the stop rule differently for a broad sweep versus a deep engagement?
    A sweep can tolerate misses - a missed hit costs calls and you are only choosing where to dig. A deep engagement cannot tolerate false hits, because each premature stop abandons a line of attack, so the rule there should be tight and early successes should be reviewed.

saying these in an interview costs you the question

  • Ships scorer output straight into tickets or a risk register with no transcript and no reader
  • Treats the scorer as infrastructure with no named owner and no test of its own
  • Uses the same stop rule for a broad sweep and a deep single-target engagement without thinking about it
  • Reports 'no findings' on a clean run without checking the scorer can recognise success at all
  • Swaps the scorer or its backing model mid-programme and keeps comparing the numbers as if nothing changed

context