skip to content

In an attacker-model jailbreak loop, one scoring model usually both tells the loop when to stop rewriting a prompt and decides which transcripts enter the report. Why does that double duty make its errors more damaging than the same error rate in an offline evaluation, and what would you change?

level: middleimportance: should knowfreq 52%

answer

  1. label is control flow, not just data
  2. errors compound across turns
  3. recall-biased signal, precision-biased gate
  4. second stage runs offline, no target calls
  5. persist score, not just the boolean

basics

~20 s

Because the scorer is not just measuring the run, it is steering it. A false success stops a thread that had not actually broken through; a false refusal keeps a solved thread spending budget or prunes a promising line. Fix it by separating the loop's stop signal from the report's gate.

solid answer

~60 s

In an offline evaluation the labels are applied to a fixed set of responses, so an error mislabels one item and the rest of the data is untouched. In a loop the label is consumed as control flow, so an error changes what gets generated next. That is the difference between noise added to a measurement and a bias in the search. Concretely: a false success ends a thread early, and everything that thread would have tried is never attempted. A false refusal keeps a thread alive past a real hit, so the budget is spent re-attacking something already broken, and in a branching search a wrong low score can prune the branch that was one rewrite from working. Errors therefore compound across turns rather than averaging out. The change is to stop using one component for two jobs. Use a cheap, deliberately recall-biased scorer as the loop's control signal so promising threads survive, then apply a stricter second-stage scorer plus human adjudication to the stored transcripts before anything is called a finding. Because the second stage runs offline over persisted text, it costs no extra target calls and can be re-run when the criteria change.

go deeper

for a junior

Should recognise the scorer's output is used both to stop the loop and to build the report, and that one error therefore has two effects.

for a middle

Explains that a control-flow label changes what gets generated next, so errors bias the search rather than averaging out, and proposes splitting signal from gate.

for a senior

Names the opposite cost asymmetries in the two roles, specifies the recall-biased in-loop signal plus precision-biased offline gate, and lists what must be persisted for re-scoring.

for a principal

Owns the trade: extra scoring cost and a longer triage queue against defensible findings, and sets where human adjudication sits in the pipeline.

The useful framing is that one scoring model is being asked to occupy two roles whose cost asymmetries point in opposite directions, and a single component at a single threshold cannot serve both. ## Role one: the loop's control signal Here the label answers *keep refining this thread, or stop?* The orchestrator reads it as control flow. That single fact is what separates this from an offline evaluation, where labels are applied to a fixed set of responses that already exist. Offline, a wrong label corrupts one row and the rest of the dataset is untouched — the error is noise on a measurement. In a loop, the attacker model's next prompt is conditioned on the label, so a wrong label changes *which responses come into existence*. Two runs of the same campaign under two scorer thresholds do not differ by a relabelling of the same data; they explore different regions of prompt space. Errors therefore compound down a thread instead of averaging out across it. In this role: - a **false refusal** is cheap in truth terms and expensive in budget — you keep attacking something already broken, burning attacker, target and scorer calls to the turn cap; - a **false success** is expensive in coverage — the thread stops and the attempts it would have made never exist. In a branching search the same error at a pruning decision removes a whole subtree rather than one item, so the loss is multiplied by the branch factor. Only one of those two is visible afterwards. The wasted turns show up as spend; the deleted subtree shows up as nothing. ## Role two: the report's gate Here the label answers *does this count as a finding?* and the asymmetry flips: | role | false positive costs | false negative costs | |---|---|---| | in-loop signal | lost coverage, unrecoverable | wasted turns, visible in spend | | report gate | triage hours, then credibility | a missed finding you never learn about | Wire one component at one threshold into both and you are forced to pick a compromise that is wrong for at least one job. Tuned for precision, the loop kills promising threads. Tuned for recall, the report ships unverified hits. ## The decoupled design Cheap, deliberately **recall-biased** scorer as the loop's stop signal, so threads survive; stricter, more expensive **precision-biased** scorer as a second stage over the persisted transcripts; human adjudication above that; and every attempt stored with prompt, full response, raw score (not the boolean), attack strategy, turn index and the reason the thread ended. The economics are what make this work. The in-loop scorer runs once per turn per thread, so its unit cost is multiplied by the whole search; a run with a few thousand attempts calls it a few thousand times, in-line with the loop's latency. The second stage runs once per stored attempt, offline, **and never touches the target** — no target tokens, no attacker tokens, no rate-limit backoff, no fresh authorisation to attack a live system. That is why the second stage can afford to be the expensive model with the long rubric, and why it can be re-run for free when the success criterion changes mid-engagement. ## What this costs you, honestly A second scoring pass over every stored attempt, not just the hits. Storage for full transcripts of a run whose non-hits outnumber its hits by one or two orders of magnitude. And a triage queue that is now *longer* than the raw hit list, by design, because the recall-biased signal deliberately admits more. Engineer hours move from re-running campaigns to adjudicating text. That is the trade, and it is the right one whenever the artefact is a report somebody acts on. ## Where the number misleads A decoupled pipeline produces two counts, and they will differ substantially. The in-loop signal's hit count is not a finding count and must never be quoted as one — it is a queue length. The failure mode in practice is a dashboard that reads the loop's counter because that is the number the tool emits, so the report inherits the recall-biased threshold that was chosen for search behaviour and had nothing to do with what counts as a violation. Equally, a low second-stage count is not evidence the run was thorough: it says nothing about threads the signal stopped early or pruned. ## What to check in a pipeline someone else built Is the same scorer object wired into both the termination condition and the reporting path? Is the raw score persisted or only the boolean — if only the boolean, any threshold change requires re-querying the target. Are non-hit transcripts kept at all? And does anything record *why* each thread ended, so that when you audit you can tell a scored stop from budget exhaustion? Without that last field the two most important failure modes look identical in the store.

  • Why can the second-stage scorer afford to be much more expensive than the in-loop one?
    It runs once over stored transcripts instead of once per turn inside the loop, and it never calls the target, so its cost does not multiply with the search.
  • What should the loop persist so a later re-score is possible?
    The full prompt and response for every attempt, hit or not, the scorer's label and raw score or confidence, the attack strategy, the turn index, and the reason the thread ended.
  • If you can only afford one scorer, which way do you bias it and why?
    Toward recall in the loop, so threads are not killed prematurely, and then accept that the finding list is a triage queue rather than a result.

saying these in an interview costs you the question

  • Treating scorer errors as random noise that averages out over a large run.
  • Wiring one scorer at one threshold into both the stopping rule and the report with no second stage.
  • Persisting only the boolean label, which makes threshold changes require a re-run against the target.
  • Claiming the fix is simply a better scorer prompt, with no separation of the two roles.

context