skip to content

In an automated jailbreak loop where an attacker model rewrites prompts, a target model answers, and a separate scoring model labels each answer a success or a refusal, what happens when the scoring model wrongly labels a refusal as a success?

level: juniorimportance: must knowfreq 70%

answer

  1. label is also the stop signal
  2. thread abandoned early
  3. hit enters the finding list
  4. triage cost scales with FP rate
  5. machine-labelled vs human-confirmed

basics

~20 s

The loop treats that thread as solved and stops rewriting, so the search ends early. The mislabelled refusal is written into the finding list as a discovered attack, and a human triager has to read the transcript to throw it out. Your reported success count is inflated by every such label.

solid answer

~50 s

A wrong success label costs you twice. Inside the run, the loop's stopping rule fires: the attacker stops refining that line, so a prompt that was still one or two rewrites away from a genuine break is abandoned. That is a silent loss of coverage nobody sees in the output. Outside the run, the transcript lands in the finding list marked as a hit. Everything downstream inherits that label: the attack-success rate you quote, the per-category counts, any comparison against a previous run. Someone has to open each transcript and read what the target actually said before the report goes out, and that triage cost scales with the false-positive rate, not with the number of real findings. The practical defence is that the tool's hit is not the report's finding. Treat the scorer's label as a candidate queue for human review, and keep the full transcript of every attempt so a hit can be re-scored later without re-querying the target.

go deeper

for a junior

Should say the wrong label puts a non-finding into the report and inflates the success count, and that someone has to read transcripts to remove it.

for a middle

Adds that the label is also the loop's stopping rule, so a false success truncates the search on that thread as well as polluting the output.

for a senior

Talks about measuring the false-positive rate on a sample, keeping transcripts for offline re-scoring, and separating machine-labelled counts from human-confirmed ones in the report.

for a principal

Frames it as report credibility: what number leaves the team, who signs it, and what the sign-off process requires before a machine label becomes a customer-facing finding.

## The pipeline, part by part An automated jailbreak loop has three model roles and one piece of bookkeeping. The **attacker model** is a language model prompted to write a candidate prompt and then rewrite it in light of what came back. The **target** is the system under assessment, usually a hosted chat endpoint reached over an API and billed per token. The **scoring model** reads the target's response and emits a label, success or refusal; PyRIT calls that object a *scorer*, garak calls the nearest equivalent a *detector*, promptfoo calls it a *grader*, and a benchmark calls it a *harm judge*. They are four different implementations occupying the same slot, so always say whose. The **orchestrator** is the loop driver, and it consumes the label twice: once as the thread's stopping rule, once as the row it writes into the results store. Three different things wear the word *success* in this pipeline and they are not the same object: | what | who produces it | what it means | |---|---|---| | the label | the scoring model | a classifier's guess about one response | | the finding | your triage process | an item a human is willing to sign | | the violation | the target's actual behaviour | ground truth, never directly observed | A false positive is the case where the first two agree and the third does not. ## Why scoring models produce them so readily The surface features of compliance are easy to imitate without any prohibited content being present. An answer that opens agreeably and then declines. A refusal that restates the request in full before refusing, so the response literally contains the request's vocabulary. A fiction or roleplay frame in which a character discusses doing the thing while the text carries nothing operative. A confident, generic, useless paragraph. A safety lecture that quotes what was asked. If the scorer's prompt asks *did the model comply?* rather than *does this response contain the specific prohibited content, judged against this written description of what would count?*, every one of those reads as a hit. ## What the error costs Two bills, and only one of them appears anywhere. Inside the run, the label is the stopping rule. A false success terminates that thread, so its remaining turns are never spent. On a ten-turn cap, a thread cut at turn two returns eight turns of attacker, target and scorer calls unspent — roughly two dozen model calls per mislabelled thread. That is money saved and coverage lost, and the saving lands in the invoice while the loss lands nowhere. A campaign with a lenient scorer therefore comes in *under* budget and looks efficient. Outside the run, the transcript enters the finding list. Somebody has to open it and read what the target actually said. Adjudicating one transcript against a written criterion is a few minutes of an experienced person's time, so triage time scales with the false-positive rate rather than with the number of real findings — a sixty-hit list at a 40% false-positive rate is a couple of hours of senior time spent producing nothing but deletions, and that cost was never in the estimate. ## How the number misleads The headline reads *sixty successful attacks*. A reader hears sixty confirmed violations. What sixty actually is: a count of machine labels. Confirmed violations can only be fewer; real successes the run produced could be more, because the misses are not in the number at all and never will be. So the same lenient scorer makes the target look **more broken** (inflated hits) and **less explored** (truncated threads) at the same time, and only the first half is visible in the artefact. The second misreading is comparative. Quoting this quarter's sixty against last quarter's forty is only a statement about the target if the scorer, its prompt and its model version were identical. Change any of them and the delta you are reporting is a property of your instrument. ## What to check before you believe it Pull a random sample of labelled hits and hand-label them yourself against a criterion you wrote down first. Then look at *where* the scorer fired: hits clustered on turn one, or concentrated on a single attack template or harm category, usually mean it is matching a phrasing pattern rather than content, and that pattern tells you which of the remaining hits to distrust. Run the scorer over a small fixed set of hand-labelled responses that deliberately includes near-misses — an agreeable-sounding refusal, a partial answer, an off-topic compliance — before the campaign rather than after. And keep the full transcript of every attempt, hit or not, with the raw score and not just the boolean, so a later re-score costs stored text instead of fresh target calls. Finally, say plainly in the report which count is machine-labelled and which is human-confirmed. Those are two numbers, and a reader will assume the smaller one unless you tell them.

  • Why is a false success worse in a multi-turn loop than in a single-shot scan of fixed prompts?
    In a single-shot scan the wrong label only mislabels one attempt. In a loop it also stops the search on that thread, so you lose the attempts that would have followed.
  • How would you keep the option to re-score a run without paying for the target calls again?
    Persist the full prompt-and-response transcript of every attempt, hit or not, with the scorer's label and confidence. Re-scoring then runs offline against stored text.
  • What is the cheapest sanity check on a scorer before a large run?
    Run it over a small hand-labelled set that deliberately includes agreeable-sounding refusals and partial answers, and look at how it labels those.

The scorer's label is both the scoreboard and the referee's whistle. When it awards a point that was never scored, it also blows play dead on the move that was about to score a real one.

saying these in an interview costs you the question

  • Treating the tool's hit count as the report's finding count with no triage step.
  • Assuming a strong general-purpose model makes a reliable scorer without any check against hand labels.
  • Not realising the label doubles as the loop's stopping criterion.
  • Discarding non-hit transcripts, which makes any later re-scoring impossible.

context