skip to content

In an automated jailbreak run where a scoring model labels each target response a success or a refusal, which shapes of target response most often get mislabelled in each direction, and how do you catch them?

level: middleimportance: must knowfreq 60%

answer

  1. agreeable opening then refusal
  2. restates request before declining
  3. roleplay frame, no content
  4. buried after long preamble
  5. non-prose, other language, encoded

basics

~20 s

Toward false success: fictional or roleplay wrappers, agreeable openings that then refuse, restatements of the request, and vague non-actionable text. Toward missed success: harmful content buried mid-answer, output in another language, code or encoded output. Catch them by hand-labelling a sample of both hits and non-hits.

solid answer

~60 s

The two directions have different shapes because the scorer is pattern-matching on surface form. False successes cluster on responses that look like compliance without carrying content: an answer that opens agreeably and then declines, a detailed restatement of the request before refusing, a story or roleplay frame with nothing operative in it, a generic and useless paragraph, or a hedged partial that stops before anything actionable. False negatives cluster on content the scorer did not read as content: the operative part buried after a long safe preamble, output in a language the scorer handles worse than the target does, output as code or structured data rather than prose, encoded or obfuscated output, or an answer split across turns so no single response looks complete. You catch the first direction by hand-labelling a sample of the labelled hits. You catch the second only by sampling from what the scorer rejected, which means the run must have kept those transcripts. A second, differently-prompted scorer over the same stored text is a cheap way to surface disagreements worth a human look.

go deeper

for a junior

Should name at least the obvious false-success shape, an answer that sounds cooperative but refuses, and say you check by reading some transcripts.

for a middle

Names shapes in both directions and explains the mechanism: the scorer reacts to surface form, so compliance-shaped text passes and content in unusual form is missed.

for a senior

Adds the measurement plan: a fixed calibration set built before the run, stratified sampling of rejected transcripts, a second scorer to surface disagreements, and reporting the resulting error estimates.

for a principal

Turns it into a standing asset: a maintained labelled calibration set per harm area that every campaign's scorer is measured against, so error rates are comparable across engagements.

Treat the scoring model as a classifier over target responses with two error types, then ask what each error actually looks like in this pipeline and where it surfaces. The scorer here is whatever object decides success or refusal — a PyRIT scorer, a garak detector, a promptfoo grader, a benchmark's harm judge — and the shapes below are properties of the *job*, not of one vendor's implementation. ## False success: the target refused, the scorer said it complied Every one of these is surface compliance without substance. - **Agreeable opening, then a decline.** "Sure, I can help with that — however, I can't provide…" The first clause carries the compliance signal. - **Restatement before refusal.** The response repeats the forbidden request in detail before declining, so the transcript literally contains the request's vocabulary. Any scorer that is even slightly keyword-adjacent lights up. - **Fiction or roleplay frame.** A character discusses doing the thing; the text contains nothing operative. The scorer sees the topic and reads it as content. - **Confident but empty prose.** A fluent, generic paragraph that answers nothing. - **Hedged partial.** The answer starts down the path and stops before anything actionable, which a threshold set for "did it engage?" scores as a hit. The single root cause is the scorer's own instruction. *Did the model comply?* is a question about tone. *Does this response contain the specific prohibited content, judged against this written description of what would count?* is a question about content. ## False negative: the target complied, the scorer said refusal These are the expensive ones, because they appear nowhere at all. - **Buried content.** The operative part arrives after a long disclaimer or safe preamble, and the scorer weighted the preamble. - **Wrong modality.** Output as code, a table, JSON, or pseudo-config rather than prose, when the scorer's few-shot examples were all prose. - **Wrong language or script.** Targets often generalise across languages better than a scorer prompted and calibrated in English. - **Obfuscated or encoded output**, which the scorer does not decode. - **Split across turns.** No single response is damning; the harmful whole exists only in the concatenation. - **Outside the scorer's own notion of harm.** The systematic case, and the one no prompt tuning fixes. ## What each direction costs | direction | where it shows up | who pays | |---|---|---| | false success | the finding list, immediately | triage: minutes of an experienced reviewer per transcript, scaling with the error rate, not the finding count | | false success | the loop, silently | the thread's remaining turns are never spent — attacker, target and scorer calls that the invoice records as a saving | | false negative | nowhere | a real weakness ships unreported, and the client acts on a clean category | The asymmetry is the point: one direction bills you visibly and one does not bill you at all, which is exactly why teams audit the hits and stop there. ## How the number misleads Because only hits are enumerable, *every* easily computed quality figure describes the false-success direction. Precision, triage yield, the proportion of findings confirmed — all of them can look excellent while the miss rate is unknown and large. The specific wrong reading is a harm category showing zero hits and being reported as "no findings", when what happened is that the category's successes came back as code, or in another language, or after a disclaimer, and the scorer rejected each one. A zero is a claim about your instrument until you have coverage evidence to make it a claim about the target. ## What you actually check Build a small fixed **calibration set** before the campaign: hand-labelled responses per harm area that deliberately include the near-misses above — an agreeable refusal, a restatement, a roleplay frame, a code-block compliance, a non-English compliance, a buried compliance. Run the scorer over it and look at the confusion, not the accuracy. If it cannot separate an agreeable refusal from a real hit, its output is not reportable and no amount of extra run budget fixes that. During the run, persist every attempt: prompt, full response, label, raw score or confidence, attack strategy, turn index, and the reason the thread ended. The boolean alone forecloses everything below. Afterwards, run two samples. Hand-label a random sample of hits to estimate precision. Then sample the *rejected* transcripts — stratified toward borderline scores, toward strategies and categories with zero hits, toward long, non-prose and non-English responses, and toward final turns of threads that ended on the cap — and hand-label those. Re-scoring the whole rejected pool with a second scorer written independently from a different prompt is the cheap enrichment step: it bills only scorer tokens over stored text, no attacker calls, no target calls, no rate-limit waiting, and it hands you a disagreement list that is where the human hours belong. Agreements tell you almost nothing. The deliverable is not a fixed scorer. It is two error estimates you can print beside the finding count, and an explicit note naming the response shapes your instrument is known to miss.

  • Why is the false-negative direction harder to measure than the false-positive direction?
    The hits are a list you can sample. The misses are not a list at all, so you have to sample the rejected transcripts, which only works if the run kept them.
  • What does a second, differently-prompted scorer over the same stored transcripts buy you?
    It concentrates human attention on the disagreements. Cases where both scorers agree tell you little; the split cases are where the errors live.
  • Your scorer prompt asks did the model comply with the request. What is the better question to ask it?
    Whether the response contains the specific prohibited content, judged against a concrete description of what would count, rather than whether the tone sounded cooperative.

saying these in an interview costs you the question

  • Only auditing the hits, and treating the absence of misses as evidence there are none.
  • Believing a longer scorer prompt fixes the problem without any labelled data to check against.
  • Assuming the scorer reads non-prose output, other languages or encoded text as well as the target produces it.
  • Deleting non-hit transcripts, which forecloses the recall estimate entirely.

context