skip to content

In a branching jailbreak search, candidate prompts are checked for whether they still pursue the behaviour under test, and ones judged to have drifted off-topic are discarded before they are ever sent to the model under test. Why is that off-topic pruning step there, and what does it cost you?

level: middleimportance: should knowfreq 38%

answer

  1. keeps attacker from softening the ask
  2. prevents cheap compliant non-findings
  3. false pruning leaves no evidence
  4. indirect framings look most off-topic
  5. log discards, sample by hand

basics

~20 s

It stops the attacker model wandering into prompts that no longer ask for the behaviour you are testing, so the queries you pay for buy relevant attempts. The cost is false pruning: an oblique, apparently unrelated framing is often exactly what slips past a filter, and a strict on-topic check kills it one turn early.

solid answer

~60 s

Attacker models drift. Told to rewrite a prompt until it succeeds, they will happily soften it into something the target answers freely — and you get a compliant response to a question you never wanted to ask. The on-topic check exists to keep the search anchored to the behaviour under test and to stop you paying target queries for candidates that could not count as a hit even if answered. The cost is asymmetric and easy to miss. A pruned candidate is never sent, so **you never learn what it would have done** — off-topic pruning produces no evidence of its own errors, unlike a scoring call that at least logs a response you can re-read. And the candidates it cuts are precisely the indirect framings — the hypothetical, the translated, the role-shifted — that carry the highest chance of getting past a filter, because they look least like the request. Tune it to be lenient about surface form and strict only about whether an answer would still constitute the behaviour you are hunting.

go deeper

for a junior

Knows candidates can be discarded before being sent and that this saves queries.

for a middle

Explains that the check stops the attacker softening the request into a benign one, and names false pruning of indirect framings as the cost.

for a senior

Separates the on-topic gate from the compliance scorer, instruments discards, and validates strictness with a pruning-off comparison arm.

for a principal

Owns the fact that this single judgment sets both the run's cost and its recall, and decides how much evidence the team keeps about what was never tried.

### Two different judgments, one loop A branching attacker-model search makes two decisions about every candidate, at two different moments, and collapsing them is the most common design error in this leaf. - The **on-topic check** runs *before* the candidate is sent. It asks: is this rewrite still trying to elicit the behaviour we set out to test? A candidate that fails is discarded and never costs a target query. This is what the question is about. - The **scorer** (or judge) runs *after* a response comes back. It asks: did the target comply? The on-topic check may be an LLM call, an embedding-similarity threshold against the seed behaviour, or a keyword rule. Which of the three you pick decides both what the gate costs and what it wrongly kills. ### Why the gate exists An attacker model told to "rewrite until the scorer reports compliance" has a trivially available cheat: make the request benign. It softens the ask, the target answers cheerfully, the scorer sees a compliant response, and the run logs a hit. The finding is worthless — you demonstrated that a model will answer a harmless question. This is reward hacking of the scorer, and the on-topic gate is the constraint that makes the search objective mean what you intended. Secondarily it saves target queries, by killing candidates before they are spent. ### What it costs Not nothing, and the sign of the saving is worth doing on paper. Let *t* be the cost of a target query plus its scoring call, *g* the cost of one on-topic check, and *p* the fraction of candidates the gate discards. The gate pays for itself only when **p*t > g**. If the gate is an LLM of the same class as the target — the tempting choice, because it is the only thing that understands intent — then g is comparable to a target call, and unless the gate is pruning a large fraction you have added cost while removing candidates. An embedding check makes g nearly free and moves the argument entirely onto recall. There is also engineer time: someone must read the discard log, and if nobody does, the gate is unsupervised. ### Where it fails, and how the number misleads **False pruning is silent.** A pruned candidate is never sent, so you have no response to re-read, no score to appeal, and no artefact in the run log unless you deliberately wrote one. Every other component in the loop leaves evidence of its mistakes; this one does not. **It cuts exactly the promising family.** Indirect framings — hypothetical, fictional, role-shifted, translated, encoded — are effective precisely because they do not look like the original request. A keyword or embedding rule scores them as drift. A translated or encoded candidate may be unreadable to the gate, which then defaults to discard, and you have excluded an entire attack family from the search without ever deciding to. **Multi-turn setups die at turn one.** A line that establishes innocuous context and only lands the request several turns later looks off-topic on the setup turn. Judged turn by turn, it never survives to pay off. **The gate can refuse.** If the on-topic check is an aligned chat model asked to rate a deliberately aggressive candidate, it may decline to answer. Parsing a refusal as "off-topic" prunes the strongest candidates first, and the run never mentions it. **The denominator moves under you.** Attack success rate is usually reported as confirmed hits over candidates *sent*. Tighten the gate and the denominator shrinks faster than the numerator, so the reported rate goes **up** while the run finds fewer behaviours. Two runs with different gate strictness are not comparable on that metric at all. ### What to check Log every discard with the gate's verbatim verdict rather than dropping it, and hand-sample fifty. Look for clustering by framing style — clustering is the signature of a rule matching surface form rather than intent. Count how often the gate output failed to parse and where those went. Run a small holdout arm with the gate disabled and compare **hits per target query**, not hits per candidate, because that is the metric the gate cannot inflate. And report attack success rate against candidates *generated*, with the pruned count stated alongside, so a change in gate strictness is visible instead of flattering.

  • Without an on-topic constraint, what does an attacker model optimising for compliance tend to converge on?
    A benign version of the request. The target answers happily, the scorer reads compliance, and the run reports hits that are not hits at all.
  • How would you detect that your on-topic check is over-pruning?
    Log discards instead of dropping them, sample them, and look for clustering on one framing style; then run a small arm with pruning off and compare hits per query.

Reporting success as hits per prompt actually sent, after a gate has quietly thrown away the awkward candidates, is like a school reporting its exam pass rate after excusing the weakest students from sitting. Tighten the filter and the rate improves while the number of people who can actually pass does not.

saying these in an interview costs you the question

  • Conflates the on-topic check with the compliance scoring step — they answer different questions at different points in the loop.
  • Treats pruning as pure cost saving with no recall consequence.
  • Discards pruned candidates without logging them, then claims the search covered the behaviour.
  • Writes an on-topic rule that matches surface keywords, which cuts exactly the oblique framings worth testing.

context