skip to content

Attacker-Model Loops

A second model writes the attacks and a scoring model rules on them, so the query bill and the judge's mistakes both land in your findings. Interviewers ask who decided a run counted as a success.

on this pageshow

explore

questions

14

In an automated jailbreak search where an attacker model proposes a prompt, a target model answers it, and a separate scoring model rates that answer, how many inference calls does one refinement pass cost, and why does that arithmetic decide the turn cap you set?

level: juniorimportance: must knowfreq 72%

answer

  1. three calls per pass
  2. attacker, target, judge
  3. seeds x restarts x turns x 3
  4. transcript grows the attacker prompt
  5. target leg is the metered one

basics

~20 s

One pass costs three calls: the attacker writes a candidate, the target answers it, and the scoring model rates the answer. A cap of twenty turns is therefore sixty calls per thread, not twenty. Multiply by restarts and seed goals before you start, because the target leg is usually the metered one.

solid answer

~50 s

Each pass is three inferences with different prices and latencies. The **attacker** reads the last candidate, the target's reply and the rating, then writes a new candidate. The **target** — the system under test — answers it. The **scoring model** reads that answer and returns a rating, often with a rationale the attacker sees next turn. So the run's size is roughly `seed goals x restarts x turns x 3 calls`, and the turn cap is a budget decision before it is a quality knob: going from ten turns to forty quadruples spend on every thread, including the ones that were never going to land. Cost is worse than linear in turns, because the attacker's prompt carries the growing transcript. When attacker and judge run on hardware you own and only the target is metered, the shape flips: wall-clock, not spend, becomes the binding limit — but the cap still exists to bound it.

go deeper

for a junior

Should say the pass is attacker, target and scorer — three calls — and that the cap multiplies that.

for a middle

Adds the full multiplication across seeds and restarts, and notes that token cost grows with the transcript so late turns cost more.

for a senior

Talks about which leg is metered, rate limits as the real ceiling, and a global call ceiling that kills a runaway run independently of per-thread caps.

for a principal

Frames the cap as spend policy for the engagement and insists on a measured dry run before a wide launch.

## What a "pass" actually contains An attacker-model jailbreak loop has three roles, and it is worth being pedantic about them because the arithmetic below is nothing but a headcount of the roles. - The **attacker model** is an LLM you prompt to write candidate prompts against a stated goal; it is not the thing under test. - The **target** is the thing under test, and it is the whole deployed stack — model weights, system prompt, any input filter, any output filter — not just the weights, because a candidate that the input filter blocks has been refused by the target as surely as one the model declines. - The **judge** (also called the scorer, the grader or the detector, depending on whose tool you are in) is a third model asked to read the target's answer and rate whether it constitutes a success against that goal. One refinement pass runs all three: propose, send, score. Three billed inferences, not one. Some loops run four legs. A cheap refusal classifier in front of the judge, a summariser that compresses the transcript once it outgrows the attacker's context window, a separate paraphraser between attacker and target — each is another call per turn. Count legs by reading the loop's code, not by assuming three. ## The arithmetic Total calls are approximately `goals x restarts x turns x legs`. Twenty seed goals, three restarts each, a cap of fifteen turns, three legs is 2,700 calls before anyone opens a transcript. Raising that cap from fifteen to forty takes the same run to 7,200. The cap is a **spend multiplier** applied to every thread — including the ones that were never going to land, which is most of them — before it is a quality knob. ## Tokens do not scale like calls Most attacker loops feed the accumulated conversation back to the attacker each turn — its own previous candidates, the target's replies, the judge's ratings and rationales — so that it does not repeat a framing that already failed. The attacker's input therefore grows roughly linearly with the turn index, which makes the attacker leg's cost per thread roughly **quadratic in the cap**. Estimating a run's spend as `calls x the first turn's cost` under-reads a deep run badly. The cheap correction is to measure tokens per leg at turn one and again at the cap during a pilot, and price the run off the average of the two. ## Where the money sits Usually the **target leg**: it is the hosted, metered system you are testing, it is the one leg you cannot substitute, and you pay per token both ways. The attacker is frequently an open-weights model on hardware you already own, and the judge is a short call over a single answer. That asymmetry decides which knob is worth turning. - Scoring only every second or third turn saves almost nothing and blinds the loop, because the rating is the signal the attacker's next rewrite is conditioned on. - Cutting the turn cap removes a target call from every thread directly. ## Where the number misleads Three ways, and all three are common. 1. First, price often is not the binding constraint — the endpoint's requests-per-minute or tokens-per-minute allowance is. A run that costs eighty dollars but takes thirty hours because of a low rate limit is a schedule problem wearing a budget's clothes, and the attacker and judge sit idle while the target queues. 2. Second, a per-thread turn cap does not bound total spend: concurrency, restarts and retries multiply it, and retried calls after a 429 or a timeout are billed but almost never appear in the estimate. 3. Third, and worst, the figure people actually quote is **cost per hit**, which has calls in the numerator and judged successes in the denominator. Loosen the judge's threshold and cost per hit falls without one extra real success. Cost per hit is comparable across runs only when the scoring configuration is identical. ## What I check before launching - A dry run at a cap of two turns across the full seed set, which gives real per-leg token counts and confirms the wiring bills what I think it bills. - A read of the loop for extra legs. - The target's rate limit, daily quota and whether retries count against it. - Whether the loop terminates the moment the judge scores a hit or keeps optimising past it. - And a hard **global ceiling** on calls or spend that kills the whole run independently of any per-thread cap, with an alert partway, so one misconfigured seed set cannot burn the engagement's quota in an afternoon.

  • The attacker and judge run on your own hardware and only the target is a paid endpoint. What changes about how you set the cap?
    Spend stops being the binding constraint and wall-clock plus the target's rate limit take over. The cap now bounds run time and quota use, and you can afford a slightly larger attacker prompt if it raises hits per target call.
  • Why does raising the turn cap cost more than proportionally?
    Because the attacker is normally shown the accumulated transcript so it does not repeat failed candidates. Prompt tokens grow every turn, so the last turns are the most expensive ones in the thread.

A turn cap is not a length, it is a multiplier: every turn you add gets paid for three times over — once by the attacker, once by the target, once by the judge — on every thread in the run, including the ones that were never going to land.

saying these in an interview costs you the question

  • Counting a refinement pass as a single model call.
  • Estimating cost from the first turn's tokens and ignoring that the attacker prompt carries the transcript.
  • Proposing to save money by scoring only occasionally, which starves the loop of its signal.
  • Launching a wide run with no global call ceiling.

context

open as a page

In an automated jailbreak loop where an attacker model rewrites prompts, a target model answers, and a separate scoring model labels each answer a success or a refusal, what happens when the scoring model wrongly labels a refusal as a success?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The loop treats that thread as solved and stops rewriting, so the search ends early. The mislabelled refusal is written into the finding list as a discovered attack, and a human triager has to read the transcript to throw it out. Your reported success count is inflated by every such label.

open as a page

A branching jailbreak search expands each surviving candidate prompt into b children, prunes each level back to at most w survivors, and repeats for d levels from one seed behaviour. Roughly how many candidate prompts reach the model under test per seed, and how does that count change if you remove the width cap?

level: middleimportance: must knowfreq 46%

basics

~20 s

With the width cap you send about w times b candidates per level, so roughly w times b times d prompts per seed behaviour, linear in depth. Remove the cap and every child branches again, so the count grows like b to the power of d. The cap is what keeps a deep search affordable.

open as a page

You must choose the maximum number of turns for an attacker-model jailbreak loop (an attacker proposes a prompt, the target answers, a scoring model rates the answer). How do you pick that cap from evidence rather than guessing, and why do extra turns buy less and less?

level: middleimportance: must knowfreq 58%

basics

~20 s

Run a pilot with a generous cap and record, for every thread that succeeded, the turn on which it first succeeded. Most hits land early and the tail is thin. Put the cap where new hits stop accruing, then spend the freed calls on more restarts and more seed goals instead of deeper threads.

open as a page

In an automated jailbreak run where a scoring model labels each target response a success or a refusal, which shapes of target response most often get mislabelled in each direction, and how do you catch them?

level: middleimportance: must knowfreq 60%

basics

~20 s

Toward false success: fictional or roleplay wrappers, agreeable openings that then refuse, restatements of the request, and vague non-actionable text. Toward missed success: harmful content buried mid-answer, output in another language, code or encoded output. Catch them by hand-labelling a sample of both hits and non-hits.

open as a page

In a branching jailbreak search, candidate prompts are checked for whether they still pursue the behaviour under test, and ones judged to have drifted off-topic are discarded before they are ever sent to the model under test. Why is that off-topic pruning step there, and what does it cost you?

level: middleimportance: should knowfreq 38%

basics

~20 s

It stops the attacker model wandering into prompts that no longer ask for the behaviour you are testing, so the queries you pay for buy relevant attempts. The cost is false pruning: an oblique, apparently unrelated framing is often exactly what slips past a filter, and a strict on-topic check kills it one turn early.

open as a page

In an attacker-model jailbreak loop, one scoring model usually both tells the loop when to stop rewriting a prompt and decides which transcripts enter the report. Why does that double duty make its errors more damaging than the same error rate in an offline evaluation, and what would you change?

level: middleimportance: should knowfreq 52%

basics

~20 s

Because the scorer is not just measuring the run, it is steering it. A false success stops a thread that had not actually broken through; a false refusal keeps a solved thread spending budget or prunes a promising line. Fix it by separating the loop's stop signal from the report's gate.

open as a page

A branching attacker-model jailbreak search runs against a metered chat endpoint with a fixed query allowance. It exhausts the allowance while only the first third of your seed behaviours have been searched at all; the rest were never attempted. How do you diagnose that, and what do you change?

level: seniorimportance: should knowfreq 33%

basics

~20 s

The allowance is global and the search walks seeds one at a time, so early seeds spend everything. Confirm it by counting queries per seed and per level. Fix it by dividing the allowance into a per-seed cap, capping frontier width, and stopping a seed early once it hits or clearly stalls.

open as a page

You have a fixed query budget against one target and must split it between running each attacker-model jailbreak thread for more turns and starting more sequential restarts from fresh seed framings. How do you decide the split, and what does each side buy?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Depth buys refinement of one framing; restarts buy independent draws at the same goal. Since most hits land in early turns, past the knee of the turn-to-first-hit curve a restart returns more hits per call than another turn. Set the cap at that knee and put the remaining budget into restarts.

open as a page

In a propose-score-rewrite jailbreak loop, one thread has run fifteen turns with the scoring model's rating flat and the target repeating essentially the same refusal. Which signals tell you the thread is stalled, and what do you do with its remaining turns?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Flat ratings, near-duplicate candidates from the attacker, and an unchanging refusal from the target together say the thread has converged on nothing. Stop it early, keep the transcript, and return the remaining turns to the pool for a fresh restart from a different seed framing rather than more rewrites of a losing line.

open as a page

You inherit a finished automated jailbreak run that reports sixty successful attacks against a chat endpoint, every one labelled by a scoring model. How do you measure that scorer's error rate in both directions before the report goes out?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Hand-label a random sample of the sixty labelled hits to get precision. For the other direction you need the stored transcripts of everything the scorer rejected: sample those, weighted toward borderline confidence, and hand-label. If only hits were kept, say recall is unmeasured rather than guessing.

open as a page

You have a fixed number of queries against the model under test for an engagement and around two hundred seed behaviours to search with a branching attacker-model loop. How do you decide between a shallow wide pass over all of them and a deep search on a chosen few, and what makes that decision defensible to the team reading the report?

level: principalimportance: should knowfreq 24%

basics

~20 s

Run shallow across everything first, then spend the remainder deep on the behaviours that showed partial movement. Wide answers which behaviours are reachable at all and gives an honest coverage denominator; deep answers how hard a specific one is. Decide by which claim the report has to support, and write the split down beforehand.

open as a page

Your automated attacker-model jailbreak loop ran to its turn cap and restart count against a target and produced no successful attempt. What can you legitimately write in the report, and what would you refuse to claim?

level: principalimportance: should knowfreq 30%

basics

~20 s

You can write that no attempt succeeded within the turn cap and restart count used, against this seed set, judged by this scoring model, on this target build and date. You cannot write that the target is jailbreak-resistant or that no attack exists. A capped search reports a budget, not a property.

open as a page

Your automated jailbreak loop uses the same model deployment as both the target under attack and the scoring model that decides which responses count as successes. As the engagement lead, what is your policy on that arrangement, and what do you require before a machine-labelled finding list is signed off?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

A scorer that shares the target's blind spots misses exactly the outputs the target should have refused, so the most important successes go unreported. Prefer an independent scorer, add deterministic detectors, measure the scorer against a fixed human-labelled set, and require human adjudication before any finding is signed off.

open as a page