In an automated jailbreak search where an attacker model proposes a prompt, a target model answers it, and a separate scoring model rates that answer, how many inference calls does one refinement pass cost, and why does that arithmetic decide the turn cap you set?
answer
- three calls per pass
- attacker, target, judge
- seeds x restarts x turns x 3
- transcript grows the attacker prompt
- target leg is the metered one
basics
~20 sOne pass costs three calls: the attacker writes a candidate, the target answers it, and the scoring model rates the answer. A cap of twenty turns is therefore sixty calls per thread, not twenty. Multiply by restarts and seed goals before you start, because the target leg is usually the metered one.
solid answer
~50 sEach pass is three inferences with different prices and latencies. The **attacker** reads the last candidate, the target's reply and the rating, then writes a new candidate. The **target** — the system under test — answers it. The **scoring model** reads that answer and returns a rating, often with a rationale the attacker sees next turn. So the run's size is roughly `seed goals x restarts x turns x 3 calls`, and the turn cap is a budget decision before it is a quality knob: going from ten turns to forty quadruples spend on every thread, including the ones that were never going to land. Cost is worse than linear in turns, because the attacker's prompt carries the growing transcript. When attacker and judge run on hardware you own and only the target is metered, the shape flips: wall-clock, not spend, becomes the binding limit — but the cap still exists to bound it.
go deeper
Should say the pass is attacker, target and scorer — three calls — and that the cap multiplies that.
Adds the full multiplication across seeds and restarts, and notes that token cost grows with the transcript so late turns cost more.
Talks about which leg is metered, rate limits as the real ceiling, and a global call ceiling that kills a runaway run independently of per-thread caps.
Frames the cap as spend policy for the engagement and insists on a measured dry run before a wide launch.
## What a "pass" actually contains An attacker-model jailbreak loop has three roles, and it is worth being pedantic about them because the arithmetic below is nothing but a headcount of the roles. - The **attacker model** is an LLM you prompt to write candidate prompts against a stated goal; it is not the thing under test. - The **target** is the thing under test, and it is the whole deployed stack — model weights, system prompt, any input filter, any output filter — not just the weights, because a candidate that the input filter blocks has been refused by the target as surely as one the model declines. - The **judge** (also called the scorer, the grader or the detector, depending on whose tool you are in) is a third model asked to read the target's answer and rate whether it constitutes a success against that goal. One refinement pass runs all three: propose, send, score. Three billed inferences, not one. Some loops run four legs. A cheap refusal classifier in front of the judge, a summariser that compresses the transcript once it outgrows the attacker's context window, a separate paraphraser between attacker and target — each is another call per turn. Count legs by reading the loop's code, not by assuming three. ## The arithmetic Total calls are approximately `goals x restarts x turns x legs`. Twenty seed goals, three restarts each, a cap of fifteen turns, three legs is 2,700 calls before anyone opens a transcript. Raising that cap from fifteen to forty takes the same run to 7,200. The cap is a **spend multiplier** applied to every thread — including the ones that were never going to land, which is most of them — before it is a quality knob. ## Tokens do not scale like calls Most attacker loops feed the accumulated conversation back to the attacker each turn — its own previous candidates, the target's replies, the judge's ratings and rationales — so that it does not repeat a framing that already failed. The attacker's input therefore grows roughly linearly with the turn index, which makes the attacker leg's cost per thread roughly **quadratic in the cap**. Estimating a run's spend as `calls x the first turn's cost` under-reads a deep run badly. The cheap correction is to measure tokens per leg at turn one and again at the cap during a pilot, and price the run off the average of the two. ## Where the money sits Usually the **target leg**: it is the hosted, metered system you are testing, it is the one leg you cannot substitute, and you pay per token both ways. The attacker is frequently an open-weights model on hardware you already own, and the judge is a short call over a single answer. That asymmetry decides which knob is worth turning. - Scoring only every second or third turn saves almost nothing and blinds the loop, because the rating is the signal the attacker's next rewrite is conditioned on. - Cutting the turn cap removes a target call from every thread directly. ## Where the number misleads Three ways, and all three are common. 1. First, price often is not the binding constraint — the endpoint's requests-per-minute or tokens-per-minute allowance is. A run that costs eighty dollars but takes thirty hours because of a low rate limit is a schedule problem wearing a budget's clothes, and the attacker and judge sit idle while the target queues. 2. Second, a per-thread turn cap does not bound total spend: concurrency, restarts and retries multiply it, and retried calls after a 429 or a timeout are billed but almost never appear in the estimate. 3. Third, and worst, the figure people actually quote is **cost per hit**, which has calls in the numerator and judged successes in the denominator. Loosen the judge's threshold and cost per hit falls without one extra real success. Cost per hit is comparable across runs only when the scoring configuration is identical. ## What I check before launching - A dry run at a cap of two turns across the full seed set, which gives real per-leg token counts and confirms the wiring bills what I think it bills. - A read of the loop for extra legs. - The target's rate limit, daily quota and whether retries count against it. - Whether the loop terminates the moment the judge scores a hit or keeps optimising past it. - And a hard **global ceiling** on calls or spend that kills the whole run independently of any per-thread cap, with an alert partway, so one misconfigured seed set cannot burn the engagement's quota in an afternoon.
- The attacker and judge run on your own hardware and only the target is a paid endpoint. What changes about how you set the cap?Spend stops being the binding constraint and wall-clock plus the target's rate limit take over. The cap now bounds run time and quota use, and you can afford a slightly larger attacker prompt if it raises hits per target call.
- Why does raising the turn cap cost more than proportionally?Because the attacker is normally shown the accumulated transcript so it does not repeat failed candidates. Prompt tokens grow every turn, so the last turns are the most expensive ones in the thread.
A turn cap is not a length, it is a multiplier: every turn you add gets paid for three times over — once by the attacker, once by the target, once by the judge — on every thread in the run, including the ones that were never going to land.
saying these in an interview costs you the question
- Counting a refinement pass as a single model call.
- Estimating cost from the first turn's tokens and ignoring that the attacker prompt carries the transcript.
- Proposing to save money by scoring only occasionally, which starves the loop of its signal.
- Launching a wide run with no global call ceiling.