skip to content

In an attacker-model-in-the-loop jailbreak search — one model rewrites the prompt, the system under test answers, a third model rates the answer — which endpoints are billed on every iteration, and how does the total scale when you widen the search into parallel branches?

level: middleimportance: should knowfreq 50%

answer

  1. attacker, target, judge — three meters
  2. branches x depth x three
  3. transcript grows, tokens grow
  4. cheap judge is the false economy
  5. calls per confirmed finding

basics

~20 s

Three calls per iteration: the attacker model that writes the next prompt, the system under test that answers it, and the judge that rates the answer. Totals scale as branches times depth times three, so widening the search multiplies spend on all three meters, not only the target's.

solid answer

~50 s

Each iteration touches **three metered endpoints**, and operators routinely cost only one of them. - The **attacker model** generates the next candidate prompt, usually with the prior refusal in its context, so its input grows as the conversation deepens. - The **system under test** answers — the only meter people remember. - The **judge model** reads the answer and decides whether it counted, which is the call that turns a transcript into a result. Running `b` branches to depth `d` gives roughly `3 × b × d` calls, and the attacker's and judge's token counts grow with transcript length, so cost grows faster than the call count suggests. The three meters are also independently rate-limited: the search can stall on the judge while the target's quota sits unused. The simple accounting breaks when a branch is pruned early or a hit terminates it. Then depth is an upper bound, and the useful figure is calls actually spent per confirmed hit, not the worst case.

go deeper

for a junior

Should notice that more than the target model is being called — an attacker model and a judging model are billed too.

for a middle

Should give the branches-times-depth-times-three shape and note that transcript growth makes later iterations costlier.

for a senior

Should discuss where to spend per call, why a cheap judge is a false economy, and per-role rate limiting, caps and early termination.

for a principal

Should reduce the run to calls per confirmed deduplicated finding and use that to decide between widening and deepening.

### The mechanism, one iteration at a time An attacker-model-in-the-loop search assigns three distinct roles to three model calls, and confusing them is where the cost estimate goes wrong. 1. **The attacker model** receives the objective and, usually, the transcript of what has already been tried and refused, and emits the next candidate prompt. 2. **The system under test** receives that candidate and answers. This is the only call most people count. 3. **The judge model** — a scorer, a grader, a harm classifier, depending on whose framework you are in — reads the answer and decides whether it counted. This is the call that converts a transcript into a *result*. Run `b` parallel branches to at most `d` refinement steps and the worst case is roughly `b x d` iterations at three calls each, so about `3 x b x d` billed calls. Widening the search multiplies all three meters, not just the target's. ### What it costs Token cost is worse than linear in `d` whenever the attacker and the judge see a growing transcript: the tenth iteration of a branch is more expensive than the first, for the same nominal call count. Cost per call also differs sharply across roles — a frontier attacker model and a small judge can be an order of magnitude apart — so a call count is not a bill. Where to spend is a real decision: - **The target call** is fixed by the engagement. You attack what you were given. - **The attacker call** is where paying more per call frequently *reduces* total spend. A stronger proposer needs fewer iterations, and every iteration removed saves a call on all three meters at once. - **The judge call** is the one operators cheapen, and it is the most dangerous to cheapen, for reasons that are not about money at all. ### Where the number misleads **`3 x b x d` is an upper bound that is simultaneously an undercount.** It is an upper bound because branches terminate early on a hit and get pruned when the judge score stops improving, so real runs finish well under it. It is an undercount because retries on throttled, filtered or malformed responses are billed and invisible in the formula, and because the successful branches are the expensive ones — a refusal is a short cheap completion, whereas the long compliant answer you were hoping for costs the most tokens on both the target and the judge. **The judge silently owns every headline number.** It decides what a hit is, which means it sets attack-success rate, it sets when a branch stops, and it sets what lands in the report. A loose judge inflates hits, and the cost does not appear on the API bill — it appears downstream as triage hours and as report items that die under review. A strict judge terminates campaigns early with a clean, confidently wrong "no findings". Neither failure is visible in the spend. **Billed does not mean the model was reached.** If a deployment's input filter rejects a request before generation, you have paid and learned something about the *filter*, not the model. Counting those in your coverage overstates what the campaign tested. **Per-role rate limits are independent.** A single shared retry policy makes the slowest of the three throttle the other two, so a run can sit idle on judge capacity while the target's quota goes unused, and the wall-clock cost bears no relation to the call count. ### What to check Set a **global** cap on total calls, not only a per-branch depth cap. Terminate a branch on a confirmed hit and prune branches whose judge scores have plateaued. Cache identical attacker proposals so convergent branches are not paid for twice on three meters. Handle rate limits per role. Then report the one figure that lets someone decide the next spend: **calls per confirmed, deduplicated finding, broken down by role** — with the judge's success criterion written down beside it, because that criterion is the definition the whole number rests on.

  • Where is spending more per call likely to reduce total cost?
    On the attacker model. Better proposals cut the number of iterations, and every iteration removed saves a call on all three meters.
  • Two branches converge on nearly the same prompt. What should the loop do?
    Detect the duplication and prune one, or cache the response. Otherwise you pay three meters twice for one result, and the report inherits duplicate items.
  • The search reports many hits but triage discards most. Which endpoint do you look at first?
    The judge. A judge with a loose success criterion inflates hits, and the cost shows up downstream as triage work rather than as API spend.

You are not sitting in one taxi with one meter running; you are paying for three, and every extra branch you open hails another three. The bill people misread is the one where they only ever glanced at the meter in the car they were riding in.

saying these in an interview costs you the question

  • Counts only the calls to the system under test.
  • Cheapens the judging model to save cost without considering mislabelled hits and missed hits.
  • Gives a flat per-iteration cost, ignoring that the attacker and judge see a growing transcript.
  • Sets a per-branch depth cap but no global cap on total calls.

context