skip to content

Iterative Refinement

A propose-score-rewrite loop needs only query access, but each pass pays an attacker, a target and a judge, and a stalled thread burns budget in silence. Interviewers ask how you cap and restart one.

on this pageshow

explore

questions

5

In an automated jailbreak search where an attacker model proposes a prompt, a target model answers it, and a separate scoring model rates that answer, how many inference calls does one refinement pass cost, and why does that arithmetic decide the turn cap you set?

level: juniorimportance: must knowfreq 72%

answer

  1. three calls per pass
  2. attacker, target, judge
  3. seeds x restarts x turns x 3
  4. transcript grows the attacker prompt
  5. target leg is the metered one

basics

~20 s

One pass costs three calls: the attacker writes a candidate, the target answers it, and the scoring model rates the answer. A cap of twenty turns is therefore sixty calls per thread, not twenty. Multiply by restarts and seed goals before you start, because the target leg is usually the metered one.

solid answer

~50 s

Each pass is three inferences with different prices and latencies. The **attacker** reads the last candidate, the target's reply and the rating, then writes a new candidate. The **target** — the system under test — answers it. The **scoring model** reads that answer and returns a rating, often with a rationale the attacker sees next turn. So the run's size is roughly `seed goals x restarts x turns x 3 calls`, and the turn cap is a budget decision before it is a quality knob: going from ten turns to forty quadruples spend on every thread, including the ones that were never going to land. Cost is worse than linear in turns, because the attacker's prompt carries the growing transcript. When attacker and judge run on hardware you own and only the target is metered, the shape flips: wall-clock, not spend, becomes the binding limit — but the cap still exists to bound it.

go deeper

for a junior

Should say the pass is attacker, target and scorer — three calls — and that the cap multiplies that.

for a middle

Adds the full multiplication across seeds and restarts, and notes that token cost grows with the transcript so late turns cost more.

for a senior

Talks about which leg is metered, rate limits as the real ceiling, and a global call ceiling that kills a runaway run independently of per-thread caps.

for a principal

Frames the cap as spend policy for the engagement and insists on a measured dry run before a wide launch.

## What a "pass" actually contains An attacker-model jailbreak loop has three roles, and it is worth being pedantic about them because the arithmetic below is nothing but a headcount of the roles. - The **attacker model** is an LLM you prompt to write candidate prompts against a stated goal; it is not the thing under test. - The **target** is the thing under test, and it is the whole deployed stack — model weights, system prompt, any input filter, any output filter — not just the weights, because a candidate that the input filter blocks has been refused by the target as surely as one the model declines. - The **judge** (also called the scorer, the grader or the detector, depending on whose tool you are in) is a third model asked to read the target's answer and rate whether it constitutes a success against that goal. One refinement pass runs all three: propose, send, score. Three billed inferences, not one. Some loops run four legs. A cheap refusal classifier in front of the judge, a summariser that compresses the transcript once it outgrows the attacker's context window, a separate paraphraser between attacker and target — each is another call per turn. Count legs by reading the loop's code, not by assuming three. ## The arithmetic Total calls are approximately `goals x restarts x turns x legs`. Twenty seed goals, three restarts each, a cap of fifteen turns, three legs is 2,700 calls before anyone opens a transcript. Raising that cap from fifteen to forty takes the same run to 7,200. The cap is a **spend multiplier** applied to every thread — including the ones that were never going to land, which is most of them — before it is a quality knob. ## Tokens do not scale like calls Most attacker loops feed the accumulated conversation back to the attacker each turn — its own previous candidates, the target's replies, the judge's ratings and rationales — so that it does not repeat a framing that already failed. The attacker's input therefore grows roughly linearly with the turn index, which makes the attacker leg's cost per thread roughly **quadratic in the cap**. Estimating a run's spend as `calls x the first turn's cost` under-reads a deep run badly. The cheap correction is to measure tokens per leg at turn one and again at the cap during a pilot, and price the run off the average of the two. ## Where the money sits Usually the **target leg**: it is the hosted, metered system you are testing, it is the one leg you cannot substitute, and you pay per token both ways. The attacker is frequently an open-weights model on hardware you already own, and the judge is a short call over a single answer. That asymmetry decides which knob is worth turning. - Scoring only every second or third turn saves almost nothing and blinds the loop, because the rating is the signal the attacker's next rewrite is conditioned on. - Cutting the turn cap removes a target call from every thread directly. ## Where the number misleads Three ways, and all three are common. 1. First, price often is not the binding constraint — the endpoint's requests-per-minute or tokens-per-minute allowance is. A run that costs eighty dollars but takes thirty hours because of a low rate limit is a schedule problem wearing a budget's clothes, and the attacker and judge sit idle while the target queues. 2. Second, a per-thread turn cap does not bound total spend: concurrency, restarts and retries multiply it, and retried calls after a 429 or a timeout are billed but almost never appear in the estimate. 3. Third, and worst, the figure people actually quote is **cost per hit**, which has calls in the numerator and judged successes in the denominator. Loosen the judge's threshold and cost per hit falls without one extra real success. Cost per hit is comparable across runs only when the scoring configuration is identical. ## What I check before launching - A dry run at a cap of two turns across the full seed set, which gives real per-leg token counts and confirms the wiring bills what I think it bills. - A read of the loop for extra legs. - The target's rate limit, daily quota and whether retries count against it. - Whether the loop terminates the moment the judge scores a hit or keeps optimising past it. - And a hard **global ceiling** on calls or spend that kills the whole run independently of any per-thread cap, with an alert partway, so one misconfigured seed set cannot burn the engagement's quota in an afternoon.

  • The attacker and judge run on your own hardware and only the target is a paid endpoint. What changes about how you set the cap?
    Spend stops being the binding constraint and wall-clock plus the target's rate limit take over. The cap now bounds run time and quota use, and you can afford a slightly larger attacker prompt if it raises hits per target call.
  • Why does raising the turn cap cost more than proportionally?
    Because the attacker is normally shown the accumulated transcript so it does not repeat failed candidates. Prompt tokens grow every turn, so the last turns are the most expensive ones in the thread.

A turn cap is not a length, it is a multiplier: every turn you add gets paid for three times over — once by the attacker, once by the target, once by the judge — on every thread in the run, including the ones that were never going to land.

saying these in an interview costs you the question

  • Counting a refinement pass as a single model call.
  • Estimating cost from the first turn's tokens and ignoring that the attacker prompt carries the transcript.
  • Proposing to save money by scoring only occasionally, which starves the loop of its signal.
  • Launching a wide run with no global call ceiling.

context

open as a page

You must choose the maximum number of turns for an attacker-model jailbreak loop (an attacker proposes a prompt, the target answers, a scoring model rates the answer). How do you pick that cap from evidence rather than guessing, and why do extra turns buy less and less?

level: middleimportance: must knowfreq 58%

basics

~20 s

Run a pilot with a generous cap and record, for every thread that succeeded, the turn on which it first succeeded. Most hits land early and the tail is thin. Put the cap where new hits stop accruing, then spend the freed calls on more restarts and more seed goals instead of deeper threads.

open as a page

You have a fixed query budget against one target and must split it between running each attacker-model jailbreak thread for more turns and starting more sequential restarts from fresh seed framings. How do you decide the split, and what does each side buy?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Depth buys refinement of one framing; restarts buy independent draws at the same goal. Since most hits land in early turns, past the knee of the turn-to-first-hit curve a restart returns more hits per call than another turn. Set the cap at that knee and put the remaining budget into restarts.

open as a page

In a propose-score-rewrite jailbreak loop, one thread has run fifteen turns with the scoring model's rating flat and the target repeating essentially the same refusal. Which signals tell you the thread is stalled, and what do you do with its remaining turns?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Flat ratings, near-duplicate candidates from the attacker, and an unchanging refusal from the target together say the thread has converged on nothing. Stop it early, keep the transcript, and return the remaining turns to the pool for a fresh restart from a different seed framing rather than more rewrites of a losing line.

open as a page

Your automated attacker-model jailbreak loop ran to its turn cap and restart count against a target and produced no successful attempt. What can you legitimately write in the report, and what would you refuse to claim?

level: principalimportance: should knowfreq 30%

basics

~20 s

You can write that no attempt succeeded within the turn cap and restart count used, against this seed set, judged by this scoring model, on this target build and date. You cannot write that the target is jailbreak-resistant or that no attack exists. A capped search reports a budget, not a property.

open as a page