skip to content

Automated Jailbreak Generation

You will learn the tools and algorithms — GCG, PAIR, TAP, AutoDAN — that search for jailbreaks automatically, and their compute/transferability trade-offs. Interviewers ask about these to test whether you can scale red-teaming beyond hand-written prompts.

on this pageshow

explore

questions

page 1 of 2

In an automated jailbreak search where an attacker model proposes a prompt, a target model answers it, and a separate scoring model rates that answer, how many inference calls does one refinement pass cost, and why does that arithmetic decide the turn cap you set?

level: juniorimportance: must knowfreq 72%

answer

  1. three calls per pass
  2. attacker, target, judge
  3. seeds x restarts x turns x 3
  4. transcript grows the attacker prompt
  5. target leg is the metered one

basics

~20 s

One pass costs three calls: the attacker writes a candidate, the target answers it, and the scoring model rates the answer. A cap of twenty turns is therefore sixty calls per thread, not twenty. Multiply by restarts and seed goals before you start, because the target leg is usually the metered one.

solid answer

~50 s

Each pass is three inferences with different prices and latencies. The **attacker** reads the last candidate, the target's reply and the rating, then writes a new candidate. The **target** — the system under test — answers it. The **scoring model** reads that answer and returns a rating, often with a rationale the attacker sees next turn. So the run's size is roughly `seed goals x restarts x turns x 3 calls`, and the turn cap is a budget decision before it is a quality knob: going from ten turns to forty quadruples spend on every thread, including the ones that were never going to land. Cost is worse than linear in turns, because the attacker's prompt carries the growing transcript. When attacker and judge run on hardware you own and only the target is metered, the shape flips: wall-clock, not spend, becomes the binding limit — but the cap still exists to bound it.

go deeper

for a junior

Should say the pass is attacker, target and scorer — three calls — and that the cap multiplies that.

for a middle

Adds the full multiplication across seeds and restarts, and notes that token cost grows with the transcript so late turns cost more.

for a senior

Talks about which leg is metered, rate limits as the real ceiling, and a global call ceiling that kills a runaway run independently of per-thread caps.

for a principal

Frames the cap as spend policy for the engagement and insists on a measured dry run before a wide launch.

## What a "pass" actually contains An attacker-model jailbreak loop has three roles, and it is worth being pedantic about them because the arithmetic below is nothing but a headcount of the roles. - The **attacker model** is an LLM you prompt to write candidate prompts against a stated goal; it is not the thing under test. - The **target** is the thing under test, and it is the whole deployed stack — model weights, system prompt, any input filter, any output filter — not just the weights, because a candidate that the input filter blocks has been refused by the target as surely as one the model declines. - The **judge** (also called the scorer, the grader or the detector, depending on whose tool you are in) is a third model asked to read the target's answer and rate whether it constitutes a success against that goal. One refinement pass runs all three: propose, send, score. Three billed inferences, not one. Some loops run four legs. A cheap refusal classifier in front of the judge, a summariser that compresses the transcript once it outgrows the attacker's context window, a separate paraphraser between attacker and target — each is another call per turn. Count legs by reading the loop's code, not by assuming three. ## The arithmetic Total calls are approximately `goals x restarts x turns x legs`. Twenty seed goals, three restarts each, a cap of fifteen turns, three legs is 2,700 calls before anyone opens a transcript. Raising that cap from fifteen to forty takes the same run to 7,200. The cap is a **spend multiplier** applied to every thread — including the ones that were never going to land, which is most of them — before it is a quality knob. ## Tokens do not scale like calls Most attacker loops feed the accumulated conversation back to the attacker each turn — its own previous candidates, the target's replies, the judge's ratings and rationales — so that it does not repeat a framing that already failed. The attacker's input therefore grows roughly linearly with the turn index, which makes the attacker leg's cost per thread roughly **quadratic in the cap**. Estimating a run's spend as `calls x the first turn's cost` under-reads a deep run badly. The cheap correction is to measure tokens per leg at turn one and again at the cap during a pilot, and price the run off the average of the two. ## Where the money sits Usually the **target leg**: it is the hosted, metered system you are testing, it is the one leg you cannot substitute, and you pay per token both ways. The attacker is frequently an open-weights model on hardware you already own, and the judge is a short call over a single answer. That asymmetry decides which knob is worth turning. - Scoring only every second or third turn saves almost nothing and blinds the loop, because the rating is the signal the attacker's next rewrite is conditioned on. - Cutting the turn cap removes a target call from every thread directly. ## Where the number misleads Three ways, and all three are common. 1. First, price often is not the binding constraint — the endpoint's requests-per-minute or tokens-per-minute allowance is. A run that costs eighty dollars but takes thirty hours because of a low rate limit is a schedule problem wearing a budget's clothes, and the attacker and judge sit idle while the target queues. 2. Second, a per-thread turn cap does not bound total spend: concurrency, restarts and retries multiply it, and retried calls after a 429 or a timeout are billed but almost never appear in the estimate. 3. Third, and worst, the figure people actually quote is **cost per hit**, which has calls in the numerator and judged successes in the denominator. Loosen the judge's threshold and cost per hit falls without one extra real success. Cost per hit is comparable across runs only when the scoring configuration is identical. ## What I check before launching - A dry run at a cap of two turns across the full seed set, which gives real per-leg token counts and confirms the wiring bills what I think it bills. - A read of the loop for extra legs. - The target's rate limit, daily quota and whether retries count against it. - Whether the loop terminates the moment the judge scores a hit or keeps optimising past it. - And a hard **global ceiling** on calls or spend that kills the whole run independently of any per-thread cap, with an alert partway, so one misconfigured seed set cannot burn the engagement's quota in an afternoon.

  • The attacker and judge run on your own hardware and only the target is a paid endpoint. What changes about how you set the cap?
    Spend stops being the binding constraint and wall-clock plus the target's rate limit take over. The cap now bounds run time and quota use, and you can afford a slightly larger attacker prompt if it raises hits per target call.
  • Why does raising the turn cap cost more than proportionally?
    Because the attacker is normally shown the accumulated transcript so it does not repeat failed candidates. Prompt tokens grow every turn, so the last turns are the most expensive ones in the thread.

A turn cap is not a length, it is a multiplier: every turn you add gets paid for three times over — once by the attacker, once by the target, once by the judge — on every thread in the run, including the ones that were never going to land.

saying these in an interview costs you the question

  • Counting a refinement pass as a single model call.
  • Estimating cost from the first turn's tokens and ignoring that the attacker prompt carries the transcript.
  • Proposing to save money by scoring only occasionally, which starves the loop of its signal.
  • Launching a wide run with no global call ceiling.

context

open as a page

In an automated jailbreak loop where an attacker model rewrites prompts, a target model answers, and a separate scoring model labels each answer a success or a refusal, what happens when the scoring model wrongly labels a refusal as a success?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The loop treats that thread as solved and stops rewriting, so the search ends early. The mislabelled refusal is written into the finding list as a discovered attack, and a human triager has to read the transcript to throw it out. Your reported success count is inflated by every such label.

open as a page

You may reach the language model you are contracted to attack only through a metered chat API that returns text — no weights, no logprobs. Which classes of automated jailbreak search can still run against it, and which class is ruled out before you start?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Any search that only needs to send text and read replies still runs: template and mutation pools, and attacker-model loops that rewrite a prompt after each refusal. Gradient-guided suffix optimisation is ruled out — it needs backward passes through the model's weights. Without weights you can only optimise against a local surrogate and hope it transfers.

open as a page

An automated jailbreak search has produced three thousand prompts its judge marked successful, but they are heavy rewrites of two underlying shapes. Why is the raw count of successful prompts a poor signal for whether to keep spending the remaining queries?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The raw count measures how often the search repeated itself, not how much it found. Thousands of near-identical rewrites are one weakness discovered many times. The signal that matters is distinct successful templates: when that stops rising while queries keep being spent, the remaining allowance is buying duplicates.

open as a page

In an automated jailbreak tool that mutates prompt templates, what is the seed corpus, why are templates that already elicited a violation the usual seeds, and what access to the model under test does this approach need?

level: juniorimportance: must knowfreq 58%

basics

~20 s

The seed corpus is the set of prompt templates the tool starts from, usually ones that already made the model comply. Mutation edits them by rewording, filling slots differently, or splicing two together to make many near variants. It needs only query access: send a prompt, read the reply. No weights, no gradients.

open as a page

In a population-based (genetic) prompt search that mutates and crosses over candidate jailbreak prompts against a target model, what job does the fitness function do, and why is choosing it the operator's decision rather than the algorithm's?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The fitness function scores each candidate prompt's result and decides which prompts survive to be mutated and crossed over. The search is indifferent to meaning: it maximises whatever number you hand it. So the operator's choice of that number defines what the run can possibly discover.

open as a page

A teammate proposes running a jailbreak method that optimises the prompt by taking gradients of a refusal-related loss with respect to the input tokens, and aiming it at a vendor's hosted chat endpoint you can only send text to and read text back from. Why can that class of method not run there at all?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Computing that gradient needs a backward pass through the model's weights. A hosted endpoint returns generated text, not derivatives with respect to your input tokens. With no weights loaded on hardware you control, there is nothing to differentiate, so the search never starts. Against that endpoint you are limited to query-driven attacks.

open as a page

A white-box token search appends a tuned suffix to a request and marks a trial successful when the model's reply begins with a preselected affirmative phrase. Why is that success criterion only a proxy, and what do you check before recording the trial as a real result?

level: juniorimportance: must knowfreq 60%

basics

~20 s

Because the check only reads the reply's first few tokens. A model can open with the agreed phrase and then refuse, stall, or produce useless text. The search optimises exactly that prefix, so it overfits it. Before recording a hit, read or grade the whole continuation.

open as a page

You ran a gradient-based search against an open-weights model on your own GPUs to produce an adversarial suffix, then sent that frozen string unchanged to a hosted chat endpoint whose weights you cannot inspect. What does it mean to say the attack "transferred", and why is a failure at the hosted endpoint less informative than a failure on the local model?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Transfer means the string you optimised on the local model still produces the disallowed behaviour on the hosted endpoint, with no re-optimisation. Failure there is uninformative because you get no gradients and no internals back, only the reply text. On the local model a failure still shows you the loss and lets you keep searching.

open as a page

A branching jailbreak search expands each surviving candidate prompt into b children, prunes each level back to at most w survivors, and repeats for d levels from one seed behaviour. Roughly how many candidate prompts reach the model under test per seed, and how does that count change if you remove the width cap?

level: middleimportance: must knowfreq 46%

basics

~20 s

With the width cap you send about w times b candidates per level, so roughly w times b times d prompts per seed behaviour, linear in depth. Remove the cap and every child branches again, so the count grows like b to the power of d. The cap is what keeps a deep search affordable.

open as a page

You must choose the maximum number of turns for an attacker-model jailbreak loop (an attacker proposes a prompt, the target answers, a scoring model rates the answer). How do you pick that cap from evidence rather than guessing, and why do extra turns buy less and less?

level: middleimportance: must knowfreq 58%

basics

~20 s

Run a pilot with a generous cap and record, for every thread that succeeded, the turn on which it first succeeded. Most hits land early and the tail is thin. Put the cap where new hits stop accruing, then spend the freed calls on more restarts and more seed goals instead of deeper threads.

open as a page

In an automated jailbreak run where a scoring model labels each target response a success or a refusal, which shapes of target response most often get mislabelled in each direction, and how do you catch them?

level: middleimportance: must knowfreq 60%

basics

~20 s

Toward false success: fictional or roleplay wrappers, agreeable openings that then refuse, restatements of the request, and vague non-actionable text. Toward missed success: harmful content buried mid-answer, output in another language, code or encoded output. Catch them by hand-labelling a sample of both hits and non-hits.

open as a page

Gradient-guided adversarial-suffix search optimises a token string by backpropagating through a language model. What must you already hold before such a search can start, and which resource does its cost meter actually run on?

level: middleimportance: must knowfreq 60%

basics

~20 s

You need the weights locally, the matching tokenizer, and enough GPU memory to run forward and backward passes, plus a target string whose likelihood the search maximises. Its cost is GPU-hours, not API calls: many gradient steps and batched candidate evaluations per suffix, with nothing billed per request.

open as a page

You want a live saturation signal for an automated jailbreak search: distinct successful templates per thousand queries. How do you count 'distinct', and why does clustering the successful prompts by surface text similarity mislead you in both directions?

level: middleimportance: must knowfreq 60%

basics

~20 s

Group the successful prompts and count groups per thousand queries spent. Surface-text similarity misleads both ways: a search mutates wording heavily while keeping one mechanism, splitting one template into many groups, and two unrelated mechanisms that share boilerplate collapse into one. Cluster on the mechanism and on what the target actually did.

open as a page

A template-mutation jailbreak run seeded with three prompt templates that previously worked generates 4,000 variants against a chat assistant and reports 400 violating responses. How do you write that 400 up honestly?

level: middleimportance: must knowfreq 62%

basics

~20 s

It counts variants, not distinct weaknesses. Those 400 are near-copies of three seeds, so write it as three template families that still reproduce under rewording, with a reproduction rate each. The denominator is variants you generated, not the model's behaviours or its attack surface, so it supports no breadth claim at all.

open as a page

A genetic prompt search scores each candidate by the absence of refusal phrasing in the target model's reply, and after many generations almost every survivor scores near the maximum. Why is that result usually worthless, and what does the population most likely contain?

level: middleimportance: must knowfreq 66%

basics

~20 s

Because not refusing is not the same as complying. The cheapest way to avoid refusal wording is to stop asking for anything harmful, so the search drifts toward fluent, harmless prompts that get chatty answers. The population is full of false successes carrying no payload.

open as a page

Before a gradient-guided prompt search can take its first optimisation step against an open-weights chat model, what must the team actually have on hand, and why does that same list have to be assembled again for every additional model in scope?

level: middleimportance: must knowfreq 58%

basics

~20 s

You need the specific checkpoint's weights on a machine you control, its tokenizer and chat template, permission to hold those weights, and a GPU with room for the model plus backward-pass activations. The gradients belong to that one weight tensor, so nothing carries over: each new model means fresh weights, fresh setup and fresh GPU hours.

open as a page

In a gradient-guided adversarial suffix search against a model whose weights you hold locally, why must the optimisation objective be a fixed target completion string rather than a compliance verdict from a judge model, and how does the choice of that string change the run?

level: middleimportance: must knowfreq 55%

basics

~20 s

Gradients need a differentiable number. Cross-entropy on a fixed target completion gives one; a judge's yes or no is a discrete label from another model with no usable gradient. The string you pick sets the difficulty: a short generic opener is easy but weak evidence, a long specific one is slower and rarer.

open as a page

Before spending GPU hours optimising a jailbreak string against a locally run open-weights model in the hope it fires at a hosted endpoint you cannot inspect, how do you choose which local model to optimise against, and why do practitioners often optimise against several local models at once?

level: middleimportance: must knowfreq 48%

basics

~20 s

Pick a local model as close as you can guess to the hidden one: similar tokenizer, similar chat formatting, similar safety tuning. The string is optimised for those exact tokens and that exact refusal behaviour. Optimising against several local models at once forces the string onto behaviour they share instead of one model's quirks, which usually survives the jump better.

open as a page

A genetic prompt search uses an automated judge model to rate whether the target model's reply is harmful, and uses that rating as fitness. Over a long run the rating climbs steadily, but manual spot-checks of the top candidates find nothing actually harmful. What is happening, and how do you change the setup?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The search is optimising the judge, not the target. Any judge has blind spots, and selection pressure finds them: replies that look harmful to the rater but are not. Fix it by verifying top candidates with a different checker that never drove selection, and by holding out judged examples to measure the rater.

open as a page

A suffix produced by a token-level search is a run of unrelated characters and word fragments no person would type. A deployment adds an input check that rejects prompts whose per-token perplexity under a small language model is far above normal traffic. Why does that check defeat this class of result cheaply, and what does it cost the defender?

level: seniorimportance: must knowfreq 55%

basics

~20 s

The search optimises token probabilities, not readability, so the winning string is statistically bizarre — exactly what a perplexity score measures. Scoring it needs one pass through a tiny model, far cheaper than the search that produced it. The cost is false positives on legitimately odd input and a threshold that must be tuned per traffic mix.

open as a page

A quarter ago your report recorded that a gradient-optimised suffix produced the disallowed behaviour on roughly two of five attempts against a hosted chat endpoint. Re-running the same frozen string today, it almost never works. What are the plausible causes, and what should the original entry have recorded so this is diagnosable at all?

level: seniorimportance: must knowfreq 42%

basics

~20 s

The endpoint changed under you: a new model build, an altered system prompt, or a filter added in front or behind it. Your string did not decay. The entry should have stamped the date, the endpoint and any build identifier returned, decoding settings, the number of attempts, and how a hit was judged. Without those, nothing is attributable.

open as a page

In a branching jailbreak search, candidate prompts are checked for whether they still pursue the behaviour under test, and ones judged to have drifted off-topic are discarded before they are ever sent to the model under test. Why is that off-topic pruning step there, and what does it cost you?

level: middleimportance: should knowfreq 38%

basics

~20 s

It stops the attacker model wandering into prompts that no longer ask for the behaviour you are testing, so the queries you pay for buy relevant attempts. The cost is false pruning: an oblique, apparently unrelated framing is often exactly what slips past a filter, and a strict on-topic check kills it one turn early.

open as a page

In an attacker-model jailbreak loop, one scoring model usually both tells the loop when to stop rewriting a prompt and decides which transcripts enter the report. Why does that double duty make its errors more damaging than the same error rate in an offline evaluation, and what would you change?

level: middleimportance: should knowfreq 52%

basics

~20 s

Because the scorer is not just measuring the run, it is steering it. A false success stops a thread that had not actually broken through; a false refusal keeps a solved thread spending budget or prunes a promising line. Fix it by separating the loop's stop signal from the report's gate.

open as a page

In an attacker-model-in-the-loop jailbreak search — one model rewrites the prompt, the system under test answers, a third model rates the answer — which endpoints are billed on every iteration, and how does the total scale when you widen the search into parallel branches?

level: middleimportance: should knowfreq 50%

basics

~20 s

Three calls per iteration: the attacker model that writes the next prompt, the system under test that answers it, and the judge that rates the answer. Totals scale as branches times depth times three, so widening the search multiplies spend on all three meters, not only the target's.

open as a page

In a genetic search over jailbreak prompts, why does a strictly binary fitness signal (1 if the attempt clearly succeeded, 0 otherwise) tend to stall the search, and what shape of signal do operators use instead?

level: middleimportance: should knowfreq 48%

basics

~20 s

Because almost every early candidate scores 0, so selection has nothing to rank and the search becomes random. Operators use a graded signal: partial credit for a reply that engages, moves toward the requested content, or omits a refusal, so gradual improvement is visible and can be selected for.

open as a page

An adversarial suffix optimised against a locally run open-weights model fires reliably at one vendor's hosted chat endpoint but does nothing at another vendor's hosted endpoint of comparable capability. What mechanisms explain the difference, and how would you work out which one is responsible without any access to either system's internals?

level: middleimportance: should knowfreq 38%

basics

~20 s

Comparable capability does not mean comparable internals. Different tokenizers re-split your string, different safety tuning gives a different refusal to suppress, hidden system prompts differ, and one endpoint may be a pipeline with classifiers around the model. You separate them from the outside by comparing the shape of the failures: wording, variation, latency, and whether any output streamed at all.

open as a page

A branching attacker-model jailbreak search runs against a metered chat endpoint with a fixed query allowance. It exhausts the allowance while only the first third of your seed behaviours have been searched at all; the rest were never attempted. How do you diagnose that, and what do you change?

level: seniorimportance: should knowfreq 33%

basics

~20 s

The allowance is global and the search walks seeds one at a time, so early seeds spend everything. Confirm it by counting queries per seed and per level. Fix it by dividing the allowance into a per-seed cap, capping frontier width, and stopping a seed early once it hits or clearly stalls.

open as a page

You have a fixed query budget against one target and must split it between running each attacker-model jailbreak thread for more turns and starting more sequential restarts from fresh seed framings. How do you decide the split, and what does each side buy?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Depth buys refinement of one framing; restarts buy independent draws at the same goal. Since most hits land in early turns, past the knee of the turn-to-first-hit curve a restart returns more hits per call than another turn. Set the cap at that knee and put the remaining budget into restarts.

open as a page

In a propose-score-rewrite jailbreak loop, one thread has run fifteen turns with the scoring model's rating flat and the target repeating essentially the same refusal. Which signals tell you the thread is stalled, and what do you do with its remaining turns?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Flat ratings, near-duplicate candidates from the attacker, and an unchanging refusal from the target together say the thread has converged on nothing. Stop it early, keep the transcript, and return the remaining turns to the pool for a fresh restart from a different seed framing rather than more rewrites of a losing line.

open as a page

showing 1–30 of 47