skip to content

Operating a Campaign

Once a search is running someone still matches the method to the access on hand and decides when the yield has died. Interviewers ask because an undeduplicated success list inflates a report.

on this pageshow

explore

questions

10

You may reach the language model you are contracted to attack only through a metered chat API that returns text — no weights, no logprobs. Which classes of automated jailbreak search can still run against it, and which class is ruled out before you start?

level: juniorimportance: must knowfreq 70%

answer

  1. access audit before algorithm choice
  2. backward pass needs weights
  3. text-only leaves loop and templates
  4. logprobs turn black box grey
  5. rate limit, not just cost

basics

~20 s

Any search that only needs to send text and read replies still runs: template and mutation pools, and attacker-model loops that rewrite a prompt after each refusal. Gradient-guided suffix optimisation is ruled out — it needs backward passes through the model's weights. Without weights you can only optimise against a local surrogate and hope it transfers.

solid answer

~50 s

Split the search classes by **what they need to read from the model**. - **Gradient-guided token search** computes a loss on a desired continuation and backpropagates into the embedding table to pick token substitutions. It needs the parameters themselves, so a text-only metered endpoint rules it out entirely. - **Attacker-model-in-the-loop search** reads only the reply text: one model proposes a prompt, the system under test answers, a judge rates the answer, repeat. Text access is enough. - **Template and mutation pools** need no model on your side at all — you send seeded variants and grade the replies. The constraint therefore moves from GPU memory to the request meter and the rate limit. The clean split breaks where the endpoint leaks more than text: per-token log-probabilities or a numeric moderation confidence turn the black box grey and let a score-guided search hill-climb instead of sampling blindly.

go deeper

for a junior

Should say that anything requiring gradients needs the weights locally, and that a text-only key still allows template pools and attacker-model loops.

for a middle

Should classify by the signal each class reads — gradients, a numeric score, or reply text only — and note that a log-probability or moderation score changes the answer.

for a senior

Should audit access before choosing: rate limits, whether the target is the model or the whole deployed system, and whether sending drafts to a third-party attacker model is even permitted.

for a principal

Should frame it as which capabilities the group can rely on across engagements, given that most grant only a metered key, and where surrogate-plus-transfer is worth funding.

Method selection on an engagement is an **access audit first and an algorithm choice second**. Before you name a search, write down exactly what the endpoint hands back, because that one fact eliminates a whole class of methods and prices the two that remain. ### The three things a deployment can give you - **Parameters** — the weight tensors, the matching tokenizer, and the model config, on a disk you control. With these you can run a *forward pass* (tokens in, a next-token probability distribution out) and a *backward pass* (a loss differentiated with respect to the inputs or the weights). - **Scores** — numbers that describe the reply without being the reply: per-token log-probabilities, a top-k distribution over next tokens, or a moderation classifier's per-category confidence. - **Text** — the completion string, and perhaps a stop reason. Nothing else. A metered chat API almost always gives you the third, sometimes the second, and never the first. ### What each search class actually reads **Gradient-guided discrete token search** — the published GCG family — defines a loss as the negative log-likelihood of a chosen target continuation given the prompt plus an appended suffix, then backpropagates that loss into the embedding positions of the suffix. The gradient does not pick a token; it *ranks* candidate substitutions at each position, and the algorithm then evaluates a batch of the top-ranked candidates with real forward passes and keeps whichever actually lowered the loss. Both halves of that loop require the parameters in memory you control. An HTTPS endpoint is a function you may *evaluate*; it is not a function you may *differentiate*. There is no backward pass through a POST. Text-only access rules this class out before a line of code is written. **Attacker-model-in-the-loop search** — the PAIR and tree-of-attacks families — reads reply text only. One model proposes a candidate prompt, the system under test answers, a judge model rates the answer, and the attacker rewrites with that feedback in its context. Every signal it consumes is a string, so a text-only key is sufficient. **Template and mutation pools** need no model on your side at all. A curated seed corpus is expanded by mechanical transforms — paraphrase, encoding, role framing, language switch — and each variant's reply is graded by a keyword rule or a small classifier. ### What each costs | class | the meter | what you actually spend | |---|---|---| | gradient search | GPU-hours and VRAM | nothing per attempt; hours to days of a machine you own, per suffix | | attacker loop | three billed endpoints | attacker, target and judge calls every iteration, with token cost rising as the transcript grows | | template pool | one billed call per variant | near-zero engineering; the maintained pool is the asset | Note that the binding limit moves from GPU memory to **the request meter and the rate limit**, which are two different limits. A campaign can be fully funded and still impossible, because a per-minute ceiling turns a depth-heavy loop into a fortnight of wall-clock. ### Where the reading goes wrong - **"Black box" and "text-only" are not the same claim.** If the endpoint returns log-probabilities or a moderation confidence, you hold a continuous objective, and blind accept/reject sampling can be replaced by score-guided hill-climbing that typically reaches a hit in far fewer requests. Calling the target "black box" and stopping there leaves that on the table. - **A same-named local checkpoint is not the target.** The deployed product is a stack: a hidden system prompt, possibly retrieval and tools, an input classifier, an output filter, and a served checkpoint that may be tuned or quantised differently. None of that is in a download. - **A zero-finding pool run reads as robustness and is not.** The denominator of a template pool is the pool. "No hits across 400 templates" means nobody seeded the thing that works, not that nothing works. - **Billed is not tested.** A request rejected by an input filter before generation still costs money and still fails to tell you anything about the model, so a per-request cost figure overstates the coverage it bought. ### What to check before committing Does the response body carry log-probabilities or a moderation score, or only text? What are the per-minute and per-day request ceilings, and does the planned depth fit inside them? Is the object under test the model or the deployed system, and does the scope statement say which? Is sending client-derived prompts to a third-party attacker or judge model contractually permitted at all — because that single clause can eliminate the loop class and leave you with pools.

  • The endpoint returns per-token log-probabilities alongside the reply. Which search class does that unlock?
    Score-guided black-box search: you can rank candidate prompts by a numeric objective and hill-climb, instead of accepting or rejecting on the reply text alone. It still does not give you gradients through the weights.
  • You have weights for a model with the same name as the hosted one. Is that white-box access to the target?
    No. The deployment adds a system prompt, possibly retrieval and tool use, and input/output filtering, and may serve a differently tuned checkpoint. The local copy is a surrogate, so anything found on it is a candidate to be tested against the real endpoint.
  • What is the cheapest search class, and what does the cheapness cost you?
    A fixed template and mutation pool with a keyword or rule check. It costs no second model and is fully reproducible, but it only finds what someone already thought to seed, and it does not adapt to the refusal it just received.

saying these in an interview costs you the question

  • Plans a gradient-based suffix optimisation against a hosted API endpoint.
  • Treats 'black box' and 'text-only' as identical, missing that log-probabilities or moderation scores are a usable signal.
  • Assumes a locally downloaded checkpoint reproduces the deployed system, ignoring system prompt, retrieval and output filtering.
  • Names an attack method without ever stating what access it assumes.

context

open as a page

An automated jailbreak search has produced three thousand prompts its judge marked successful, but they are heavy rewrites of two underlying shapes. Why is the raw count of successful prompts a poor signal for whether to keep spending the remaining queries?

level: juniorimportance: must knowfreq 72%

basics

~20 s

The raw count measures how often the search repeated itself, not how much it found. Thousands of near-identical rewrites are one weakness discovered many times. The signal that matters is distinct successful templates: when that stops rising while queries keep being spent, the remaining allowance is buying duplicates.

open as a page

Gradient-guided adversarial-suffix search optimises a token string by backpropagating through a language model. What must you already hold before such a search can start, and which resource does its cost meter actually run on?

level: middleimportance: must knowfreq 60%

basics

~20 s

You need the weights locally, the matching tokenizer, and enough GPU memory to run forward and backward passes, plus a target string whose likelihood the search maximises. Its cost is GPU-hours, not API calls: many gradient steps and batched candidate evaluations per suffix, with nothing billed per request.

open as a page

You want a live saturation signal for an automated jailbreak search: distinct successful templates per thousand queries. How do you count 'distinct', and why does clustering the successful prompts by surface text similarity mislead you in both directions?

level: middleimportance: must knowfreq 60%

basics

~20 s

Group the successful prompts and count groups per thousand queries spent. Surface-text similarity misleads both ways: a search mutates wording heavily while keeping one mechanism, splitting one template into many groups, and two unrelated mechanisms that share boilerplate collapse into one. Cluster on the mechanism and on what the target actually did.

open as a page

In an attacker-model-in-the-loop jailbreak search — one model rewrites the prompt, the system under test answers, a third model rates the answer — which endpoints are billed on every iteration, and how does the total scale when you widen the search into parallel branches?

level: middleimportance: should knowfreq 50%

basics

~20 s

Three calls per iteration: the attacker model that writes the next prompt, the system under test that answers it, and the judge that rates the answer. Totals scale as branches times depth times three, so widening the search multiplies spend on all three meters, not only the target's.

open as a page

The system you are contracted to attack is a hosted chat product you can only send requests to, but you hold open weights of a similar model plus GPU hours. Under what conditions is running a white-box suffix search on your local copy a defensible plan, and what must you reserve request allowance for?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Only when you treat the result as a candidate, not a finding: a suffix optimised against local weights is tuned to that copy's tokenizer and parameters. Reserve enough requests against the real product to test transfer, and expect most candidates to fail there, especially through an unseen system prompt and output filter.

open as a page

An automated jailbreak search has spent two-thirds of its query allowance against a metered endpoint and has produced no new distinct successful template for the last fifth of those queries, while judged successes keep arriving. How do you decide between stopping, reseeding, and letting it run?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Stop paying for repeats. First check the flat stretch is a real plateau, not a stalled judge or a collapsed search. If it is real, spend the remaining allowance on a different seed set or objective rather than more of the same, and record the queries spent and the last new template so the stop is evidence.

open as a page

You reported that an automated jailbreak search had saturated on a target. A colleague then ran the same search class against the same target with a different seed corpus and objective and found successful templates yours never produced. What did your saturation measurement actually measure, and how should it have been worded?

level: seniorimportance: should knowfreq 46%

basics

~20 s

It measured that one search stopped finding new templates, not that the target has none left. Saturation is relative to the seeds, objective, scoring step and mutation operators that run used; change any of them and the reachable region changes. Word it as this search exhausted, with those settings named, never as an absence of weaknesses.

open as a page

You own automated jailbreak campaigns across a dozen deployed targets, run by several operators, and you must publish one stopping rule based on distinct-template yield. What do you standardise, what do you deliberately leave to the operator, and what goes wrong if you get that split backwards?

level: principalimportance: should knowfreq 34%

basics

~20 s

Standardise the definition of a distinct template, the scoring method, and the requirement that every stop cite a yield curve and the queries spent. Leave the threshold and the reseed decision to the operator, since targets differ. Get it backwards and teams tune the clustering until every campaign saturates exactly on schedule.

open as a page

You lead a red-team group whose engagements almost always grant only a metered text API for the system under test, and rarely model weights. Which automated jailbreak search capabilities do you standardise on, and what do you give up by not building the white-box one?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Standardise on what most engagements can actually run: template and mutation pools for cheap breadth, and an attacker-model loop for depth, both needing only text access. Keep white-box search as an occasional capability. You lose the ability to attack open-weights deployments at their strongest and to produce transferable candidates cheaply.

open as a page