In an automated jailbreak tool that mutates prompt templates, what is the seed corpus, why are templates that already elicited a violation the usual seeds, and what access to the model under test does this approach need?
answer
- seeds = templates that already worked
- mutate, send, keep survivors
- query access only, no weights
- depth around seeds, never new families
- check the variant still asks for it
basics
~20 sThe seed corpus is the set of prompt templates the tool starts from, usually ones that already made the model comply. Mutation edits them by rewording, filling slots differently, or splicing two together to make many near variants. It needs only query access: send a prompt, read the reply. No weights, no gradients.
solid answer
~50 sThe **seed corpus** is the starting set of prompt templates a mutation-based jailbreak tool copies and edits. It is seeded with templates that have already produced a violation because this search gets no gradient and often no score beyond pass/fail on the reply, so it needs a starting point already close to something that works; mutating random text almost never lands. Mutation operators are textual: reword a sentence, substitute the filler slot that carries the request, change formatting or language, splice two templates together. Each variant is one query. Access required is the weakest of any jailbreak search: the ability to send a prompt and read the response. That is why it runs against a hosted endpoint you have no weights for. The price of that cheapness is that everything produced lives in the neighbourhood of what you seeded — it buys depth around known families, never a family you did not put in.
code
text · 9 linesseeds = load_templates("corpus/known_working/") # each = <template text with a request slot>
survivors = []
for generation in range(N):
for t in sample(seeds + survivors, k):
v = mutate(t) # reword / re-slot / splice / reformat
r = target.send(v) # only capability needed: prompt in, reply out
if counted_as_violation(r) and still_asks_for_the_thing(v):
survivors.append(v)
report(group_by_seed_lineage(survivors))go deeper
Should say the tool starts from templates known to work, edits them into many variants, and only needs to send prompts and read replies.
Adds why successful seeds are needed — the pass/fail signal is flat elsewhere — and that mutation can break the request, so survivors must be re-read and re-run.
Frames the scope honestly: the corpus defines the run, unseeded families produce silence, and triage must collapse variants back to seed lineage before anything is reported.
Positions the technique as cheap regression monitoring rather than discovery, and insists the report state which families were seeded so leadership does not read variant volume as assurance.
## The pieces, named - A *template* is a prompt with a hole in it: fixed framing text plus a *slot* where the actual request goes. - A *seed corpus* is the set of templates the run starts from — in most fuzzers it is literally a directory you point the run at, and it is the single most consequential input the operator supplies. - A *mutation operator* is a function from one prompt string to another: reword a sentence, swap a synonym, re-fill the slot, change the formatting or the language, splice the opening of one template onto the tail of another. - A *generation* is one pass in which the tool samples parents, applies operators, sends the children, and keeps whatever survived. - The *decision function* — a garak detector, a PyRIT scorer, a promptfoo grader, or a hand-rolled regular expression over the reply — is the component that turns a response into a boolean. Nothing else in the loop has an opinion about whether an attempt counted. ## The loop, step by step Load the seeds. For each generation: 1. sample parents from the seeds plus any surviving children, 2. apply one or more operators to each, 3. send each resulting variant as one request, 4. run the decision function over the reply, 5. keep the survivors. Then group the survivors by which seed they descend from and report. Read that loop twice and the load-bearing property falls out: no step in it invents structure. Every string the run ever sends is a **perturbation of a string you supplied**. ## Why the seeds are templates that already elicited a violation This search gets no gradient and usually no graded score — only pass or fail on the reply. On that signal the space is almost entirely flat: essentially every arbitrary string returns a refusal or an irrelevant answer, so there is nothing to climb. A template known to land starts the search inside the small region where the signal actually varies, and mutation then maps the plateau around it. It is the same reason a coverage-guided binary fuzzer is seeded with valid files rather than random bytes; the corpus is what makes the feedback informative at all. ## What access it needs, and why that is the whole appeal Prompt in, reply out. No weights, no gradients, no per-token log-probabilities, no fine-tuning hook. Contrast that with - white-box suffix optimisation, which needs the weights and GPU hours, or with - searches that steer on token probabilities and therefore need an API that returns them. Because template mutation needs none of it, it runs against any hosted endpoint over the same interface a customer uses, gated only by an API key and a rate limit. That is why it is usually the first automated thing a team stands up — and why its output is the easiest in the toolbox to over-read. ## What it costs **Queries** and **analyst hours**, in that order of visibility and the reverse order of size. - The query arithmetic is variants times re-sends: three seeds expanded to 4,000 variants, each re-sent three times for stability, is 12,000 requests. At a per-account limit of 60 requests a minute that is roughly three and a half hours of wall clock no matter what the tokens cost, and bursts that trip throttling stretch it further through retry backoff. - The token bill for a few hundred tokens in and out per call is often the smallest line. - The real bill is **triage**: survivors scale with the variant budget while families do not, so a bigger run buys you more rows to collapse, not more to learn. ## Where it fails, and how its number misleads Four failure modes, each of which becomes a wrong reading of the output. - *Semantic drift*: an operator paraphrases the disallowed request out of the prompt, leaving something benign the model was right to answer, and the decision function — which only ever saw that no refusal appeared — scores it as a hit. - *Decision-function error*: absence of a refusal phrase is not compliance; a fluent evasion or an on-topic non-answer passes the same test, so the tool's own false-positive rate multiplies straight into the headline. - *Nondeterminism*: with sampled decoding the same variant may land once in twenty attempts, and a binary hit conceals that entirely. - *Scope silence*: the corpus is the scope, so a family nobody seeded generates no rows, and zero rows renders on a dashboard exactly like a pass. Put plainly, the hit count has your own generator in the denominator; it is a statement about your corpus and your operators meeting the model, never about the model alone. ## What I would check before believing a run - How many genuinely distinct seeds the corpus holds and when it was last refreshed; - whether hits were deduplicated back to seed lineage before anyone counted them; - how many re-sends each candidate got and what reproduction rate it earned; - a hand-read sample of survivors to measure how often the decision function was wrong; - and an explicit list of which families were seeded, so that everything absent from it is written up as untested rather than as clean.
- Your mutation run reports a hit, but re-sending the exact same variant gets a refusal. What is going on and what do you do?Sampling nondeterminism: at non-zero temperature the same prompt yields different completions. Re-send each candidate several times and report a reproduction rate rather than a binary hit; a variant that lands once in twenty is a very different claim from one that lands nineteen times in twenty.
- Why can a mutated variant 'succeed' while telling you nothing about the model's safety?Because the mutation edited away the part that carried the disallowed request, so the model complied with something benign. Any survivor has to be re-read to confirm it still asks for the original thing before it counts.
- How does this differ from seeding a run with hand-written templates that have never been tested?Untested seeds put the search back on a flat signal — most produce nothing to mutate toward, so query spend goes into a region with no feedback. Untested templates are worth including deliberately, as an exploration slice, not as the whole corpus.
saying these in an interview costs you the question
- Says it needs model weights, gradients or logprobs — mutation search over templates needs only prompt-in, reply-out.
- Thinks 'fuzzing' means random input, so the seed list does not matter; the seed list is the entire scope of the run.
- Treats every surviving variant as a distinct vulnerability instead of one family restated.
- Never mentions re-running a hit, so a one-off sampling fluke gets reported as a reproducible break.