skip to content

A template-mutation jailbreak run seeded with three prompt templates that previously worked generates 4,000 variants against a chat assistant and reports 400 violating responses. How do you write that 400 up honestly?

level: middleimportance: must knowfreq 62%

answer

  1. tool unit = variant, report unit = family
  2. dedup by seed lineage
  3. denominator is your generator
  4. reproduction rate, not binary hit
  5. unseeded = untested, not passed

basics

~20 s

It counts variants, not distinct weaknesses. Those 400 are near-copies of three seeds, so write it as three template families that still reproduce under rewording, with a reproduction rate each. The denominator is variants you generated, not the model's behaviours or its attack surface, so it supports no breadth claim at all.

solid answer

~50 s

Deduplicate back to seed lineage first. Every variant descends from one of three templates, so the report has at most **three items**, each with a measured reproduction rate — 'family A reproduced in 310 of 1,300 rewordings, and 9 of 10 re-sends of the best variant' — not four hundred bullet points. The 10% figure is a ratio over your own generator's output: change the mutation operators or the variant budget and it moves with the model untouched, so it is comparable to nothing unless corpus and operators were frozen. Two caveats belong in the write-up. Survivors must be re-read to confirm the mutation did not delete the actual request, and re-sent several times because decoding is stochastic. And the run says nothing about families the three seeds did not represent — that silence has to be stated as untested, because a stakeholder reading '400 hits' will otherwise read it as breadth.

go deeper

for a junior

Should at least see that 400 near-copies of three seeds are not 400 separate problems and should be grouped.

for a middle

Explains that the denominator is the tool's own output, deduplicates to seed lineage, and re-sends candidates because decoding is stochastic.

for a senior

Adds validity checks — variants that lost the request, hand-sampling the automatic decision — and reports the range of rewordings that preserved the effect so the owner knows whether a string block will hold.

for a principal

Insists the write-up carries the untested-family line beside the hit count, and sets a house rule for how such numbers may be compared across runs and versions.

**Read the two numbers literally first.** 4,000 is how many strings your generator emitted. 400 is how many of those strings the decision function marked as violations. Neither number contains a fact about the model on its own; both are outputs of a pipeline whose inputs were three templates you chose, a set of mutation operators you configured, a variant budget you set, and a decoding temperature the endpoint used. Any honest write-up starts by saying so. **The unit problem.** The tool's unit is a variant. The report's unit is a triaged, deduplicated family. Converting one into the other is essentially the whole job here. Group survivors by the seed they descend from — most fuzzers keep lineage, and if yours does not, that is a defect to fix before the next run. Then collapse within each family: two survivors that differ only in synonym choice or sentence order are one item, because they fix as one change. Three seeds means at most three report items. A report that lists 400 rows manages to be simultaneously alarming and useless: it makes the assistant look catastrophically broken while giving the owner no discrete thing to fix. **How the ratio misleads.** 400/4000 is 10%, and 10% invites every wrong reading available. It is not an attack-success rate for the model, not a share of the attack surface, and not comparable to another team's number. The reason is that the denominator is your own generator, and the ratio moves for reasons that have nothing to do with the assistant: | Change, model untouched | Effect on 400/4000 | |---|---| | Milder operators (variants stay near a working seed) | Rises | | Aggressive operators (variants drift off the request) | Falls | | Seed three easy templates instead of three hard ones | Rises | | Double the variant budget | Hit count roughly doubles; ratio roughly flat | | Loosen the decision function to any non-refusal | Rises, on false positives | | Lower the decoding temperature | Moves either way, and the run becomes less reproducible to compare against | A number that four knobs on your side can swing is a measurement of the knobs. Cross-version trending is only meaningful with corpus, operators, variant count, re-send count and decoding settings all pinned and versioned alongside the results — and even then you compare per-family reproduction rates, not the headline total. **Validity checks that must run before anything is reported.** Re-send each candidate several times and report a reproduction rate: a variant landing nineteen times in twenty and one landing once in twenty are different claims, and a binary hit erases the difference. Re-read each survivor to confirm the mutation did not paraphrase the disallowed request out of the prompt — semantic drift is the classic false positive here, and the decision function cannot see it because all it observed was the absence of a refusal. Hand-check a random sample rather than trusting the automatic decision on all 400; whatever error rate you measure on the sample is the error bar on your headline, and quoting the headline without it is the single most common overclaim in this kind of report. **What it costs to do properly.** The 4,000 sends were the cheap part. Three re-sends each turns 4,000 requests into 12,000, and against a metered, rate-limited endpoint that is hours of wall clock rather than money. The expensive line is analyst time, and it scales with survivors, not with families: 400 rows to read, sample and collapse is a day of work that produces three findings. Budget it explicitly, because a manager who priced the run by tokens will be surprised. **What to hand the owner.** Per family: the seed's structure, the measured reproduction rate, one representative variant, and the range of rewordings that preserved the effect. That last item is the genuinely valuable output of the whole run, because it tells the owner whether a string-level block can hold or whether the fix has to be behavioural. Then, as a separate and explicit line: which families were in the corpus and therefore which were not tested at all. A run of this kind produces evidence of presence only. On everything unseeded it is silent, and silence is not a pass — but on a slide, silence and a pass look identical, which is exactly why the line has to be written.

  • The owner blocks the exact strings of all 400 variants. What has actually been fixed?
    The measured neighbourhood of three seeds, at string level. The next mutation generation from the same seeds is the cheapest possible test of whether the fix generalised — if fresh rewordings land again, the block was on the surface form, not the behaviour.
  • Your next run against the same model reports 120 hits instead of 400. What can you conclude?
    Nothing, unless the corpus, mutation operators, variant budget and decoding settings were identical. If any of those moved, the drop measures your generator. Even with everything pinned, sampling variance across families means you compare per-family reproduction rates, not the headline total.
  • How would you present the same run to an executive who wants one number?
    Give families reproduced out of families seeded, plus a stated count of families not in the corpus. That keeps the denominator visible and stops variant volume from being read as breadth.

The 10% is like a hit rate from a fishing trip where you also chose the pond, the bait and the number of casts: it tells you about your tackle at least as much as about the fish. Change any of your own settings and the number moves while the water is untouched.

saying these in an interview costs you the question

  • Reports '400 vulnerabilities' or files 400 tickets from one three-seed run.
  • Quotes the 10% figure as an attack-success rate for the model, comparable to another team's number.
  • Never re-sends a candidate, so stochastic one-offs enter the report as reproducible.
  • Presents the run as coverage of the model's jailbreak surface without stating which families were seeded.
  • Assumes any surviving variant still contains the disallowed request.

context