Your multi-turn red-team harness generated each follow-up prompt with an attacker model rather than reading a fixed script, and it succeeded against a hosted chat endpoint. What goes in the finding so the result can be re-run, given that re-running the harness produces different prompts every time?
answer
- replay the result, don't re-run the search
- verbatim turns = the reproduction artefact
- generator config = provenance section
- capture mid-conversation retrieval and tool results
- record what decided it was a hit
basics
~20 sShip the realised conversation, not the generator. Store every turn verbatim in order as a fixed replay script a reader can send as-is, plus the target's decoding settings and system prompt. Record the attacker-model configuration separately, as provenance for how the conversation was found, not as the reproduction steps.
solid answer
~50 sSeparate two things people merge: **replaying the attack** and **re-searching for it**. Replay is what the finding owes a reader: a fixed artefact — the exact sequence of turns that worked, verbatim, with the target's system prompt and decoding settings, playable by hand with no attacker model involved. Anyone can execute it and get a yes or no. Re-search is what your harness does, and it does not reproduce the same way — a generator run again finds a different path, or none. Record its configuration (the strategy, the turn budget, the stopping condition, the criterion that decided the attempt counted, and the attacker model's settings) as provenance, in a section labelled as such. Keep them apart because they answer different questions. Replay answers "is this defect still there". Provenance answers "how much search did finding it cost". Filing the generator config as reproduction steps hands the developer a job they cannot do.
go deeper
Attaches the winning transcript, which is the right instinct, but may leave out the system prompt, decoding settings and mid-conversation tool results.
Explicitly separates the verbatim replay from the generator configuration and knows the generator will not retrace its own path.
Also captures mid-conversation retrieval and tool results, names the criterion that decided the attempt was a hit, and flags app state the transcript cannot carry.
Requires both sections in every adaptive-harness finding and uses the search-runs-to-hit ratio as an org-level signal about adversary cost.
### A search procedure does not reproduce; its result does An adaptive multi-turn harness is a loop: an **attacker model** reads the target's last reply, writes the next prompt, sends it to the **target**, and a **scorer** decides whether the objective has been met or the loop continues. Because the attacker model is itself sampling, two runs of the same configuration explore different branches. Re-running it is not a re-test — it is a fresh search over the same space. The finding therefore has to carry the *result* as a first-class artefact and treat the procedure as history. ### The replay artefact Everything a reader needs to send the conversation by hand, with no harness and no attacker model: - every turn in order, verbatim, **including the target's replies** — a reader needs them to see where the conversation was steered, and to recognise when their own replay has diverged; - the target's system prompt in force and its effective decoding settings; - any tool call results and retrieved documents that came back mid-conversation. These matter more than people expect: if turn four only worked because one particular document was retrieved, a replay that retrieves something else fails and reads as a fix; - an explicit statement of any state the transcript cannot carry — a session id, a memory store, an uploaded file. Silently shipping a transcript that cannot stand alone manufactures a false "does not reproduce". ### The provenance section, kept separate The strategy, the turn budget, the stopping condition, the scorer that decided the attempt counted, the attacker model's identity and decoding settings, and how many independent search runs were launched versus how many found a path. Label it provenance, not reproduction steps. ### What the search cost, and why the finding should say Each turn typically costs three model calls, not one: attacker generation, target completion, scorer evaluation. A ten-turn budget over twenty parallel search runs is therefore up to six hundred calls for one objective. The scorer is often the largest per-token line item, because a transcript-level judge re-reads the whole growing conversation on every turn, so its input tokens grow quadratically in turn count while the attacker's and target's grow linearly. Wall clock is dominated by the serial turn structure — twenty runs of ten turns is minutes to tens of minutes even fully parallelised. In money that is typically single-digit dollars per objective on hosted endpoints and rather more if the target is a large frontier model; the number worth writing down is calls-per-objective, because it is the one that scales with your objective list. ### Where the numbers mislead - **Reporting the hit without the search ratio.** "The harness found a working path" is compatible with one hit in twenty runs and with nineteen in twenty, and those describe very different adversary effort. The harness already knows the ratio; omitting it flatters the finding. - **Trusting the harness's own scorer.** Many scorers reduce to "the reply contains no refusal string". Absence of a refusal is not presence of the harmful behaviour, and that substitution is the single largest source of inflated hit counts in adaptive tooling. Record what decided the hit, and its version — a replay verified under a *different* criterion is not verifying the same claim. - **Giving the attacker model information a real adversary lacks.** If the harness fed it the target's system prompt or internal documentation to steer with, the path it found may be unreachable from outside. That is a caveat on exposure, and it belongs in the provenance section. - **Re-running the harness to check a fix.** A search that finds nothing today has resampled, not verified. Fix verification runs against the replay artefact; the harness re-run is a separate, weaker signal. ### What you would check Replay the stored transcript into a fresh session, by script rather than by hand, and confirm the target's replies match materially rather than exactly — they will not be identical, and demanding that they are is its own false negative. Have a human read the final reply against the finding's stated harm claim rather than trusting the scorer. Confirm whether the mid-conversation retrieval results are stable, since if they are not, your replay has a hidden input. Confirm what the attacker model was allowed to see. And before filing, hand the replay to somebody with no access to your harness and ask them to execute it — if they cannot, what you have written is provenance wearing a reproduction section's label.
- The replay needs a session or uploaded file the transcript does not contain. What do you do?State it as a stated limit of the replay and attach or describe the missing state. Silently shipping a transcript that cannot stand alone produces a false 'does not reproduce'.
- Why record how many search runs found the path?It separates a fragile one-off from a path the harness finds most times, which is what tells a reader how much adversary effort the defect actually costs.
Re-running an adaptive harness is like sending a different person into the maze: they may find another exit or none at all, and neither outcome tells you whether the exit you found is still open.
saying these in an interview costs you the question
- Files 'run the harness with this config' as the reproduction steps.
- Ships only the final winning prompt, dropping the earlier turns that set it up.
- Omits mid-conversation retrieval or tool outputs that the success depended on.
- Claims the harness is deterministic because its config file is fixed.
- Leaves unstated that the replay needs session state the reader does not have.