You are handing an AI red-team finding to an application team that does not have your scanning harness installed. What must the evidence package contain so they can reproduce the failure themselves?
answer
- runs without your tool
- request, settings, response, criterion
- N out of M attempts
- clean-machine dry run
- appendix, not payload
basics
~20 sShip a self-contained reproduction: the exact request they must send, the endpoint and generation settings it was sent with, the response you observed, and a plain statement of what makes that response a failure. Add how often it happened out of how many attempts. No harness install, no red-team-only dependency.
solid answer
~50 sThe package has to survive being opened by someone who has never heard of your tooling. That means four things, at minimum: 1. **A harness-free reproduction.** The literal request body, headers that matter, endpoint, and the sampling settings used — expressed as something they can paste into their own client, not as "run my scanner with this config". 2. **The observed evidence.** The full request/response exchange as captured, timestamped, with the model or deployment identifier they can compare against what is live now. 3. **The failure criterion.** One sentence saying what in that response makes it a failure. Without it the developer argues about whether the output was actually bad instead of fixing it. 4. **The rate.** How many attempts produced it out of how many run, because a single transcript against a sampled model is an existence proof, not a frequency. Anything else — your run logs, your scanner's internal identifiers, your scoring internals — is an appendix, not the package.
go deeper
Names the essentials: the exact input, the settings, the captured output, and steps that do not require the red-team tool.
Adds the explicit failure criterion and the attempt count, and separates the package from the appendix of raw run output.
Designs the handoff so it survives tool churn and provider drift, and dry-runs it on a clean machine before sending.
Standardises the package shape across the team so findings from different tools and operators arrive in one reviewable format engineering already trusts.
A red-team finding is not a document; it is a request that somebody else change code. The package therefore has exactly one job: let a person who has never seen your tooling produce the same failure on their own machine and recognise it as a failure. Every rule below follows from that, and the two ways it fails are the ticket closed as *unreproducible* and the ticket closed as *works as intended*. ## Where the material comes from, and why it cannot ship as-is Your scanner already holds the raw material. garak writes a JSONL report in which each row is one attempt — the prompt sent, the generator's output, and the verdict of the detector that judged it. promptfoo writes an eval record per test case carrying the assertion's result. PyRIT keeps conversation pieces in its memory store, one row per turn. Those artefacts are the truth of your run, and none of them is a reproduction. They are keyed by the tool's own probe, plugin and target names; they assume the tool's config file; their identifiers drift between releases. Hand one over and the developer's first task is to install and learn your instrument before they can see the bug at all. ## The four parts that must survive the trip 1. **A harness-free reproduction.** The literal request — endpoint, method, headers that matter, exact body — plus the decoding settings that were in force: temperature, top-p or top-k, max tokens, stop sequences, and the system prompt if the harness supplied one. Expressed as a curl invocation or a fifteen-line script against the same endpoint, it carries no red-team dependency and stays valid after your tool's next release. 2. **The observed evidence.** Request and response as captured, verbatim and untruncated, with a capture timestamp and the model or deployment identifier you actually hit. Hosted endpoints change behind stable aliases; without that identifier neither side can later tell a fix apart from a provider-side shift. 3. **The failure criterion, in one sentence.** In your run something decided the attempt counted: a detector, an assertion, a scorer, or your own read. The developer inherits none of it. "Counts as a failure if the response supplies actionable steps toward the objective" is a sentence they can apply to a response your fix has not produced yet, and it is what a regression test will eventually assert on. 4. **The rate with its denominator.** "Observed in 7 of 20 attempts at these settings." One transcript is an existence proof, not a frequency. Everything else — your run logs, the tool's internal identifiers, your scoring internals — is an appendix, not the package. ## What it costs Assembling this properly is real work: mining the harness artefact for the right rows, re-issuing the case from a plain shell, writing the criterion, choosing benign controls. Budget thirty to sixty minutes per finding on top of the run, and a verification loop of twenty trials costs twenty model calls — twenty times the turn count if the attack is multi-turn — which is pennies of spend but minutes of wall clock once a per-minute rate limit is in play. That hour is the cheapest in the engagement. The alternative is a reopen cycle days later in which the developer loses an afternoon failing to reproduce, you lose an hour reconstructing what you ran, and the model behind the endpoint may have moved underneath both of you in the meantime. ## Where the number misleads The rate is the most misread thing in the package. "35% success" means seven of twenty attempts, at your decoding settings, against your deployment, on that day. It is not a user-facing incidence rate, and at that sample size it is consistent with anything from roughly one in five to more than one in two. Two specific misreads follow. First, a transcript with no denominator reads as deterministic, so the developer's single clean re-run reads as disproof rather than as one sample. Second, the count in the tool's own report frequently has a different denominator from the one a developer would measure: scanners multiply each prompt by a repeat factor — garak's `--generations` flag does exactly this — so a flagged-attempt count is attempts, not distinct inputs, and quoting it as the number of ways the feature fails inflates the finding without anybody lying. ## What to check before you send Hand the ticket text to someone who was not on the engagement, on a machine with the harness uninstalled, and watch them work. Anything they have to ask you is a hole in the package. Confirm the deployment identifier is one their platform team can resolve, confirm the criterion is a sentence rather than a score, and confirm the controls are attached — the benign inputs closest to the attack that must still succeed once the fix lands. ``` finding/ README.md # what fails, why it counts, observed 7/20 at stated settings repro.sh # plain HTTP against the same endpoint, decoding settings inline transcript.json # request + response as captured, timestamps, deployment id controls/ # benign inputs that must keep working after the fix ```
- The attack only works as a multi-turn conversation your tool drove automatically. What is the reproduction now?The verbatim conversation transcript, turn by turn, so the developer can replay it by hand or script it, plus a note on which turns are load-bearing. The tool config becomes background, not the repro.
- Why include the model or deployment identifier when the developer already knows their own system?Because hosted endpoints change behind a stable name. Recording what you hit lets both sides tell 'fixed' apart from 'the provider shifted underneath us'.
- How do you validate the package before sending it?Hand the ticket text to someone who was not on the engagement, on a machine without the harness, and see whether they can reproduce it without asking you a question.
saying these in an interview costs you the question
- Reproduction steps that begin with installing the red-team tool.
- Attaching the entire scan log and calling it evidence.
- No statement of what makes the output a failure — just 'the model said this'.
- One transcript with no attempt count, presented as if the behaviour were deterministic.
- Screenshots of a chat window instead of the request that produced it.