skip to content

Developer Evidence Package

A developer fixes what they can reproduce, so the package has to stand on its own without your harness installed. Interviewers ask because a finding nobody can re-run gets closed as unreproducible.

on this pageshow

explore

questions

5

You are handing an AI red-team finding to an application team that does not have your scanning harness installed. What must the evidence package contain so they can reproduce the failure themselves?

level: juniorimportance: must knowfreq 52%

answer

  1. runs without your tool
  2. request, settings, response, criterion
  3. N out of M attempts
  4. clean-machine dry run
  5. appendix, not payload

basics

~20 s

Ship a self-contained reproduction: the exact request they must send, the endpoint and generation settings it was sent with, the response you observed, and a plain statement of what makes that response a failure. Add how often it happened out of how many attempts. No harness install, no red-team-only dependency.

solid answer

~50 s

The package has to survive being opened by someone who has never heard of your tooling. That means four things, at minimum: 1. **A harness-free reproduction.** The literal request body, headers that matter, endpoint, and the sampling settings used — expressed as something they can paste into their own client, not as "run my scanner with this config". 2. **The observed evidence.** The full request/response exchange as captured, timestamped, with the model or deployment identifier they can compare against what is live now. 3. **The failure criterion.** One sentence saying what in that response makes it a failure. Without it the developer argues about whether the output was actually bad instead of fixing it. 4. **The rate.** How many attempts produced it out of how many run, because a single transcript against a sampled model is an existence proof, not a frequency. Anything else — your run logs, your scanner's internal identifiers, your scoring internals — is an appendix, not the package.

go deeper

for a junior

Names the essentials: the exact input, the settings, the captured output, and steps that do not require the red-team tool.

for a middle

Adds the explicit failure criterion and the attempt count, and separates the package from the appendix of raw run output.

for a senior

Designs the handoff so it survives tool churn and provider drift, and dry-runs it on a clean machine before sending.

for a principal

Standardises the package shape across the team so findings from different tools and operators arrive in one reviewable format engineering already trusts.

A red-team finding is not a document; it is a request that somebody else change code. The package therefore has exactly one job: let a person who has never seen your tooling produce the same failure on their own machine and recognise it as a failure. Every rule below follows from that, and the two ways it fails are the ticket closed as *unreproducible* and the ticket closed as *works as intended*. ## Where the material comes from, and why it cannot ship as-is Your scanner already holds the raw material. garak writes a JSONL report in which each row is one attempt — the prompt sent, the generator's output, and the verdict of the detector that judged it. promptfoo writes an eval record per test case carrying the assertion's result. PyRIT keeps conversation pieces in its memory store, one row per turn. Those artefacts are the truth of your run, and none of them is a reproduction. They are keyed by the tool's own probe, plugin and target names; they assume the tool's config file; their identifiers drift between releases. Hand one over and the developer's first task is to install and learn your instrument before they can see the bug at all. ## The four parts that must survive the trip 1. **A harness-free reproduction.** The literal request — endpoint, method, headers that matter, exact body — plus the decoding settings that were in force: temperature, top-p or top-k, max tokens, stop sequences, and the system prompt if the harness supplied one. Expressed as a curl invocation or a fifteen-line script against the same endpoint, it carries no red-team dependency and stays valid after your tool's next release. 2. **The observed evidence.** Request and response as captured, verbatim and untruncated, with a capture timestamp and the model or deployment identifier you actually hit. Hosted endpoints change behind stable aliases; without that identifier neither side can later tell a fix apart from a provider-side shift. 3. **The failure criterion, in one sentence.** In your run something decided the attempt counted: a detector, an assertion, a scorer, or your own read. The developer inherits none of it. "Counts as a failure if the response supplies actionable steps toward the objective" is a sentence they can apply to a response your fix has not produced yet, and it is what a regression test will eventually assert on. 4. **The rate with its denominator.** "Observed in 7 of 20 attempts at these settings." One transcript is an existence proof, not a frequency. Everything else — your run logs, the tool's internal identifiers, your scoring internals — is an appendix, not the package. ## What it costs Assembling this properly is real work: mining the harness artefact for the right rows, re-issuing the case from a plain shell, writing the criterion, choosing benign controls. Budget thirty to sixty minutes per finding on top of the run, and a verification loop of twenty trials costs twenty model calls — twenty times the turn count if the attack is multi-turn — which is pennies of spend but minutes of wall clock once a per-minute rate limit is in play. That hour is the cheapest in the engagement. The alternative is a reopen cycle days later in which the developer loses an afternoon failing to reproduce, you lose an hour reconstructing what you ran, and the model behind the endpoint may have moved underneath both of you in the meantime. ## Where the number misleads The rate is the most misread thing in the package. "35% success" means seven of twenty attempts, at your decoding settings, against your deployment, on that day. It is not a user-facing incidence rate, and at that sample size it is consistent with anything from roughly one in five to more than one in two. Two specific misreads follow. First, a transcript with no denominator reads as deterministic, so the developer's single clean re-run reads as disproof rather than as one sample. Second, the count in the tool's own report frequently has a different denominator from the one a developer would measure: scanners multiply each prompt by a repeat factor — garak's `--generations` flag does exactly this — so a flagged-attempt count is attempts, not distinct inputs, and quoting it as the number of ways the feature fails inflates the finding without anybody lying. ## What to check before you send Hand the ticket text to someone who was not on the engagement, on a machine with the harness uninstalled, and watch them work. Anything they have to ask you is a hole in the package. Confirm the deployment identifier is one their platform team can resolve, confirm the criterion is a sentence rather than a score, and confirm the controls are attached — the benign inputs closest to the attack that must still succeed once the fix lands. ``` finding/ README.md # what fails, why it counts, observed 7/20 at stated settings repro.sh # plain HTTP against the same endpoint, decoding settings inline transcript.json # request + response as captured, timestamps, deployment id controls/ # benign inputs that must keep working after the fix ```

  • The attack only works as a multi-turn conversation your tool drove automatically. What is the reproduction now?
    The verbatim conversation transcript, turn by turn, so the developer can replay it by hand or script it, plus a note on which turns are load-bearing. The tool config becomes background, not the repro.
  • Why include the model or deployment identifier when the developer already knows their own system?
    Because hosted endpoints change behind a stable name. Recording what you hit lets both sides tell 'fixed' apart from 'the provider shifted underneath us'.
  • How do you validate the package before sending it?
    Hand the ticket text to someone who was not on the engagement, on a machine without the harness, and see whether they can reproduce it without asking you a question.

saying these in an interview costs you the question

  • Reproduction steps that begin with installing the red-team tool.
  • Attaching the entire scan log and calling it evidence.
  • No statement of what makes the output a failure — just 'the model said this'.
  • One transcript with no attempt count, presented as if the behaviour were deterministic.
  • Screenshots of a chat window instead of the request that produced it.

context

open as a page

A developer runs the reproduction from your AI red-team finding once, gets a polite refusal, and closes the ticket as unreproducible — but the behaviour is real and intermittent. What should the evidence package have contained to prevent that outcome?

level: middleimportance: must knowfreq 47%

basics

~20 s

State up front that the target is sampled, so one attempt proves nothing. Ship the observed count out of total attempts, the sampling settings you used, a script that repeats the request and tallies how many responses met the criterion, and a written verification rule such as: any hit in twenty attempts is still a failure.

open as a page

In your AI red-team harness, whether an attempt counted as a failure was decided by an automated model-based grader with a threshold. The application team will re-run the reproduction without that grader. What do you put in the handoff so their verdict matches yours?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Translate the grader into something they can apply. Write the decision rule in plain words with its threshold, attach the grader's output on the cases you sent, include labelled examples just either side of the line, and give a deterministic check they can assert on. If the verdict truly needs the grader, ship it pinned as a small dependency.

open as a page

One attack family in your red-team run flagged forty-plus attempts against the same chat feature. You will not attach all forty to the developer ticket. How do you choose which attempts go into the evidence package and which stay out?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Pick the set that defines the fix boundary, not the prettiest example. One canonical reproduction, plus a few variants chosen because each defeats a different plausible cheap fix, plus benign inputs that must keep working after the change. Link the full log as an appendix and say why each attached case is there.

open as a page

The core evidence for an AI red-team finding is a model response containing genuinely harmful, actionable content, and the fix requires developers to reproduce it. Your issue tracker is readable by most of the company. How do you package the finding?

level: principalimportance: should knowfreq 29%

basics

~20 s

Split the package. The broadly readable ticket carries a characterisation of the harm, the criterion, the rate and the fix requirement. The reproduction input and the full response go to an access-controlled store, referenced by identifier and hash, granted to the engineers doing the fix. Agree the split, the retention and the deletion date in advance.

open as a page