What must you record about an AI red-team run so that repeating it months later is a comparison against the filed result rather than a fresh experiment?
answer
- record config, not just rates
- prompt hash + guard versions
- decoding params and seed honoured?
- the judge is a versioned dependency
- keep raw transcripts for re-scoring
basics
~20 sRecord the whole configuration, not the numbers: the concrete model identifier and any version metadata returned, decoding parameters and seeds, a hash of the system prompt and context, the guardrail versions and thresholds, the exact attack prompts and harness version, the deciding component's version, trials per prompt, and the raw transcripts.
solid answer
~50 sA rate without its configuration is not a result. The run manifest needs five groups: 1. **Target identity** — the model identifier requested, whatever version metadata came back, the endpoint and any gateway in the path. 2. **Sampling settings** — temperature, top-p, max tokens and any seed the API honours, since comparing across different decoding settings measures the settings. 3. **The application around the model** — hashes of the system prompt and tool descriptions, the retrieval corpus state, and the guardrail stack: classifiers, versions, thresholds, fail open or closed. 4. **The instrument** — harness version, the exact attack prompt set (plus generator and seed if generated), and the version of the component that decides whether an attempt counted. 5. **The evidence** — trials per attack prompt and the raw transcripts, not just aggregate rates. The most-missed item is that last decider. If it is itself a model, it drifts too, and a rerun can show a changed rate with the target untouched.
go deeper
Lists the obvious fields: model identifier, date, prompts used, and the pass/fail numbers with trial counts.
Covers decoding parameters, system-prompt and attack-set hashes, guard versions and thresholds, and keeps raw transcripts rather than only aggregates.
Includes the deciding component as a versioned dependency, re-scores old transcripts with the current decider to separate judge drift from target drift, and records cost so a rerun can be budgeted.
Makes the manifest a programme artefact — every filed result carries its fingerprint and expires when the fingerprint changes — and negotiates the fields the platform team must emit for it.
The manifest exists to make a future disagreement decidable. Six months from now two people will point at two numbers and argue about what changed; everything not written down at run time becomes opinion. A success rate with no configuration attached is not a result, it is an anecdote with a decimal point. ### What goes in it A workable shape, written by the harness itself and stored beside the raw transcripts: ```yaml run_id: 2026-04-11-full-suite target: endpoint: chat.internal/v1/completions model_requested: <pinned dated identifier> model_metadata_returned: <as returned, verbatim> gateway: <proxy name/version, if any> sampling: temperature: 0.0 top_p: 1.0 max_tokens: 1024 seed: 12345 # if honoured; record that it was not, if not application: system_prompt_sha256: <hash> tool_defs_sha256: <hash> retrieval_snapshot: <corpus version or date> guards: input_classifier: <name/version, thresholds, fail-open?> output_classifier: <name/version, thresholds> instrument: harness_version: <version> attack_prompt_set_sha256: <hash> generated_by: <generator + seed, if dynamic> decider: <the component judging a hit: version; if a model, its id + prompt hash> execution: trials_per_prompt: 20 attempted / completed: 4400 / 4400 metered_cost: <units> ``` Five groups, and each earns its place. **Target identity** says what answered, including the gateway, because a proxy in the path can rewrite metadata and route requests. **Sampling settings** matter because comparing two runs at different temperatures measures the temperature. **The application around the model** — prompt hash, tool definitions, retrieval corpus state, guard versions and thresholds — is where most silent change actually happens. **The instrument** is the half people forget: the harness, the exact attack prompt set, and the component that decides whether an attempt counted. **The evidence** is trials per prompt plus the raw transcripts, because aggregates cannot be re-examined. ### What it buys *Comparability*: a rerun whose manifest differs is a different experiment, and the manifest tells you which line moved. *Re-scoring*: with transcripts retained you can re-decide the OLD run with the CURRENT judge, which cleanly separates judge drift from target drift — the single most common false regression in this work. *Budgeting*: recorded request counts and metered cost are what let you tell a stakeholder what an out-of-cadence rerun will cost before you call one. ### What it costs Recording is cheap in money — a full suite's transcripts are tens of megabytes — and expensive in policy. Raw transcripts of a red-team run contain the generated harmful content the run elicited, so they need encryption, access control, a named retention window and a defensible answer about who may read them. The cost of *not* recording is larger and lands later: an unreproducible number means a full metered rerun to settle an argument, at exactly the moment someone is waiting on a release decision. ### Where the number misleads A rate quoted without its denominator is the commonest trap: "12% success" is 3 of 25 or 120 of 1000, and those two carry wildly different intervals while printing the same digits. Next is the partially completed run — rate limits, a timeout or a truncated sweep — reported against the full-suite denominator, which understates the rate and overstates the coverage simultaneously. Then judge drift: if the decider is itself a model or a versioned classifier, the rate can move with the target untouched, and without transcripts you cannot tell. Then the regenerated prompt set: a harness that samples fresh attack prompts each run changes the measuring instrument between runs, so no two rates share a denominator. Finally the false determinism claim — a `seed` parameter recorded in the manifest but never honoured by the endpoint is worse than no seed at all, because it converts an unknown into a wrong belief. ### What you check Recompute the hashes at run time rather than copying them forward from the previous manifest — a stale copied hash is how a changed prompt gets recorded as unchanged. Confirm the response metadata was captured verbatim rather than the value you requested. Verify the seed was honoured by issuing one duplicate request. Check that attempted and completed counts match, and that the denominator in any reported rate is the completed count. And generate the manifest from the harness, never by hand: a manifest a human types is a manifest that describes what someone believed they ran.
- Why keep raw transcripts when you already have the pass/fail rate?So you can re-decide the old run with the current judging component. If the rate moves only on re-scoring, the judge drifted, not the target — and without transcripts you cannot tell those apart.
- You must drop one field to keep the manifest small. Which goes last, and which would you never drop?Latency and similar telemetry can go. Never drop the model identifier and metadata, the system-prompt hash, the guard versions and the attack-set hash — those four are what make any later number comparable.
saying these in an interview costs you the question
- Filing only aggregate success rates with no configuration and no transcripts.
- Regenerating the attack prompt set on every run, so no two runs share a denominator.
- Treating the component that decides whether an attempt counted as fixed infrastructure with no version.
- Claiming reproducibility from a seed parameter without checking the endpoint honoured it.