You run a public harmful-behaviour benchmark such as HarmBench against a hosted chat endpoint and save the resulting attack-success rate for later comparison. What do you record alongside the number so it can still be interpreted in six months?
answer
- run manifest next to the rate
- endpoint plus resolved revision
- decoding settings and error accounting
- judge identity and threshold
- keep transcripts to re-judge later
basics
~20 sRecord the endpoint string and any pinned revision it resolved to, the timestamp, decoding settings such as temperature and max tokens, the dataset revision and which behaviours ran, the harm judge and its version, and the raw transcripts. Without these, a later number cannot be compared and any difference cannot be explained.
solid answer
~50 sSave a **run manifest** next to the rate, because the rate alone is uninterpretable later. Minimum contents: - **What was tested:** endpoint string, any pinned revision identifier the provider exposes, region or deployment if relevant, and whether a system prompt was attached. - **How it was driven:** decoding settings, max tokens, retry and timeout policy, concurrency, and how errors, empty responses and truncations were counted. - **What was asked:** benchmark dataset revision and the exact behaviour subset, since a rate over a different denominator is not a comparison. - **Who scored it:** the harm judge's identity and revision, and its threshold. - **The evidence:** raw prompts and responses. Transcripts are the highest-value item. With them you can re-judge an old run under today's scorer, which separates judge drift from model drift. Without them, the run can never be re-examined and only the aggregate survives — and the aggregate is exactly the part that cannot explain itself.
go deeper
Names timestamp, endpoint, decoding settings, dataset revision and judge, and knows the raw outputs should be kept.
Adds error and retry accounting, the system prompt hash, and explains that the denominator must match for two rates to be comparable.
Has the harness emit the manifest automatically, captures any served-revision metadata the provider returns, and sets a deliberate retention policy for sensitive transcripts.
Makes the manifest a required artefact for any safety number cited in a decision, so comparisons across teams and quarters are auditable rather than anecdotal.
### Why the rate alone is uninterpretable later A benchmark result is a compound claim: *this endpoint, driven this way, judged this way, on this date, over this behaviour set, produced this rate.* Every clause you fail to write down is a clause the next reader has to guess, and guessing is how two runs that were never comparable end up on the same trend line. The fix is unglamorous: emit a **run manifest** from the harness alongside every number, and keep the raw transcripts. ### What the manifest must contain ```yaml run_id: 2026-Q3-safety-suite-014 started_at: <iso-8601 timestamp> target: endpoint: <provider endpoint string> pinned_revision: <snapshot id, or 'alias-floating'> served_revision_seen: <id returned in response metadata, if any> system_prompt_sha256: <hash, or 'none'> sampling: temperature: 0.0 max_tokens: 512 attempts_per_behaviour: 5 success_rule: any_attempt_counts_behaviour on_error: retry_3_then_count_as_non_hit dataset: name: <benchmark name> revision: <commit or release id> behaviours_run: 210 judge: name: <harm judge> revision: <id> threshold: 0.5 harness: template_sha256: <hash of the prompt-building code> results: attack_success_rate: 0.14 attempts: 1050 errors: 3 transcripts: <restricted evidence store path> ``` Four fields people skip and later regret. **The served revision, not just the alias.** An endpoint name is a routing label; the weights and the serving-side filtering behind it can be replaced without it changing. If the provider echoes a model or snapshot identifier in response metadata, capture it — that is the only in-band evidence of what actually answered. **Error accounting.** A run where forty attempts timed out and were silently dropped reports a rate over a smaller denominator. Next quarter, with timeouts counted as non-hits instead, the rate falls with no change in behaviour whatsoever. Say explicitly which convention you used. **The success rule and attempts per behaviour.** "A behaviour counts as compromised if any of five attempts landed" and "the fraction of individual attempts that landed" are different statistics computed from the same transcripts, and the first is systematically larger. **Hashes of the system prompt and the templating code.** A harness change then shows up as a one-line manifest diff instead of as a mysterious movement in the rate. ### What it costs The manifest itself is free — a few dozen lines emitted by the harness, no extra model calls. The real costs are storage and handling of the transcripts. A suite of a couple of hundred behaviours at five attempts produces on the order of a thousand prompt/response pairs, a few megabytes of text: trivial to store, but it is a corpus of attempted harmful elicitations, so it belongs in the same restricted evidence store as the rest of the engagement, with an access list and a deliberately chosen retention period. Budget the review time too — someone must confirm the transcripts are complete and correctly redacted before they are archived. Keep the manifest even after the transcripts expire: a manifest with no transcripts still tells you whether the next run is comparable, which is most of the value. ### Where an unmanifested number misleads The failure is always the same shape: a rate moves, and the movement is read as the endpoint getting safer or riskier. With no manifest, that reading cannot be defended, because a changed judge revision, a revised behaviour set, a different temperature, a new retry policy, or a different success rule each produce the identical symptom. Worse, the aggregate is the one artefact that cannot explain itself, and it is the only one that survives into a slide deck. A number with no manifest is not a weak measurement; it is an uninterpretable one, and treating it as weak-but-usable is how it gets cited anyway. ### What you would check Before archiving a run: does the manifest name the judge revision and threshold; does it distinguish the alias from any served revision; does it state the error convention and the success rule; are the transcripts stored and readable; does the behaviour count match the dataset revision. Before comparing two runs: diff the two manifests first and only then look at the rates. Most apparent drift is found in that diff in under a minute.
- Why is the harm judge's revision as important as the model's?Because the judge decides which attempts count. A stricter or retrained judge moves the rate with no change in the endpoint at all, and only a recorded judge revision lets you rule that out.
- Which single artefact most improves a run's future value?The transcripts. They let you re-score an old run under a current judge, which is the only way to separate judge drift from model drift after the fact.
saying these in an interview costs you the question
- Saves only the aggregate rate and a model name.
- Does not record the harm judge or its threshold.
- Drops errored attempts silently, changing the denominator between runs.
- Keeps no transcripts, making the run impossible to re-judge.
- Records the alias only, never the revision the provider actually served.