How do you keep an AI red-team report from being reused months later as a standing safety certificate for a system that has since changed — a swapped model, an edited system prompt, a new tool integration — and what goes into the deliverable to make that explicit?
answer
- fingerprint the tested configuration
- list invalidating changes per section
- validity window + re-test triggers
- non-determinism: same config, different result
- coverage ledger across engagements
basics
~20 sBound the result to the exact system you tested. Record the configuration you exercised, state that the conclusions hold only for it, and name the changes that void them: a model swap, a system-prompt edit, a new tool or data source. Add a validity window and explicit re-test triggers.
solid answer
~60 sA report describes one configuration at one moment, but it gets filed as a permanent artefact and cited in questionnaires long after the system has moved. Three things stop that. **Fingerprint the tested system.** An appendix that pins what was actually driven: model identifier and serving mode, system-prompt and policy-file digests, guardrail configuration, retrieval index state, the tool and integration list, and the endpoint or environment. Conclusions are stated as holding for that fingerprint. **Name the invalidating changes.** Do not leave "if the system changes" to the reader's judgement — list the specific events that void which conclusions: a model or serving-tier change, any edit to the system prompt or policy text, a new tool or data source reaching context, a guardrail config change, a new entry point. **Give it a validity window and re-test triggers**, tied to those events rather than to the calendar alone. Organisationally, back it with a rule that any assurance claim must cite a specific run and its fingerprint. Nobody re-reads a caveat; a citation requirement forces the lookup.
go deeper
Says the report should state which version or configuration was tested and its date.
Adds a configuration appendix — model, prompts, guardrail settings, tools, retrieval state — and a re-test recommendation.
Maps invalidating changes to the specific conclusions they void, and states that a clean result was probabilistic and sample-bounded even on an unchanged system.
Moves the guarantee into process: a coverage ledger across engagements, a rule that assurance claims cite a run and fingerprint, and an explicit stance on provider-side drift under a stable model identifier.
The failure mode here is procurement, not testing. A competent report from March gets attached to a security questionnaire in November, for a system whose system prompt was edited in April and whose model was swapped in June. Nobody lied at any point. The deliverable simply had no expiry, and nothing in it made the change detectable to the person quoting it. **Fingerprint what you actually drove** An appendix that pins the identifiers determining behaviour, so that every conclusion can be phrased against it rather than against “the system”: ```text APPENDIX A — TESTED CONFIGURATION model provider id + serving tier + any version string exposed sampling temperature / top-p, if pinned during the run system prompt SHA-256 of the exact text used tool definitions SHA-256 of the tool/function description bundle; enabled tool list guardrail config product + version + thresholds per category retrieval index identity + snapshot date + document count environment endpoint, tenant, auth mode, dates and times of each run ``` Digests rather than the text itself: they let anyone re-compute and compare without the appendix carrying the prompt into a document that circulates. The same appendix makes the engagement reproducible, which is the same property viewed from the other end. **Name the invalidating changes, mapped to conclusions** “Results may not apply if the system changes” is unactionable, because it delegates the judgement to the person least able to make it. List the events and say what each one kills: ```text VALIDITY Results describe Appendix A only. Voided by: model or serving-tier change (all sections) | system-prompt or policy edit (4, 6) | new tool or data source reaching context (5, 6) | guardrail config or threshold change (7) | new entry point (coverage declaration, 3). Re-test: on any trigger above, or 6 months, whichever comes first. ``` Mapping triggers to sections means a small change costs a partial re-test rather than discarding the report — which is what makes the rule survive contact with a team that ships weekly. **What it costs** The fingerprint is an hour at the end of the engagement, mostly digest computation and asking two questions of the platform team. A triggered partial re-test is the real expense: re-driving one section is typically half a day of setup plus the model calls for that section's attempts — hundreds of dollars, not tens of thousands — whereas a full re-run is the original engagement again. That asymmetry is the argument that sells the section mapping: without it, every prompt edit nominally invalidates everything, so in practice teams invalidate nothing. **Where the number misleads — two things a fingerprint cannot fix** First, an unchanged configuration does not give an unchanged result. Generation is non-deterministic and attack success is probabilistic, so a re-run scoring worse is not evidence of regression until you know the sampling variance. The arithmetic worth carrying: zero hits in 50 attempts is consistent with a true success rate up to roughly 6% (the rule of three, upper bound ≈ 3/n). A clean row was always an interval, never a zero, and per-item attempt counts are what let a reader see how wide that interval is. Second, hosted models change beneath a stable identifier. The provider may alter weights, serving stack or system-side policy without changing the string you recorded, so a matching model id proves less than it looks. Say what you could pin, and state provider-side drift as an explicit assumption whose only detection method is re-testing. **Make the process carry it, not the prose** Caveats do not survive being quoted; lookups do. Three mechanisms that work: a coverage ledger maintained across engagements rather than per report, so the organisation tracks which surfaces have ever been driven and when; a rule that any external assurance claim must cite a run identifier and its fingerprint, which forces someone to open the artefact; and a short front-of-report block written for whoever reads only page one. The report can make the lookup possible; only the process can make it happen. **What I check before sending** Can a reader re-compute every digest in Appendix A from artefacts they control? Does each trigger name the sections it voids rather than the whole document? Does every clean result carry its attempt count, so its interval is visible? Is provider drift stated as an assumption rather than assumed away? And is there any sentence in the report that would still read as an endorsement six months from now if quoted with the date stripped off?
- Why is a date-based expiry alone insufficient?Systems change on deploys, not on calendars. A prompt edit the week after delivery can void a conclusion the report says is valid for six months, so triggers have to be change-based with the date as a backstop.
- You cannot pin the weights of a hosted model behind a stable identifier. What do you record?Everything you can pin — identifier, serving mode and any version string the provider exposes, dates and times of the run, sampling settings — and state provider-side drift as an explicit assumption, with a note that re-testing is the only way to detect it.
saying these in an interview costs you the question
- A report with no record of the configuration that was tested
- A generic 'results may not apply if the system changes' with nothing named
- Treating a re-run on an unchanged configuration as guaranteed to reproduce
- Assuming a stable hosted-model identifier means unchanged behaviour
- Relying on caveats in prose to survive being quoted into a questionnaire