In an AI red-team engagement, why is pasting the full successful attack transcript — the prompts plus the model's harmful completion — into the circulated report a bad default, when that transcript is exactly what proves the finding?
answer
- proof and payload in one artefact
- report distribution always grows
- tiered deliverable: claim vs evidence
- scoped access, not a wider report
- check appendices and screenshots
basics
~20 sBecause the report travels much further than the evidence should. A full transcript is a validated working attack plus the harmful text itself, readable by everyone the document reaches. Keep the raw prompt and completion in a restricted evidence store, and put a description, the harm class and a pointer in the report.
solid answer
~50 sThe transcript does two jobs at once: it is the proof and it is the payload. Reports get emailed, attached to tickets, pasted into chat and dropped in shared drives, read by people who signed nothing about handling this material — so shipping the payload inside the document hands a validated working attack, and the harmful output itself, to everyone downstream of delivery. The normal shape is a **tiered deliverable**. The circulated document carries the claim, the harm class, the conditions under which the model produced it, how often it reproduced, and the verdict of whatever decided the attempt counted. The prompt-and-completion pair lives in a restricted store with named access and a retention date. The break condition is a fixer who genuinely cannot reproduce or regression-test from the description. The answer then is scoped access for that person to the evidence store — not widening the report.
go deeper
Says the transcript is dangerous to circulate, and that raw output belongs somewhere access-controlled rather than in the report body.
Describes the tiered deliverable concretely: which fields go in the report, what stays in the evidence store, and how a fixer gets access.
Adds retention as a separate axis from access, names the leak paths (tickets, chat, screenshots, tool logs), and handles the client who demands the raw proof.
Frames it as a policy the whole practice runs on — a standard report template, a classified evidence store with an owner and a retention schedule, and an exception path signed by someone outside the red team.
Two separate risks travel inside one artefact, and a candidate who sees only one of them has answered half the question. ## What the artefact actually is A red-team transcript for a language-model finding is a pair. One half is the **prompt sequence** the tester sent: possibly a single message, more often a multi-turn conversation, together with the system prompt, tool definitions and retrieved context the target was running under at the time. The other half is the **completion** the model returned. Both halves are evidence, and they carry different dangers. The prompt side is not a hypothesis. It is a sequence someone already confirmed defeats a control that is live in production right now. Handing it to a reader is handing them a working procedure against a system they may not be authorised to test. The completion side is the harmful text itself: material the client may be restricted from holding, or in some classes must not hold at all, regardless of who can open it. Access controls answer the first problem; only retention rules answer the second. Conflating them is the single most common error here. ## Why a report is the wrong container The reason is distribution, and specifically that a report's readership is not the address line. It is the transitive closure of everyone that line forwards to, plus every system the file passes through on the way: the remediation ticket it gets attached to, the vendor-management thread, the steering-committee slide, the shared drive whose search index now returns the payload for a keyword query, the assistant someone points at that drive to summarise the quarter. None of those hops is re-authorised. The set only grows, and it grows after you have stopped watching it. ## The shape that works: a tiered deliverable The circulated document carries the **claim**; a restricted store carries the **substantiation**. | goes in the report | stays in the evidence store | |---|---| | finding id, the behaviour elicited, harm class | the verbatim prompt sequence | | target build, system-prompt hash, decoding settings | the raw completion | | attack family named, never quoted | the tool's own request/response log | | attempts and successes, and what an attempt was | access by name, every read logged | | the verdict and who or what reached it | a retention date, and an owner | | a stable pointer: hash, location, access procedure | | ## What it costs Writing a characterisation that survives a challenge takes roughly twenty to forty minutes per finding on top of finding it, and a countersigning reviewer doubles the senior review time on the high-severity ones. A thirty-finding engagement therefore absorbs a day or more of effort that produces no new findings. Standing up and maintaining the evidence store is real engineering time; a supervised reproduction session for a client engineer costs an hour of two people plus scheduling. Budget these at proposal time, because the alternative — pasting the transcript in — is free, which is exactly why it is the default that has to be argued against. ## Where the reading misleads Four specific misreadings do most of the damage. First, **treating a label as a control**: "confidential" in a header and a password on the PDF change nothing, because the password travels in the same email and a screenshot leaves no label behind at all. Second, **cutting the wrong half**: redacting the model's output while keeping the prompt feels like restraint, but the prompt is the reusable half, so this inverts the risk it was meant to reduce. Third, **counting the distribution list at delivery**: the number of readers is measured once, at the moment it is smallest, and nobody re-measures it in month three. Fourth, **reading absence of a transcript as absence of proof**: a client under remediation pressure has an incentive to treat any unquoted finding as unproven, and you will lose that argument unless the report says up front what substitutes for the quotation and offers supervised access on request. ## What I would check before sending Search the deliverable and every appendix for any string from the completion. Open every screenshot, because an image of a terminal is fully readable and not searchable, so it survives text scrubbing. Check document metadata, tracked changes and comment history, where an earlier draft's quotation often still lives. Check whether an automated scanner's raw run log was attached "for completeness". Finally, resolve the pointer yourself: confirm the hash matches, the artefact is where the report says, and the named access procedure is one a real person could actually follow.
- The client's engineer says they cannot fix what they cannot reproduce, and demands the exact prompt. What do you do?Grant that named person scoped, logged access to the evidence store, or run the reproduction with them in a supervised session. You widen access to one person under a record, not the report to everyone.
- Where does the harmful completion typically leak from even when the report itself is clean?The scanner's own request/response log, CI job artefacts and console output, ticket attachments, screenshots pasted into chat, and a working directory sitting inside a cloud-synced folder.
- Does redacting the payload but keeping the attack prompt make it safe to circulate?No. The prompt is the reusable half — it is the recipe. Keeping it and cutting the output inverts the risk you were trying to reduce.
The transcript is both the photograph of the crime scene and the loaded weapon that made it. A report is a container built to circulate photographs, and everyone it circulates to ends up handling the weapon as well.
saying these in an interview costs you the question
- Treating it purely as a confidentiality problem, with no notion that some output must not be retained at all
- Believing that marking the document 'confidential' or password-protecting the PDF solves the distribution problem
- Removing the evidence entirely with no substitute, leaving a finding the client cannot evaluate
- Keeping a personal copy of the raw transcript 'in case they argue'