skip to content

Handling Harmful Output

The text that proves a jailbreak worked is also the most dangerous material in the engagement, so retention, redaction and access become decisions rather than defaults. Interviewers ask who can open your archive.

on this pageshow

explore

questions

5

In an AI red-team engagement, why is pasting the full successful attack transcript — the prompts plus the model's harmful completion — into the circulated report a bad default, when that transcript is exactly what proves the finding?

level: juniorimportance: must knowfreq 62%

answer

  1. proof and payload in one artefact
  2. report distribution always grows
  3. tiered deliverable: claim vs evidence
  4. scoped access, not a wider report
  5. check appendices and screenshots

basics

~20 s

Because the report travels much further than the evidence should. A full transcript is a validated working attack plus the harmful text itself, readable by everyone the document reaches. Keep the raw prompt and completion in a restricted evidence store, and put a description, the harm class and a pointer in the report.

solid answer

~50 s

The transcript does two jobs at once: it is the proof and it is the payload. Reports get emailed, attached to tickets, pasted into chat and dropped in shared drives, read by people who signed nothing about handling this material — so shipping the payload inside the document hands a validated working attack, and the harmful output itself, to everyone downstream of delivery. The normal shape is a **tiered deliverable**. The circulated document carries the claim, the harm class, the conditions under which the model produced it, how often it reproduced, and the verdict of whatever decided the attempt counted. The prompt-and-completion pair lives in a restricted store with named access and a retention date. The break condition is a fixer who genuinely cannot reproduce or regression-test from the description. The answer then is scoped access for that person to the evidence store — not widening the report.

go deeper

for a junior

Says the transcript is dangerous to circulate, and that raw output belongs somewhere access-controlled rather than in the report body.

for a middle

Describes the tiered deliverable concretely: which fields go in the report, what stays in the evidence store, and how a fixer gets access.

for a senior

Adds retention as a separate axis from access, names the leak paths (tickets, chat, screenshots, tool logs), and handles the client who demands the raw proof.

for a principal

Frames it as a policy the whole practice runs on — a standard report template, a classified evidence store with an owner and a retention schedule, and an exception path signed by someone outside the red team.

Two separate risks travel inside one artefact, and a candidate who sees only one of them has answered half the question. ## What the artefact actually is A red-team transcript for a language-model finding is a pair. One half is the **prompt sequence** the tester sent: possibly a single message, more often a multi-turn conversation, together with the system prompt, tool definitions and retrieved context the target was running under at the time. The other half is the **completion** the model returned. Both halves are evidence, and they carry different dangers. The prompt side is not a hypothesis. It is a sequence someone already confirmed defeats a control that is live in production right now. Handing it to a reader is handing them a working procedure against a system they may not be authorised to test. The completion side is the harmful text itself: material the client may be restricted from holding, or in some classes must not hold at all, regardless of who can open it. Access controls answer the first problem; only retention rules answer the second. Conflating them is the single most common error here. ## Why a report is the wrong container The reason is distribution, and specifically that a report's readership is not the address line. It is the transitive closure of everyone that line forwards to, plus every system the file passes through on the way: the remediation ticket it gets attached to, the vendor-management thread, the steering-committee slide, the shared drive whose search index now returns the payload for a keyword query, the assistant someone points at that drive to summarise the quarter. None of those hops is re-authorised. The set only grows, and it grows after you have stopped watching it. ## The shape that works: a tiered deliverable The circulated document carries the **claim**; a restricted store carries the **substantiation**. | goes in the report | stays in the evidence store | |---|---| | finding id, the behaviour elicited, harm class | the verbatim prompt sequence | | target build, system-prompt hash, decoding settings | the raw completion | | attack family named, never quoted | the tool's own request/response log | | attempts and successes, and what an attempt was | access by name, every read logged | | the verdict and who or what reached it | a retention date, and an owner | | a stable pointer: hash, location, access procedure | | ## What it costs Writing a characterisation that survives a challenge takes roughly twenty to forty minutes per finding on top of finding it, and a countersigning reviewer doubles the senior review time on the high-severity ones. A thirty-finding engagement therefore absorbs a day or more of effort that produces no new findings. Standing up and maintaining the evidence store is real engineering time; a supervised reproduction session for a client engineer costs an hour of two people plus scheduling. Budget these at proposal time, because the alternative — pasting the transcript in — is free, which is exactly why it is the default that has to be argued against. ## Where the reading misleads Four specific misreadings do most of the damage. First, **treating a label as a control**: "confidential" in a header and a password on the PDF change nothing, because the password travels in the same email and a screenshot leaves no label behind at all. Second, **cutting the wrong half**: redacting the model's output while keeping the prompt feels like restraint, but the prompt is the reusable half, so this inverts the risk it was meant to reduce. Third, **counting the distribution list at delivery**: the number of readers is measured once, at the moment it is smallest, and nobody re-measures it in month three. Fourth, **reading absence of a transcript as absence of proof**: a client under remediation pressure has an incentive to treat any unquoted finding as unproven, and you will lose that argument unless the report says up front what substitutes for the quotation and offers supervised access on request. ## What I would check before sending Search the deliverable and every appendix for any string from the completion. Open every screenshot, because an image of a terminal is fully readable and not searchable, so it survives text scrubbing. Check document metadata, tracked changes and comment history, where an earlier draft's quotation often still lives. Check whether an automated scanner's raw run log was attached "for completeness". Finally, resolve the pointer yourself: confirm the hash matches, the artefact is where the report says, and the named access procedure is one a real person could actually follow.

  • The client's engineer says they cannot fix what they cannot reproduce, and demands the exact prompt. What do you do?
    Grant that named person scoped, logged access to the evidence store, or run the reproduction with them in a supervised session. You widen access to one person under a record, not the report to everyone.
  • Where does the harmful completion typically leak from even when the report itself is clean?
    The scanner's own request/response log, CI job artefacts and console output, ticket attachments, screenshots pasted into chat, and a working directory sitting inside a cloud-synced folder.
  • Does redacting the payload but keeping the attack prompt make it safe to circulate?
    No. The prompt is the reusable half — it is the recipe. Keeping it and cutting the output inverts the risk you were trying to reduce.

The transcript is both the photograph of the crime scene and the loaded weapon that made it. A report is a container built to circulate photographs, and everyone it circulates to ends up handling the weapon as well.

saying these in an interview costs you the question

  • Treating it purely as a confidentiality problem, with no notion that some output must not be retained at all
  • Believing that marking the document 'confidential' or password-protecting the PDF solves the distribution problem
  • Removing the evidence entirely with no substitute, leaving a finding the client cannot evaluate
  • Keeping a personal copy of the raw transcript 'in case they argue'

context

open as a page

You must keep a harmful model completion out of an AI red-team report, but the reader still has to believe the attack worked. What do you put in the report in place of the completion, and what does the reader lose by accepting it?

level: middleimportance: must knowfreq 52%

basics

~20 s

Replace it with an attested description: what the output contained at the granularity of the harm claim, its length and structure, who or what judged it a success, how many attempts succeeded, the target's version and configuration, and a hash plus location of the sealed artefact. The reader loses independent verification and must trust your judgement.

open as a page

Mid-engagement, an AI red-team target emits output that falls into a class your organisation must not retain at all — not merely restrict. What do you do in the next few minutes, and how does the finding still reach the report?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Stop that line of testing, do not copy or forward the output, and treat the tool's own log as holding it too. Trigger the pre-agreed escalation to your named legal and trust-and-safety contact, isolate and destroy the artefacts under a witnessed record, and carry the finding as metadata plus a two-person attestation.

open as a page

An automated LLM red-team scanner writes every request and response, including successful harmful completions, to log files in its working directory. Where do those files realistically end up during an engagement, and what do you change before the first run?

level: seniorimportance: should knowfreq 44%

basics

~20 s

They spread: CI job artefacts and console output, the runner's disk after the job, ticket attachments, screenshots in chat, cloud-synced home folders, laptop backups. Before the first run, move the working directory onto encrypted, unsynced storage in a per-engagement folder, stop artefact upload or restrict and expire it, and schedule destruction.

open as a page

Continuous automated LLM red-team suites make successful harmful outputs pile up as an archive that doubles as the regression set proving fixes hold. How do you decide what is kept, for how long, and who may open it?

level: principalimportance: should knowfreq 28%

basics

~20 s

Split the artefact. Attack inputs are what regression testing needs, so keep those under access control; harmful completions are usually re-derivable by re-running and re-scoring, so keep only a verdict and a hash. Set per-class retention with default expiry, two-person access with logging, and a named owner outside the red team.

open as a page