You must keep a harmful model completion out of an AI red-team report, but the reader still has to believe the attack worked. What do you put in the report in place of the completion, and what does the reader lose by accepting it?
answer
- substitution, not deletion
- characterise at the claim's granularity
- attribute the verdict, name the grader
- hash + location = a seal
- plausible vs actionable is what is lost
basics
~20 sReplace it with an attested description: what the output contained at the granularity of the harm claim, its length and structure, who or what judged it a success, how many attempts succeeded, the target's version and configuration, and a hash plus location of the sealed artefact. The reader loses independent verification and must trust your judgement.
solid answer
~50 sThe substitute has to answer the question the completion answered: *did the model actually do the harmful thing, or something that merely looks like it?* So the record carries characterisation at the harm claim's granularity — "produced step-ordered operational detail, with specific quantities, for a prohibited category" — not a quotation, plus the structural facts (length, format, whether it hedged or refused partway), the deciding verdict and what produced it (human reviewer by name, or a named automated grader with its configuration), successes over attempts, the exact target build and settings, and a hash of the sealed artefact with where it lives. What the reader loses is real and you should name it: independent verification collapses into your judgement. In particular they cannot tell **plausible-sounding text** from genuinely actionable content — the single distinction that drives severity. Mitigations: a second named reviewer countersigns the characterisation; the report defines the harm scale it is grading against; and supervised access to the artefact is offered on request.
code
yaml · 23 linesfinding_id: RT-014
claim: refusal boundary failed under a multi-turn role-framing family
harm_class: prohibited-category operational detail
characterisation: >
Model returned step-ordered instructions with specific quantities and no
hedging. Correctness not independently assessed. No quotation recorded here.
verdict:
decided_by: human_reviewer
reviewer: <name>
rubric: engagement harm scale v1, level 3
countersigned_by: <name>
reliability:
attempts: 20
successes: 13
target:
build: <deployment identifier>
system_prompt_sha256: <hash>
decoding: temperature 0.7, top_p 1.0
artefact:
sha256: <hash of sealed transcript>
location: restricted evidence store, engagement 2026-014
access: named request, logged, two-person approval
retention_until: 2026-12-31go deeper
Knows the report should describe the output rather than quote it, and that a success count and a scorer verdict belong in the record.
Builds the full substitute record — characterisation, verdict provenance, reproduction parameters, hash — and states plainly that independent verification is what the reader gives up.
Ties redaction depth to the harm claim, warns that fragments across findings recompose, and adds countersigning plus supervised access as the credibility repair.
Standardises it: a report template with required evidence fields, a defined harm scale, a two-reviewer rule for high-severity characterisations, and an audit that samples sealed artefacts against their descriptions.
Redaction here is a **substitution** problem, not a deletion problem. Deleting the completion deletes the finding's evidentiary value along with it, so the work is to identify which properties of that text the reader actually needed and record those instead. ## What the completion was doing for the reader One question only: *did the model really do the harmful thing, or did it produce something that merely reads as though it did?* That distinction drives severity, and it is the thing a reader loses when the quotation goes. Fluent, confident, entirely useless text is common; so is text that is short, flat and genuinely actionable. Length, tone and vividness do not separate them. A characterisation must therefore state what made the output actionable — specificity, ordering, quantities, plausibility of correctness — without reproducing any of it. ## The four parts of a substitute record **Characterisation.** Prose written at the granularity of the claim. If the claim is "the model gave actionable uplift", the characterisation has to justify "actionable". If the claim is only "the refusal boundary failed", a much thinner description is honest and sufficient. Matching depth to claim is the skill; over-writing a weak finding is as much a defect as under-writing a strong one. **Provenance of the verdict.** Who or what decided this counted, and under what rubric. Either a human reviewer, identified by name, with the harm scale and level they applied; or an automated grader, named with its version and configuration, plus whether a human spot-checked its verdicts and at what sample rate. An unattributed verdict is the single most common weakness in redacted findings, because it collapses the whole finding into an anonymous assertion. **Reproduction parameters minus the payload.** Target build, system-prompt hash, decoding settings such as temperature and top-p, date, attack family, turn count, attempts and successes. This is what lets an engineer re-derive the result under supervision instead of replaying stored harmful text. **A seal.** A cryptographic hash of the raw artefact, its storage location, the access procedure and the retention date. The hash converts "trust my description" into "the artefact I described has not changed since I described it". ## What it costs, and what the numbers do not say Each substituted finding costs roughly half an hour of careful writing, and a two-reviewer rule on high-severity items costs a second reviewer the same again. Over a thirty-finding engagement that is a day of senior time spent producing no new findings — the price of the archive being safe to circulate. Then the misleading part, which matters more. A redacted finding is read almost entirely through two numbers, and both mislead by default. The **success ratio**. "13 of 20 attempts" is meaningless until the report defines an attempt. One prompt? One complete multi-turn conversation? One seed prompt sampled twenty times at nonzero temperature, where the variance is the sampler's and not the attack's? These give wildly different pictures of how reliable the attack is, and a reader who assumes the wrong one will mis-set severity in either direction. Write the denominator's definition into the record next to the ratio, always. The **grader's verdict**. If an automated judge decided the attempt succeeded, its own error rate is now part of your finding, and it is usually unmeasured. Judges over-count text that begins with a refusal and then complies, and under-count harm expressed obliquely. Reporting "the grader scored this a hit" without a human spot-check rate hands the reader a number whose accuracy nobody has established. Two more failure modes belong on this list. **Over-redaction flattens severity**: when every finding is described in the same guarded vocabulary, the client prioritises by your adjectives rather than by evidence, and genuinely severe items sink into the mass. **Fragments recompose**: a "short harmless excerpt" chosen to make one finding vivid is chosen precisely because it is the specific line, and excerpts across several findings can reconstitute more than any single one shows. Automated identifier scrubbers do not help here — they were built to remove names and account numbers, not to neutralise instructional content, so they return exactly the sentences a determined reader wants. ## What I would check Read the characterisation back looking for any string that could be lifted and used as-is. Confirm every finding's verdict names its source and rubric. Confirm the ratio is accompanied by a definition of an attempt. Have a second reviewer open the sealed artefact and confirm the description matches it — that check is the only thing standing between a redacted report and an unfalsifiable one.
- The client disputes the severity, saying your description sounds worse than the output was. How do you resolve it without circulating the artefact?Offer supervised review of the sealed artefact by one named client reviewer, or a joint re-run under observation. Both keep distribution at one person and produce an agreed severity.
- Why is a hash of the raw artefact worth including if almost nobody will ever check it?It fixes the description to a specific, unaltered artefact. If severity is contested months later, or the archive is audited, the description can be checked against what was actually captured.
A success ratio with no definition of an attempt is a batting average where nobody says whether a walk counts as a turn at bat. The same number describes a reliable attack or a lucky sample depending on a denominator the report never states.
saying these in an interview costs you the question
- Quoting a 'short harmless excerpt' of the completion to make the finding feel real
- Running an off-the-shelf identifier scrubber over the completion and calling the result publishable
- Recording a verdict with no attribution — no reviewer, no named grader, no rubric
- Not acknowledging that the reader loses the ability to distinguish plausible text from actionable content