skip to content

Mid-engagement, an AI red-team target emits output that falls into a class your organisation must not retain at all — not merely restrict. What do you do in the next few minutes, and how does the finding still reach the report?

level: seniorimportance: should knowfreq 33%

answer

  1. pre-agreed handling matrix, named contact
  2. copies: tool log, scrollback, CI, sync
  3. destroy under a witnessed record
  4. metadata + two-person attestation survives
  5. state the destruction in the report

basics

~20 s

Stop that line of testing, do not copy or forward the output, and treat the tool's own log as holding it too. Trigger the pre-agreed escalation to your named legal and trust-and-safety contact, isolate and destroy the artefacts under a witnessed record, and carry the finding as metadata plus a two-person attestation.

solid answer

~60 s

The decisive part is that this is a procedure you wrote **before** the engagement, because in the moment you will not be improvising a legal judgement. Immediate actions: stop that probe line; do not screenshot, paste, forward or 'save one copy as proof'; remember the harmful text is already on disk in the scanner's request/response log and in any terminal scrollback, so those are in scope too. Isolate the artefacts, notify the named escalation contact — legal plus a trust-and-safety owner — and follow whatever they direct, including any reporting obligation, which is their call and not yours. The finding survives as **metadata and attestation**: timestamp, target build, attack family, the verdict and who reached it, and a written statement signed by two people describing the class of output without reproducing it. Severity is then argued from the class and the reliability figures rather than from a sample. Record the destruction itself — what, when, by whom, witnessed — because an unexplained gap in the evidence chain reads worse than the incident.

go deeper

for a junior

Knows to stop, not to copy or forward the output, and to escalate to a named person rather than deciding alone.

for a middle

Adds that the material is also in the tool's log and scrollback, and that the finding continues as metadata plus a written description.

for a senior

Runs the pre-agreed procedure end to end: isolate every copy including CI, escalate to legal and trust-and-safety, destroy under a witnessed record, and document the destruction in the report so the gap reads as discipline.

for a principal

Owns the handling matrix itself — classes, storage rules, escalation and out-of-hours routes agreed with the client and counsel up front, plus a test-plan rule that stops eliciting material the organisation cannot hold.

Three things separate a competent answer from a nervous one: whether the procedure existed before the moment, whether you know how many copies already exist, and whether the finding survives the destruction. ## Pre-agreement is the whole answer Before testing begins, the engagement should carry a **handling matrix**: harm classes down one axis and, for each, whether the artefact may be stored normally, stored under restriction, or must never be retained — with a named escalation contact, a deputy, and an out-of-hours route. Without that document you are asking a tester to make, alone and under time pressure, a decision with legal consequences in a jurisdiction they have not studied. Whether something is a mandatory-report class, and to whom, is not a call a red-teamer makes on the day. The matrix does a second job that is easy to miss: it is agreed with the client and counsel in advance, so when you destroy an artefact three weeks later, the destruction reads as a control that was already in force rather than as a tester deciding unilaterally to remove evidence. That distinction is what keeps the word "tampering" out of the conversation. ## Scope of the copies The output is not only on screen. By the time you notice, it is plausibly in: the tool's raw run report or memory database; terminal scrollback and shell history; an editor's autosave or undo history; the clipboard; a CI job's uploaded artefacts and console log; a cloud-synced folder that has already replicated it; and the endpoint operator's own logs, which are not yours to purge. Isolation means taking the whole set out of circulation, not closing a window. If the material reached CI, the incident is already larger than your laptop and the pipeline's artefact store becomes part of the response. ## Immediate actions Stop that probe line. Do not screenshot, forward, paste into chat to ask a colleague, or "save one copy as proof" — that last one is the instinct that turns a handled incident into a personal-liability problem. Isolate the copies you can reach. Notify the named contact, which in practice means legal plus a trust-and-safety owner, and follow what they direct, including any reporting obligation. Then destroy under a **witnessed record**: what was destroyed, from which locations, when, by whom, on whose authority, countersigned. ## How the finding still reaches the report Destroying the artefact costs the proof, so the substitute is built deliberately. Safe-to-hold **metadata**: timestamp, target build and configuration, attack family, turn count, attempts and successes, the verdict and its source. An **attestation** written and countersigned by two people, describing the class of output and what made it that class, with no quotation and no distinguishing detail. And an explicit sentence in the report stating that the artefact was destroyed under the agreed procedure, on whose instruction, on what date. That sentence is what makes an absence read as discipline rather than as a gap. ## What it costs, and where the numbers mislead An escalation of this kind consumes hours to days of counsel's time, and it commonly pauses the affected test line for the rest of the day or longer, which is engagement hours the client is paying for. Two people are tied up in the destruction record. Budget for at least one such event on any engagement touching high-harm classes, rather than treating it as an exception that will not happen. Two readings mislead badly. First, **the reliability figure on a destroyed finding**. If the class was reached once in three hundred attempts, "1/300" is not an estimate of anything — a single observation supports no rate, and reporting it as a low rate invites the client to deprioritise a boundary failure that is real. Say it was reached, say under what conditions, and resist converting one event into a frequency. Second, **"we deleted it"**. Deleting the visible file is the smallest part; the memory store row, the scrollback, the sync replica and the backup generation all persist independently, and a destruction record that lists only the file is a record of an incomplete action. There is also a test-plan consequence worth raising unprompted: if this class is reachable, continuing to elicit it generates repeat handling incidents without adding information. Establish the boundary once, characterise it, and shift to probing whether the *control* holds under the attack family using near-boundary prompts whose outputs you are permitted to retain. ## What I would check afterwards That the destruction record exists and is countersigned. That no copy remains in backups, CI or the tool's own store. That the handling matrix is updated with the class you have just proved is reachable, so the next tester meets a procedure rather than a surprise.

  • Why must the destruction be recorded rather than simply done?
    A silent gap in the evidence chain looks like tampering later. A dated, countersigned record naming what was destroyed and on whose instruction turns the gap into a documented control.
  • How do you keep testing that area at all once you know the class is reachable?
    Shift from eliciting content to probing the control: measure whether the refusal holds under the attack family, using near-boundary prompts whose outputs you are permitted to retain.

saying these in an interview costs you the question

  • Keeping a private copy 'as proof' after the artefact was ordered destroyed
  • Deciding alone, in the moment, what the legal obligation is
  • Deleting quietly with no destruction record, so the evidence chain has an unexplained gap
  • Isolating only the on-screen copy while the scanner's raw log, scrollback and CI artefacts still hold it
  • Continuing to probe the same class after the boundary is established, generating repeat handling incidents

context