skip to content

The core evidence for an AI red-team finding is a model response containing genuinely harmful, actionable content, and the fix requires developers to reproduce it. Your issue tracker is readable by most of the company. How do you package the finding?

level: principalimportance: should knowfreq 29%

answer

  1. two tiers: characterise open, verbatim restricted
  2. hash plus pointer in the ticket
  3. split by role, not by seniority
  4. retention and deletion date
  5. over-redaction hides unverifiability

basics

~20 s

Split the package. The broadly readable ticket carries a characterisation of the harm, the criterion, the rate and the fix requirement. The reproduction input and the full response go to an access-controlled store, referenced by identifier and hash, granted to the engineers doing the fix. Agree the split, the retention and the deletion date in advance.

solid answer

~50 s

Two obligations collide. Verification needs the exact input and output; distribution control says that content should not sit in a tracker most of the company reads, get indexed, or ride along in notification email. The workable shape is a two-tier package: - **Tier 1, the ticket everyone sees**: the failure in characterised terms, why it counts, the observed rate with its denominator, the affected surface, the two-sided acceptance criterion, and a pointer plus content hash for the evidence. - **Tier 2, restricted**: the verbatim prompt and response, in a store with per-person access, an access log, and a stated retention and deletion point. Split by role: whoever must reproduce gets tier 2; whoever must prioritise does not. Redact inside tier 2 only where the redacted span is genuinely not needed to reproduce — over-redaction quietly makes the finding unverifiable, and that fails late. Agree the policy with legal and security before an engagement, not inside a ticket.

go deeper

for a junior

Recognises that the raw harmful output should not be pasted into a widely readable ticket and asks before filing.

for a middle

Separates the characterised description from the verbatim artefact and puts the artefact somewhere access-controlled.

for a senior

Designs the split by role, keeps the reproduction genuinely verifiable, and refuses redaction that silently breaks verification.

for a principal

Sets the handling policy with legal and security before the engagement — tiers, access path, retention, deletion, escalation — and makes it the default template rather than a per-ticket negotiation.

This is a governance decision dressed as a packaging question, which is why it lands on a lead rather than on the operator who found it. ## Why both naive options fail Pasting the content into the open ticket distributes it far beyond the people who need it, and the act is effectively irrevocable. Withholding it entirely leaves developers unable to reproduce, so they either close the ticket or — worse — reconstruct the attack themselves, generating the same content again in less controlled places and without your capture discipline. ## The two tiers, mechanically **Tier 1, the ticket everyone can read.** The failure in characterised terms, why it counts, the observed rate with its denominator, the affected surface, the two-sided acceptance criterion (attacks stop, benign controls still pass), the capture timestamp, the model or deployment identifier, and a cryptographic digest — a SHA-256 over the exact bytes of the restricted artefact — plus a pointer to where that artefact lives. **Tier 2, restricted.** The verbatim prompt and response, in a store with per-person grants rather than group-wide read, an access log, and a stated retention period with a named deletion owner and date. Split the tiers by *role*, not seniority. Prioritisation needs characterisation, impact, reachability and rate — tier 1 alone. Fixing needs the exact input — tier 2, for a small, named, time-bounded group. Verification needs the criterion and the repeat procedure, most of which is tier 1. Redact inside tier 2 only where the redacted span is genuinely not needed to reproduce. ## What it costs The setup is a one-time cost: a store, a grant path, a template, and a handling policy agreed with counsel and security *before* the engagement rather than improvised inside a ticket at filing time. Per finding it adds perhaps fifteen minutes and one access request. The cost of not having it is asymmetric and irreversible: a tracker post fans out into notification email, mobile push previews, integration mirrors, nightly exports, backups and a search index within seconds, and no subsequent delete recovers those copies. Fifteen minutes against an unrecallable distribution event is not a close call. ## Where the number misleads Two numbers lie here, and both lie quietly. The first is the apparent readership. A tracker shows a watcher count and a view count, and those small numbers get read as "only a few people saw it" — while the body has already been mailed in full to every watcher, mirrored into a chat integration, indexed for search and captured in an export. The view count measures the user interface, not the distribution, and it is the number people cite when arguing the content is contained. The second is verification after over-redaction. The reflex is to redact aggressively, and the cost lands weeks later: a developer re-runs a reproduction whose load-bearing span is gone, sees a clean result, and marks the finding verified. In the tracker that looks identical to a real fix — same status, same closure date, same green metric. A finding closed on a reproduction that could not have reproduced anything is worse than an open one, because it removes the item from the queue. If a span must be redacted and the reproduction depends on it, the answer is a different mechanism — a supervised reproduction session, or verification performed by the red team with the developer watching — not a ticket that merely looks complete. A digest has its own limit in the same spirit: it proves identity, never availability. If the artefact is deleted on schedule and the ticket is reopened a year later, the hash confirms nothing anyone can still read. ## What to check before you file Read tier 1 alone and ask whether a prioritiser could rank it without opening anything else. Read tier 2 alone and ask whether a named engineer could reproduce from it. Confirm the access list is named, logged and time-bounded rather than a group with permanent read. Confirm there is a deletion date and an owner for executing it, and that the deletion itself gets recorded. Have someone outside the fix group read tier 1 and confirm it carries no operative detail — characterisation should describe the harm's class and impact without becoming a compressed version of the thing you are containing. Where the content is also unlawful or reportable in your jurisdiction, that path is decided with counsel in advance, and the engagement rules should have said so before the run started.

  • Who should be on the tier-2 access list?
    The engineers who must reproduce and the reviewer who verifies the fix — named, logged, and time-bounded. Prioritisation and management can work entirely from the characterised tier-1 ticket.
  • How do you keep the restricted store from becoming an unverifiable claim in the open ticket?
    Put a content hash, capture timestamp and deployment identifier in tier 1, so anyone can later confirm the restricted artefact matches the finding without the artefact moving.

saying these in an interview costs you the question

  • Pasting actionable harmful content into a company-wide readable ticket because 'it is the evidence'.
  • Withholding the reproduction input entirely and expecting the fix to be verified anyway.
  • Redacting so heavily that nobody notices the finding can no longer be reproduced.
  • No retention or deletion plan for accumulated harmful-output evidence.
  • Deciding the handling policy ad hoc at filing time instead of before the engagement.
  • Assuming a private repository or an attachment is equivalent to access control with a log.

context