Red-Team Reporting & Operations
You will learn how to convert raw tool output into a prioritized, reproducible red-team report mapped to ATLAS and the AI RMF. Interviewers weight this heavily because the deliverable, not the exploit, is what makes AI red-teaming useful to an organization.
on this pageshowhide
explore
- Running the Engagement19 questions
- Scope and Authorisation5 questions
- Handling Harmful Output5 questions
- Disclosing to a Provider4 questions
- Planning and Declaring Coverage5 questions
- Writing the Finding21 questions
- Reproducing an Attack5 questions
- Rating Severity6 questions
- Deduplicating Variants5 questions
- Developer Evidence Package5 questions
- Published Structures8 questions
- Adversary Technique Mapping4 questions
- Governance and Risk Frameworks4 questions
- Ongoing Assurance14 questions
- Release Gates4 questions
- Regression on Model Change5 questions
- Measuring the Programme5 questions
questions
62 · 4 sectionsIn an AI red-team report, what is the difference between writing "we exercised this behaviour and observed no failures" and "we never exercised this behaviour", and why must the deliverable keep the two visibly apart?
basics
~20 sThe first is evidence; the second is a gap. A reader who sees neither statement assumes the area passed. So the report must list untested areas explicitly, in the same place as the results, rather than leaving them out. Absence of a finding is only meaningful where something was actually attempted.
During a red-team engagement on a customer application that calls a third-party hosted model endpoint, you reproduce an unsafe response. How do you decide whether the fix owner is the customer or the model provider, and what changes in the report item when the answer is the provider?
basics
~20 sAsk where the behaviour can actually be changed. If the customer's prompt assembly, retrieval data, tool permissions or output filtering could stop it, they own the fix. If only the model's weights or the provider's own safety layer could, the provider owns it, and your report must still give the customer something to apply now.
In an AI red-team engagement, why is pasting the full successful attack transcript — the prompts plus the model's harmful completion — into the circulated report a bad default, when that transcript is exactly what proves the finding?
basics
~20 sBecause the report travels much further than the evidence should. A full transcript is a validated working attack plus the harmful text itself, readable by everyone the document reaches. Keep the raw prompt and completion in a restricted evidence store, and put a description, the harm class and a pointer in the report.
Your client owns a customer-facing chat application built on a third-party hosted model API. Before you red-team that application, whose permission do you need besides the client's, and why?
basics
~20 sYou also need the model provider's. The client owns the application, but the traffic you generate lands on the provider's infrastructure under their acceptable-use terms, which often restrict deliberate attempts to defeat safety measures. Check the provider's testing policy, and get the client's written sign-off naming the accounts and endpoints you may hit.
An AI red-team deliverable states "78% coverage" with no other qualifier. What are at least three different denominators that percentage could be measuring, and what would you write in the report instead?
basics
~20 sCoverage has no meaning without its denominator. It can mean test families executed out of those the harness offers, attack surfaces reached out of the application's entry points, or harm categories exercised out of a taxonomy. Write the fraction with both numerator and denominator named, not a bare percentage.
An automated LLM red-team scan finishes with 480 rows flagged as hits against one chat endpoint. Why is that not 480 report findings, and what does a grouping key do?
basics
~20 sBecause most rows are one weakness repeated: a single attack template retried with small wording changes, and each case re-sent several times. A grouping key is the field combination you collapse rows on, such as technique plus the behaviour elicited, so each report item is one distinct failure a developer fixes once.
You are handing an AI red-team finding to an application team that does not have your scanning harness installed. What must the evidence package contain so they can reproduce the failure themselves?
basics
~20 sShip a self-contained reproduction: the exact request they must send, the endpoint and generation settings it was sent with, the response you observed, and a plain statement of what makes that response a failure. Add how often it happened out of how many attempts. No harness install, no red-team-only dependency.
Your red-team harness logs only the attack prompt and the model's final reply for each attempt against a hosted chat endpoint. Why can a colleague not re-run that attempt from the log, and what should the log capture instead?
basics
~20 sPrompt plus reply says nothing about how the reply was produced. Capture the endpoint and the model identifier the provider returned, the decoding settings you sent, the application's system prompt, every earlier turn, and a timestamp with the provider's request id. Without those, a colleague who reproduces nothing cannot tell drift from a fix.
When collapsing a red-team scan's flagged attempts into report items, what is the difference between grouping by the attack template that produced a hit and grouping by the behaviour the model produced, and what does each key hide?
basics
~20 sGrouping by template says how the model was pushed; grouping by behaviour says what it did. The template key hides that one technique unlocked several unrelated harms. The behaviour key hides that one harm is reachable by many independent routes, so a fix aimed at one route leaves the rest open.
A developer runs the reproduction from your AI red-team finding once, gets a polite refusal, and closes the ticket as unreproducible — but the behaviour is real and intermittent. What should the evidence package have contained to prevent that outcome?
basics
~20 sState up front that the target is sampled, so one attempt proves nothing. Ship the observed count out of total attempts, the sampling settings you used, a script that repeats the request and tallies how many responses met the criterion, and a written verification rule such as: any hit in twenty attempts is still a failure.
Your AI red-team report tags each delivered finding with an identifier from a published adversary technique reference. What does that tag give a reader who was not on the engagement, and what does it not give them?
basics
~20 sIt gives the finding a shared address: a reader can look the technique up, see how other reports used it, and find published mitigations. It does not give severity, likelihood or reproduction detail. Those come from your own evidence, narrative and rating. The identifier is an index entry, not a verdict.
A finding from your AI red-team engagement matches no technique in the published adversary reference you map against. What do you do with it, and how do you write that entry?
basics
~20 sMark it unmapped and say so explicitly. Describe the behaviour in your own words, name the nearest technique and state exactly why it does not fit. Forcing a fit misdescribes the finding, sends the reader to the wrong mitigations, and hides a gap the reference itself may need to cover.
A red-team run against a customer-facing chat assistant produced an attack-success rate of roughly one attempt in five on a jailbreak suite. The governance workstream wants that entered in the AI risk register as a control marked pass or fail. What is lost in that flattening, and how do you record it so the pass or fail is defensible later?
basics
~20 sYou lose the denominator, the mix of attacks that produced the rate, and the fact that the control is probabilistic rather than on or off. Record the rate, the suite it came from, the number of attempts, the target configuration and the date beside the verdict, and state the threshold that turned the rate into pass or fail.
An AI red-team report is read both by the engineers who will fix the issues and by a governance function that maintains an AI risk register. What does each of those readers need from the same finding, and what happens to the attack narrative in the register version?
basics
~20 sEngineers need reproduction: the prompt pattern, the target configuration, what the model returned, and the fix. The register needs a risk statement, which control it is evidence about, how bad it is, and an owner and date. The attack narrative usually shrinks to one line, so keep the detail in a linked annex.
One delivered finding in your AI red-team report came from a chain: untrusted retrieved content steered the assistant into a tool call, and that call moved data out. How do you map it onto a published adversary technique reference without inflating your finding count?
basics
~20 sKeep it as one finding — the reader's unit is the delivered impact, not the step count. Name a primary technique for that impact and list the enabling steps as supporting techniques inside the same entry, in order. Splitting one chain into three findings inflates the count and destroys the causal sequence a defender needs.
An AI red-team programme reports a quarterly "findings" number obtained by summing every hit its automated LLM scanners emitted. Why is that number not the number of findings in the report, and what unit should be counted instead?
basics
~20 sA scanner hit is one prompt attempt its detector marked as failed. Many hits are the same behaviour repeated across prompt variants, probe classes and reruns, and some are detector false positives. The report's unit is a triaged, deduplicated issue with a cause. Count those, and report attempts separately as volume.
Your AI red-team findings were filed against a hosted chat application whose endpoint URL and request contract never change. Which changes underneath that unchanged endpoint can silently invalidate those findings, and why does nothing alert you?
basics
~20 sThree: the provider swaps the model behind the same endpoint alias, the safety classifier in front of it is upgraded, or the product team edits the system prompt. None changes the URL or the contract, so no deploy alert or failing test fires. Your filed pass and fail results simply stop describing the live system.
If an AI red-team programme's headline number is confirmed findings per quarter, which part of its automated testing gets more investment over time, and why is that the wrong part?
basics
~20 sWhatever produces findings cheapest wins the budget. That is usually a high-yield, low-severity probe family against an easy surface. The hard work — a slow multi-turn attack on a tool-using agent — yields few items and looks unproductive. The programme drifts toward whichever instrument is noisiest, not toward where the risk is.
What must you record about an AI red-team run so that repeating it months later is a comparison against the filed result rather than a fresh experiment?
basics
~20 sRecord the whole configuration, not the numbers: the concrete model identifier and any version metadata returned, decoding parameters and seeds, a hash of the system prompt and context, the guardrail versions and thresholds, the exact attack prompts and harness version, the deciding component's version, trials per prompt, and the raw transcripts.
A red-team engagement on a customer-facing chat assistant closes with a list of confirmed findings, and your release gate offers three dispositions: block the release, ship behind a compensating control, or ship with a named acceptance. How do you decide which disposition each finding gets?
basics
~20 sDecide by finding class and consequence, not by a tool's score. Classes whose harm is irreversible or legally exposed block. A finding ships behind a compensating control only if a control already deployed demonstrably stops it. Everything else ships with a named owner accepting it in writing, with a review date.