skip to content

Running the Engagement

An AI test is agreed before it is run: what may be hit, whose terms govern it, and what happens to the harmful text you deliberately produce. Interviewers ask because an unauthorised run ends the job.

on this pageshow

explore

questions

19

In an AI red-team report, what is the difference between writing "we exercised this behaviour and observed no failures" and "we never exercised this behaviour", and why must the deliverable keep the two visibly apart?

level: juniorimportance: must knowfreq 65%

answer

  1. three states: failed, clean, untested
  2. silence reads as passed
  3. remainder section, referenced from summary
  4. not exercised vs out of scope
  5. clean rows need their basis

basics

~20 s

The first is evidence; the second is a gap. A reader who sees neither statement assumes the area passed. So the report must list untested areas explicitly, in the same place as the results, rather than leaving them out. Absence of a finding is only meaningful where something was actually attempted.

solid answer

~50 s

Two sentences that look alike on a slide mean opposite things. **"Exercised, no failures observed"** is bounded evidence: something was attempted, some scoring step decided it did not succeed, and a reader can ask how many attempts and by whose judgement. **"Never exercised"** is an admission of a gap, and it carries no evidence at all. Readers default to charity. An area a report does not mention is assumed to have been looked at and found fine — nobody reads silence as a hole. That is why a deliverable needs a remainder section in the body, referenced from the summary, naming the attack categories, entry points and harm classes that were planned but never run, and why (hours, authorisation, environment, a target that was unavailable). The cost is honest: a stated remainder makes the headline look weaker than a silent one. That is the trade you are being paid to make.

go deeper

for a junior

Says untested is not the same as clean, and that the report should list what was not tested.

for a middle

Adds the third state (out of scope, agreed at kickoff) and attaches a basis — attempt counts, entry point — to each clean result.

for a senior

Structures the deliverable so the summary cannot outrun the evidence table, and knows which untested areas came from a decision versus an operational failure.

for a principal

Sets the house template and the rule that no assurance sentence may reference an area without a row, and negotiates with clients who want the remainder removed.

“Exercised and clean” and “never exercised” are two different amounts of evidence — one bounded, one zero — and once either one becomes white space on a page, a reader cannot tell them apart. Keeping them apart is a structural problem, not a wording one. **The four states** | State | What happened | What the evidence supports | |---|---|---| | Exercised, failed | attempts ran; a scorer called some of them hits | a defect exists here, sized by the attempts behind it | | Exercised, clean | attempts ran; no attempt scored a hit | nothing was found *at this depth, on this surface, by this scorer* | | Not exercised | nobody ran it — hours, environment, credential, provisioning | unknown | | Out of scope | a named person agreed at kickoff it would not be run | unknown, and someone accepted that | Most report templates only carry the first row well. The other three collapse into absence, and absence is read as the second one. **Why the tooling pushes you towards silence** Scanners report what they ran, never what they did not. `garak` writes, per run, a `*.report.jsonl` holding one record per attempt plus a `*.hitlog.jsonl` holding only the attempts a detector scored as a hit; a probe module you never named in `garak --probes` contributes no records at all, and a behaviour that no shipped probe module targets contributes none either. `promptfoo`'s red-team summary is assembled from the plugins named in `redteam.plugins` — anything absent from that list is absent from the table. So in the raw artefacts, “we drove it and it held” is a set of rows scoring zero, while “the idea never entered the config” is no rows; by the time a human pastes a summary into a deck, both are blank. The written plan is the only place the distinction survives, and only if somebody wrote it down before the run. **What a clean row is actually worth** Zero hits is a sampling statement, and its value comes from two numbers the row rarely carries. Depth: `garak`'s `-g` / `--generations` sets how many outputs are drawn per prompt, and a probe left at the tool's default is a handful of samples per prompt, not a campaign. Judgement: `garak` pairs each probe with detectors, and a detector that never fires on the phrasing your target happens to use produces a clean row meaning “nothing matched”, which is not “nothing happened”. So an exercised/clean row needs its basis attached — entry point driven, attempts made, and what decided a response did not count. **What it costs** Producing the ledger is the cheapest line in the engagement: enumerate the planned rows at kickoff with the plan open (about an hour), keep the state current as the run proceeds (minutes a day), reconcile it against the harness output at the end (about an hour). Under a day across a two-week job, and most of it is work you need anyway to know when you are finished. The expensive part is not editorial. A summary with a stated remainder always looks worse than a silent one produced by the identical run, and whoever commissioned the engagement may push back on exactly that. That is the trade this leaf exists to teach: you pay in headline strength to buy a claim that survives being questioned. **Where the number misleads** Two readings to head off. First, an unmentioned area reads as passed — readers default to charity, and nobody interprets silence as a hole. Second, absence quietly inflates every per-category statistic in the document: “3 findings across 9 categories” implies nine categories were driven, and if six were never touched, the honest statement is three findings across three driven categories — the same numerator against a denominator three times smaller, which is a far worse result wearing a better number. **What I check before sending** - Does every planned area appear exactly once, in exactly one state? - Does each exercised/clean row carry attempts, entry point and scorer? - Does any assurance sentence reference an area that has no row behind it? - Could a reader who reads only the summary page name the untested areas? If not, the summary misleads even though every sentence in it is true. - Is “not exercised” separated from “out of scope”? The second was signed by someone at kickoff; the first is an operational loss with no owner yet, and it is the one that should generate a follow-up.

  • Why separate "out of scope" from "not exercised" when both mean no evidence?
    Out of scope was a decision someone accepted and signed at kickoff; not exercised is usually an operational loss — hours, a missing environment, a credential. The first has an owner, the second needs one.
  • A client asks you to delete the untested-areas list because it 'undermines the report'. What do you do?
    Keep it and offer to move it: retitle it as agreed limitations, put the reasons alongside, and note which items a follow-up would cover. Removing it converts an honest report into an assurance claim you cannot support.

saying these in an interview costs you the question

  • Treating an unmentioned area as implicitly out of scope
  • An executive summary that says 'no issues found' with no bound on what was looked at
  • Listing untested areas only in an appendix nobody reading the summary reaches
  • Presenting clean results with no attempt counts or entry point

context

open as a page

During a red-team engagement on a customer application that calls a third-party hosted model endpoint, you reproduce an unsafe response. How do you decide whether the fix owner is the customer or the model provider, and what changes in the report item when the answer is the provider?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Ask where the behaviour can actually be changed. If the customer's prompt assembly, retrieval data, tool permissions or output filtering could stop it, they own the fix. If only the model's weights or the provider's own safety layer could, the provider owns it, and your report must still give the customer something to apply now.

open as a page

In an AI red-team engagement, why is pasting the full successful attack transcript — the prompts plus the model's harmful completion — into the circulated report a bad default, when that transcript is exactly what proves the finding?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Because the report travels much further than the evidence should. A full transcript is a validated working attack plus the harmful text itself, readable by everyone the document reaches. Keep the raw prompt and completion in a restricted evidence store, and put a description, the harm class and a pointer in the report.

open as a page

Your client owns a customer-facing chat application built on a third-party hosted model API. Before you red-team that application, whose permission do you need besides the client's, and why?

level: juniorimportance: must knowfreq 72%

basics

~20 s

You also need the model provider's. The client owns the application, but the traffic you generate lands on the provider's infrastructure under their acceptable-use terms, which often restrict deliberate attempts to defeat safety measures. Check the provider's testing policy, and get the client's written sign-off naming the accounts and endpoints you may hit.

open as a page

An AI red-team deliverable states "78% coverage" with no other qualifier. What are at least three different denominators that percentage could be measuring, and what would you write in the report instead?

level: middleimportance: must knowfreq 55%

basics

~20 s

Coverage has no meaning without its denominator. It can mean test families executed out of those the harness offers, attack surfaces reached out of the application's entry points, or harm categories exercised out of a taxonomy. Write the fraction with both numerator and denominator named, not a bare percentage.

open as a page

You must keep a harmful model completion out of an AI red-team report, but the reader still has to believe the attack worked. What do you put in the report in place of the completion, and what does the reader lose by accepting it?

level: middleimportance: must knowfreq 52%

basics

~20 s

Replace it with an attested description: what the output contained at the granularity of the harm claim, its length and structure, who or what judged it a success, how many attempts succeeded, the target's version and configuration, and a hash plus location of the sealed artefact. The reader loses independent verification and must trust your judgement.

open as a page

You are asked to red-team a staging copy of a client's LLM application instead of production. What differences between the copy and production would stop your results transferring, and what do you record about the environment you tested?

level: middleimportance: must knowfreq 62%

basics

~20 s

Staging results transfer only as far as the configuration matches. Check the system prompt, the model identifier and routing, the retrieval corpus, tool permissions, and whether input and output filters are enabled the same way. Record all of it with the test dates, and say in the report that findings apply to that configuration.

open as a page

A client wants the executive summary of an AI red-team report to read "no critical issues found", but the engagement's hours only allowed two of the eight planned attack categories to be exercised. How do you word the conclusion and the remainder section?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Keep the true finding and bound it. Say no critical issues were found in the categories actually exercised, name them, and list the six that were not, with the reason. A summary saying only "no critical issues found" lets the reader treat untested areas as cleared, which is a claim the evidence cannot carry.

open as a page

The provider of a hosted model closes your submission as expected model behaviour rather than a vulnerability, but your client's application still exhibits the effect. What do you do with that item in the client's report?

level: seniorimportance: must knowfreq 52%

basics

~20 s

It stays in the report, re-framed. The item is no longer waiting on a fix; it is a standing property of the platform your client chose. Rate it against the client's application, describe a compensating control the client can apply themselves, and record the residual risk that is left once that control is in place.

open as a page

You are filing a model-behaviour report with the provider of a hosted chat endpoint your client builds on. What evidence does that submission need that a report of a deterministic software bug would not, and why?

level: middleimportance: should knowfreq 44%

basics

~20 s

The behaviour is probabilistic, so one transcript is not a reproduction. Give how many attempts you made and how many succeeded, the exact endpoint and model identifier, the decoding settings and any system prompt, and the date and time window. Without a rate and those conditions, the reader cannot separate your result from noise.

open as a page

An automated red-team run against a hosted assistant finishes, but roughly 30% of its attempts ended in transport errors, rate-limit rejections or empty responses rather than reaching a scored verdict. How should the coverage statement in the report treat those attempts?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Unscored is not passed. Attempts that ended in transport errors, rate-limit rejections or empty replies never reached a verdict, so they belong in neither the hit nor the clean column. Report attempted, scored and hit as three separate numbers, and re-run the lost attempts before drawing conclusions.

open as a page

Mid-engagement, an AI red-team target emits output that falls into a class your organisation must not retain at all — not merely restrict. What do you do in the next few minutes, and how does the finding still reach the report?

level: seniorimportance: should knowfreq 33%

basics

~20 s

Stop that line of testing, do not copy or forward the output, and treat the tool's own log as holding it too. Trigger the pre-agreed escalation to your named legal and trust-and-safety contact, isolate and destroy the artefacts under a witnessed record, and carry the finding as metadata plus a two-person attestation.

open as a page

An automated LLM red-team scanner writes every request and response, including successful harmful completions, to log files in its working directory. Where do those files realistically end up during an engagement, and what do you change before the first run?

level: seniorimportance: should knowfreq 44%

basics

~20 s

They spread: CI job artefacts and console output, the runner's disk after the job, ticket attachments, screenshots in chat, cloud-synced home folders, laptop backups. Before the first run, move the working directory onto encrypted, unsynced storage in a per-engagement folder, stop artefact upload or restrict and expire it, and schedule destruction.

open as a page

A client asks you to red-team an assistant whose tools can send email, file tickets and call a partner company's API on a user's behalf. How do you scope what the assistant may actually do during the test, and what goes in writing?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The assistant's tools reach systems your client does not own, so a successful injection could send real mail or write to a partner. Agree in writing which tools stay live, which are stubbed, which recipients and accounts are test-only, and who you call if an action escapes into a real third-party system.

open as a page

You plan to run an automated adversarial-prompt campaign of tens of thousands of requests against a client's production LLM endpoint, billed to a metered model-provider account. Beyond permission to test, what do you agree before you start?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Agree whose provider account and key pay for it, a spend cap, a request-rate ceiling and a stop condition. Warn that sustained adversarial traffic can trip the provider's abuse detection and throttle or suspend the account the live product uses, so ask for a separate key and a named contact who can unblock it.

open as a page

How do you keep an AI red-team report from being reused months later as a standing safety certificate for a system that has since changed — a swapped model, an edited system prompt, a new tool integration — and what goes into the deliverable to make that explicit?

level: principalimportance: should knowfreq 33%

basics

~20 s

Bound the result to the exact system you tested. Record the configuration you exercised, state that the conclusions hold only for it, and name the changes that void them: a model swap, a system-prompt edit, a new tool or data source. Add a validity window and explicit re-test triggers.

open as a page

Your team wants to publish a writeup of a weakness you found in a third-party hosted model during a client engagement. Classic vulnerability disclosure publishes when a fix ships or an agreed window expires, and here there may be neither a patch you can point at nor a party who owes you a date. How do you set the publication policy?

level: principalimportance: should knowfreq 28%

basics

~20 s

Decide the trigger yourself, in writing, before you file. With no patched version to point at, pick conditions you control: an agreed period since filing, the client's consent as the report's owner, and content that describes the class of weakness and the mitigation without carrying a working prompt anyone can paste.

open as a page

Continuous automated LLM red-team suites make successful harmful outputs pile up as an archive that doubles as the regression set proving fixes hold. How do you decide what is kept, for how long, and who may open it?

level: principalimportance: should knowfreq 28%

basics

~20 s

Split the artefact. Attack inputs are what regression testing needs, so keep those under access control; harmful completions are usually re-derivable by re-running and re-scoring, so keep only a verdict and a hash. Set per-class retention with default expiry, two-person access with logging, and a named owner outside the red team.

open as a page

The model behind a client's application is a hosted third-party service the client cannot pin or freeze. How do you write scope and findings validity so the report still means something three months later?

level: principalimportance: should knowfreq 38%

basics

~20 s

Record the exact model identifier, endpoint and dates tested, and state that findings describe that snapshot. Write a validity window and named re-test triggers into the scope: a provider model update, a system-prompt change, a new tool or corpus. Push the client toward durable application-side controls rather than one-off prompt patches.

open as a page