AI Red Team Tooling & Benchmarks
This is the purpose-built toolchain and the benchmarks used to automate AI/LLM attacks and score attacks and defenses. Interviewers ask directly which of these tools you have actually driven and how you turned findings into a report.
on this pageshowhide
explore
- PyRIT48 questions
- Execution Objects23 questions
- Driving a Run15 questions
- Writing Your Own Pieces10 questions
- garak43 questions
- Scanner Plugins25 questions
- Running a Scan10 questions
- Reading the Report8 questions
- promptfoo42 questions
- Eval Configuration15 questions
- Red-Team Generation18 questions
- Continuous Evaluation9 questions
- Adversarial Robustness Tooling49 questions
- Toolkit Abstractions20 questions
- Running an Assessment20 questions
- Counterfit9 questions
- Automated Jailbreak Generation47 questions
- White-Box Search14 questions
- Attacker-Model Loops14 questions
- Mutation and Crossover9 questions
- Operating a Campaign10 questions
- Guardrail & Moderation Testing46 questions
- The Layers Under Test15 questions
- Measuring a Guard18 questions
- Test-Set Design13 questions
- Red-Team Benchmarks & Datasets48 questions
- Standard Suites14 questions
- What the Number Is19 questions
- Reading a Leaderboard Honestly15 questions
- Red-Team Reporting & Operations62 questions
- Running the Engagement19 questions
- Writing the Finding21 questions
- Published Structures8 questions
- Ongoing Assurance14 questions
- Agent Red-Team Harnesses44 questions
- Benchmark Environments13 questions
- Instrumenting a Live Target18 questions
- Deciding What Counted13 questions
questions
429 · 9 sectionsIn PyRIT, what does a scorer do during an attack run, and what must a custom scorer you write return for the run to be able to use it?
basics
~20 sA PyRIT scorer reads the target's response and decides whether it satisfies the run's objective. A custom scorer must return what the built-in ones return: a verdict, boolean or a normalized number, plus a rationale and the identifier of the piece it scored, so the attack strategy can branch and memory can store it.
You need to run PyRIT against an internal HTTP chat endpoint that PyRIT ships no prompt target for. What does the custom prompt target you write own, and what does the rest of the PyRIT run still do for you?
basics
~20 sYour target owns everything endpoint-specific: authentication, request and response shape, timeouts and retries, and turning the reply into one response PyRIT can store. Everything else is unchanged. The attack strategy still chooses prompts, converters still transform them, scorers still judge them, and the memory store still records each exchange.
In PyRIT, an attack strategy object is constructed once with its prompt target and its scorer, while the objective is supplied each time you execute it. What does that split buy you when you have twenty objectives to test against the same endpoint, and what does each execution get of its own?
basics
~20 sYou build the strategy once with its target, scorer and any converters, then execute it in a loop, once per objective. The wiring is shared across all twenty. Each execution carries its own objective and its own conversation, so the transcript and the verdict stay separate per objective.
In PyRIT, what does a prompt converter do to a prompt before it reaches the target, and why is adding a converter on its own not an attack?
basics
~20 sA PyRIT prompt converter rewrites the outgoing prompt text after the seed prompt is chosen and before it is sent. It changes the surface form, not the request underneath. On its own it is just a transform: it tests whether that form gets through, never whether the model will actually comply.
In PyRIT, what does the conversation memory store record while a run executes, and what do you give up by running with the in-memory store instead of the durable one?
basics
~20 sPyRIT's memory records every prompt sent and every response received, turn by turn, with the scores attached, so a run can be resumed, audited and re-scored later. The in-memory option keeps that only for the process lifetime: when it exits, the transcripts are gone and nothing can be re-read or re-scored.
You want the garak LLM scanner to cover an attack family it does not ship a probe for. Beyond writing the prompt set itself, what else must you supply before a run can report a failure rate?
basics
~20 sPrompts alone produce no number. You must also supply, or point at, the thing that decides whether each reply counts as a hit, and make both discoverable to the scanner so the run loads them. Without that judging step garak collects replies but has nothing to score.
In garak, what job does a detector do, and why is it the detector rather than the probe that decides whether an attempt is recorded as a failure?
basics
~20 sA garak probe only sends prompts. The detector reads each reply and rules whether the attack succeeded. Every pass and failure in the report comes from that ruling, so the number describes the detector's opinion of the outputs, not the model's behaviour directly. Different detectors over the same replies give different numbers.
In garak, what is a generator, and what do you have to own yourself when you wire one to a deployed chat application's authenticated HTTP endpoint instead of to a model you run locally?
basics
~20 sA garak generator is the adapter that sends each probe prompt to the system under test and returns the reply text. Aimed at a deployed app, you own the wiring: the endpoint and method, the auth and tenant headers, the request body shape, and which JSON field the reply is read from.
In the garak LLM scanner, what does a probe supply to a scan, and what does a clean result across the probes you selected say about the probes you did not run?
basics
~20 sA garak probe is the plugin that supplies the prompts sent to the target for one attack family. A clean result covers only the probes you selected; it says nothing about families you never ran. Coverage means probes run out of probes available, and the default selection is not everything.
In a garak run, what does it actually mean when an attempt is marked a hit, and why is that not yet a confirmed vulnerability?
basics
~20 sIt means garak's detector, not a human, judged that one model response matched its rule for failure. The detector may be a string match or a small classifier, so it can be wrong. Open the logged attempt, read the prompt and the full response, and confirm the model really did the thing.
You ran promptfoo's red-team mode before and after editing your application's system prompt, and the failure rate moved. What has to have stayed identical between the two runs for that difference to be evidence about the prompt?
basics
~20 sThe same generated adversarial test cases, the same promptfoo plugins and strategies that produced them, the same grader deciding pass or fail, and the same target model and decoding settings. Only the system prompt may differ. If the cases were generated again, the two runs attacked different inputs and the numbers are not comparable.
In a promptfoo eval configuration, what is the difference between a deterministic assertion (substring, regex, or a small script) and a model-graded llm-rubric assertion, and what does each one cost per test case?
basics
~20 sA deterministic assertion compares the output against a rule you wrote, so it is free, instant, and gives the same verdict on every rerun. An llm-rubric assertion sends the output plus your written criterion to a grading model, costing one extra paid call per case and returning a verdict that can shift between runs.
In promptfoo, you point the HTTP provider at your own chat service, which answers with a JSON envelope. What is the job of the response transform, and what do the assertions grade if you never declare one?
basics
~20 sThe transform picks the assistant's text out of the HTTP response body and hands that string to the graders. Without one, promptfoo grades whatever the raw body serialises to: the whole envelope, metadata included. Assertions then match against wrapper fields, so the report looks plausible but describes the envelope, not the reply.
A promptfoo eval config lists 3 providers, 4 prompt variants and 25 test cases. How many target-model calls does one run make, and what should that number change about your plan?
basics
~20 spromptfoo crosses every prompt with every provider for every test case, so 3 x 4 x 25 is 300 calls per run, before any model-graded check adds its own. The matrix multiplies rather than adds, so each extra provider or prompt buys a whole column you pay for on every run.
In promptfoo's red-team mode, what does adding a plugin to the red-team configuration actually change about the suite that gets generated?
basics
~20 sEach promptfoo red-team plugin stands for one harm class. Adding one makes the generator write adversarial test cases aimed at eliciting that harm, and it supplies the grading criteria used to judge the replies. A plugin you leave out generates nothing, so that harm is never attempted and never shows up in the report.
An adversarial-robustness toolkit such as the Adversarial Robustness Toolbox groups its attack classes under headings like evasion, poisoning, extraction and inference. Your engagement grants only query access to a deployed model and explicitly forbids touching its training data or pipeline. Which headings are off the table, and what must you check before picking any class?
basics
~20 sPoisoning and backdoor classes are off the table: they need write access to training data or the training pipeline plus a retrain, which you were not granted. Evasion classes fit query access. Extraction and inference classes fit technically but produce a surrogate copy or membership claims, so confirm they are in scope and contractually allowed first.
In an adversarial-robustness toolkit such as the Adversarial Robustness Toolbox, defences ship as objects of several kinds: preprocessor, postprocessor, detector, trainer and transformer. Where does each kind sit relative to a call to the wrapped model, and which of them change the model weights?
basics
~20 sA preprocessor defence transforms inputs before the model sees them. A postprocessor alters what comes back, usually the scores. A detector flags a sample as adversarial instead of classifying it. A trainer defence retrains the weights. A transformer hands you back a new, modified model object. Only trainer and transformer touch weights.
An adversarial-attack library reports two numbers after an evaluation run: accuracy on the adversarial examples, and a mean perturbation size. Which rows does each of those two numbers average over?
basics
~20 sAdversarial accuracy is over every example you evaluated, flipped or not. The mean perturbation size is over only the examples the attack actually flipped; rows it failed on are dropped, not counted as huge perturbations. Two different denominators, so the two numbers are not about the same population.
In an adversarial-ML evaluation stack, what is the difference between an attack library (such as the Adversarial Robustness Toolbox or Foolbox) and a command-line harness that drives one (such as Counterfit)? Which of the two decides what attacks are available to you and how far you can tune them?
basics
~20 sThe attack library implements the attacks and holds the parameters; the harness only drives it, doing target setup, batching, logging and output. So the library decides which attacks exist and how tunable they are. The harness decides how conveniently you run them, and may expose fewer attacks or fewer parameters.
In Counterfit, an attack reaches a model only through a scan target you write. What must that scan target supply so a query-only attack can run against a hosted inference endpoint, and what does it deliberately not hand the attack?
basics
~20 sThe scan target wraps your endpoint as a callable: given a batch of samples it returns the model's per-class scores, or a label. You also declare the input shape and data type, the list of output classes, and a few seed samples the attack will perturb. It hands over queries only, never gradients or weights.
In an automated jailbreak search where an attacker model proposes a prompt, a target model answers it, and a separate scoring model rates that answer, how many inference calls does one refinement pass cost, and why does that arithmetic decide the turn cap you set?
basics
~20 sOne pass costs three calls: the attacker writes a candidate, the target answers it, and the scoring model rates the answer. A cap of twenty turns is therefore sixty calls per thread, not twenty. Multiply by restarts and seed goals before you start, because the target leg is usually the metered one.
In an automated jailbreak loop where an attacker model rewrites prompts, a target model answers, and a separate scoring model labels each answer a success or a refusal, what happens when the scoring model wrongly labels a refusal as a success?
basics
~20 sThe loop treats that thread as solved and stops rewriting, so the search ends early. The mislabelled refusal is written into the finding list as a discovered attack, and a human triager has to read the transcript to throw it out. Your reported success count is inflated by every such label.
You may reach the language model you are contracted to attack only through a metered chat API that returns text — no weights, no logprobs. Which classes of automated jailbreak search can still run against it, and which class is ruled out before you start?
basics
~20 sAny search that only needs to send text and read replies still runs: template and mutation pools, and attacker-model loops that rewrite a prompt after each refusal. Gradient-guided suffix optimisation is ruled out — it needs backward passes through the model's weights. Without weights you can only optimise against a local surrogate and hope it transfers.
An automated jailbreak search has produced three thousand prompts its judge marked successful, but they are heavy rewrites of two underlying shapes. Why is the raw count of successful prompts a poor signal for whether to keep spending the remaining queries?
basics
~20 sThe raw count measures how often the search repeated itself, not how much it found. Thousands of near-identical rewrites are one weakness discovered many times. The signal that matters is distinct successful templates: when that stops rising while queries keep being spent, the remaining allowance is buying duplicates.
In an automated jailbreak tool that mutates prompt templates, what is the seed corpus, why are templates that already elicited a violation the usual seeds, and what access to the model under test does this approach need?
basics
~20 sThe seed corpus is the set of prompt templates the tool starts from, usually ones that already made the model comply. Mutation edits them by rewording, filling slots differently, or splicing two together to make many near variants. It needs only query access: send a prompt, read the reply. No weights, no gradients.
A guardrail evaluation runs 500 prompts, every one of them an adversarial attack, and 12% of them reach the model unblocked. A stakeholder reads that as "12% of our traffic is getting through". Why is that reading wrong?
basics
~20 sThe 12% is conditional on the prompt already being an attack, because the corpus is 100% attacks. The denominator is attack prompts, not requests. Real traffic is almost entirely benign, so the share of all requests that are successful attacks is 12% multiplied by how rare attacks actually are.
A teammate reports "our input-moderation guard has a 3% bypass rate" from a red-team harness run. Before you repeat that number in a report, what must the number state about what was counted?
basics
~20 sIt must say what counted as a bypass in the numerator: the guard failed to flag, or the model then actually complied - two different numbers. It must say whether the denominator is attempts, distinct attacks, or attempts per attack. And it must say how many retries each attack got, because retries move the rate without the guard changing.
A guardrail test corpus for an input classifier contains only attack prompts. Why does a classifier that blocks every single input score perfectly on that corpus, and what must be added before the score means anything?
basics
~20 sEvery item in the corpus is an attack, so blocking everything means catching everything: the measured block rate hits 100% and nothing penalises refusals. The corpus can only observe one error type. Add benign prompts the product really sees, and report how many of those were wrongly blocked alongside the attack numbers.
Your red-team report says a content-moderation guard let 3% of your attack prompts through. Which configuration facts have to sit next to that number before anyone else can reproduce or compare it?
basics
~20 sThe block threshold you ran at, which categories were set to block, the guard's identity and configuration snapshot, and the date. A bypass percentage is only true at one operating point, so without those settings written down nobody can rerun your test or compare a later number to yours.
A hosted content-moderation service (such as Azure AI Content Safety or OpenAI's moderation endpoint) returns every category below your threshold for a prompt your team considers harmful. As a red teamer, what does that clean verdict actually license you to conclude, and what does it not?
basics
~20 sOnly that nothing in the service's fixed category set scored above the threshold you applied. The vendor's policy and training data are not published, so a clean verdict is not evidence of safety: the harm may fall outside the categories on offer entirely. Record it as unmeasured, not as passed.
A vendor publishes a red-team benchmark score for a model, and your product calls that same model behind a system prompt, a tool loop and a retrieval layer. What interface did that published number actually measure, and why can you not quote it as your product's safety number?
basics
~20 sThe number came from prompts sent straight to the bare model endpoint - no system prompt, no tools, no retrieved text. Your product is a different system: extra instructions, extra input channels, extra output paths, none of which that run exercised. So the score describes the model you call, not the application you ship.
You run a public harmful-behaviour benchmark such as HarmBench against a hosted chat endpoint and save the resulting attack-success rate for later comparison. What do you record alongside the number so it can still be interpreted in six months?
basics
~20 sRecord the endpoint string and any pinned revision it resolved to, the timestamp, decoding settings such as temperature and max tokens, the dataset revision and which behaviours ran, the harm judge and its version, and the raw transcripts. Without these, a later number cannot be compared and any difference cannot be explained.
A red-team benchmark reports an attack-success rate (ASR) for a model. What are the numerator and the denominator of that fraction, and what must the results table state before the number can be read at all?
basics
~20 sThe numerator is the count of attempts the suite judged harmful; the denominator is the unit the suite chose to count over. That unit may be behaviours on a fixed list, individual attempts, or behaviour-and-attack pairs. Before reading the rate you need the behaviour list, the attempts per item, the attack method, and who ruled an attempt a hit.
A jailbreak benchmark harness rules an attempt successful when the model's reply does not contain any phrase from a fixed refusal-string list ("I'm sorry", "I cannot", "As an AI"). What errors does this substring rule push into the attack-success rate it produces?
basics
~20 sIt counts by wording, not content. A reply that apologises then complies is scored a refusal; an empty, off-topic or garbled reply with no listed phrase is scored a success. Unlisted refusal wordings, other languages and paraphrases all leak through, so the rate drifts both up and down.
A red-team report says a model scored 0% attack success on a jailbreak benchmark. Why is that number alone not evidence the model is good, and what second measurement belongs beside it?
basics
~20 sA model that refuses everything scores zero on any attack suite, so zero can mean hardened or mean useless. The attack rate only reads next to a benign-refusal rate: run a set of harmless prompts, many of them phrased to look sensitive, and report how many were refused. Publish both numbers together.
In an AI red-team report, what is the difference between writing "we exercised this behaviour and observed no failures" and "we never exercised this behaviour", and why must the deliverable keep the two visibly apart?
basics
~20 sThe first is evidence; the second is a gap. A reader who sees neither statement assumes the area passed. So the report must list untested areas explicitly, in the same place as the results, rather than leaving them out. Absence of a finding is only meaningful where something was actually attempted.
During a red-team engagement on a customer application that calls a third-party hosted model endpoint, you reproduce an unsafe response. How do you decide whether the fix owner is the customer or the model provider, and what changes in the report item when the answer is the provider?
basics
~20 sAsk where the behaviour can actually be changed. If the customer's prompt assembly, retrieval data, tool permissions or output filtering could stop it, they own the fix. If only the model's weights or the provider's own safety layer could, the provider owns it, and your report must still give the customer something to apply now.
In an AI red-team engagement, why is pasting the full successful attack transcript — the prompts plus the model's harmful completion — into the circulated report a bad default, when that transcript is exactly what proves the finding?
basics
~20 sBecause the report travels much further than the evidence should. A full transcript is a validated working attack plus the harmful text itself, readable by everyone the document reaches. Keep the raw prompt and completion in a restricted evidence store, and put a description, the harm class and a pointer in the report.
Your client owns a customer-facing chat application built on a third-party hosted model API. Before you red-team that application, whose permission do you need besides the client's, and why?
basics
~20 sYou also need the model provider's. The client owns the application, but the traffic you generate lands on the provider's infrastructure under their acceptable-use terms, which often restrict deliberate attempts to defeat safety measures. Check the provider's testing policy, and get the client's written sign-off naming the accounts and endpoints you may hit.
An automated LLM red-team scan finishes with 480 rows flagged as hits against one chat endpoint. Why is that not 480 report findings, and what does a grouping key do?
basics
~20 sBecause most rows are one weakness repeated: a single attack template retried with small wording changes, and each case re-sent several times. A grouping key is the field combination you collapse rows on, such as technique plus the behaviour elicited, so each report item is one distinct failure a developer fixes once.
In a packaged agent tool world such as AgentDojo — a benchmark that ships mock tools plus tasks — where does the attacker's text actually enter the episode, and why is that different from typing it into the user's prompt?
basics
~20 sIt enters through tool output. The mock tools return data — an email body, a document, a transaction note — and the attacker text sits inside that returned content. The user's request stays benign, so the episode tests whether the agent obeys instructions it read while working, not instructions its user typed.
An agent red-team harness reports a single number: the share of injected runs in which the agent carried out the attacker's instruction. A proposed defence drives that number to nearly zero by making the agent refuse almost every request. Why does it score so well, and what else must the harness measure?
basics
~20 sAttack success rate only counts attacker-instruction executions. An agent that refuses everything never executes one, so it scores near zero: perfectly secure and completely useless. The harness must also run each task without any injection and score whether the agent finished the user's real work, then report both numbers together.
When red-teaming an agent, why do teams stand up their own tool server for the agent to call instead of pointing it at the vendor's live tool server?
basics
~20 sBecause a server you host is repeatable. You can change a tool's description or schema between runs and replay the same attempt to see what moved. A live vendor server gives one shot: it changes under you, its calls touch real data, and a result there cannot be reproduced or varied.
Before red-teaming a deployed assistant that can send email, file tickets and write CRM rows, what do you provision so that a successful attack does not touch real people or real records?
basics
~20 sProvision throwaway identities the agent acts as: a scratch mailbox, a test tenant or sandbox project, and disposable CRM records. Point outbound tools at dry-run or test endpoints. Mark everything with a unique per-run tag so you can find it later, and agree who owns cleanup.
In an agent red-team harness, why is the adversarial instruction placed in the result a tool returns rather than in the user's message, and what does a hit at each of those two boundaries actually prove?
basics
~20 sPut it in the tool result, not the user message. The agent is already trained to be wary of user requests, but treats tool output as data it fetched. The harness intercepts the call and returns attacker-controlled content in its place. A hit there shows the agent obeys untrusted retrieved content.