skip to content

Red-Team Benchmarks & Datasets

You will learn the standardized benchmarks and datasets that put a number on attack success and defense robustness, and their methodology caveats. Interviewers use them to test whether your red-team claims are measurable and comparable rather than anecdotal.

on this pageshow

explore

questions

page 2 of 2

When a judge model scores red-team transcripts to decide which attempts landed, the text it reads contains the attacker's own prompt and the target's reply. Why does that make the judge itself an attack surface, and how would you check whether your published attack-success rate has been distorted by it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

The judge reads attacker-controlled text, so the transcript is untrusted input to it. Payloads can address whatever reads them next and push a scoring decision, and encoded or obfuscated replies fall outside a trained classifier's training data. Hand-adjudicate a random sample of both scored-hit and scored-miss transcripts.

open as a page

A safety-tuned build drops your attack-success rate on the same jailbreak suite from 22% to 6%. What do you measure, and how do you design the comparison, to show that this is real hardening rather than the model simply refusing more of everything?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Re-run the frozen benign prompt set against both builds and compare refusal rates. Real hardening leaves the benign rate flat while the attack rate falls; a blanket refusal shift moves both together. Hold the prompt sets, system prompt, decoding settings and labelling rubric identical, and change only the build.

open as a page

You hold a fixed query budget of 20,000 generations against a metered chat endpoint for one engagement. Do you spend it as 400 behaviours at 50 attempts each, or 2,000 behaviours at 10 attempts each? How do you decide?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Decide from the claim the number must support. Breadth, many behaviours at few attempts, finds unexpected harm categories and gives a stable per-behaviour picture. Depth, fewer behaviours at many attempts, shows what a determined attacker eventually gets. Most engagements buy a wide cheap pass first, then spend the remainder deep where the pass looked weak.

open as a page

Another team publishes "2% attack success" on the same public harmful-behaviour list you used, where you measured 30%, and their write-up never states attempts per behaviour. What can you legitimately conclude, and what do you do to make the two comparable?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Almost nothing about the gap. Their 2% may be one attempt per behaviour against your twenty. Treat it as a lower bound, ask for attempts per behaviour, the stopping rule and what ruled a hit, and re-measure yourself at a declared attempt budget before putting the two numbers in one table.

open as a page

Two candidate input filters were each measured on the same fixed-attack jailbreak benchmark, and one reports a lower attack-success rate. What must you verify before telling the team that filter is the stronger defence to ship?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Check that both ran the same frozen prompt set and the same judge, that the gap is bigger than run-to-run variation, and that neither filter was tuned on those exact prompts. Then look at the per-behaviour breakdown and at what each filter costs in wrongly blocked legitimate traffic.

open as a page

You run a long-published, fixed-attack jailbreak benchmark against a hosted chat endpoint that the vendor has patched many times, and it reports a very low attack-success rate. What are the reasons that number can badly understate your real exposure, and what would you run alongside it?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Its prompts are public and old, so vendors have likely patched exactly those strings and they may sit in training or filter data. A low score then proves those specific attempts fail, not that the model resists new ones. Pair it with freshly generated attacks and attacks aimed at your own application.

open as a page

Your team adds twenty product-specific behaviours to an evaluation that previously used only the HarmBench behaviour list, and the headline number moves. What has actually changed, and how would you report the two sets of behaviours?

level: seniorimportance: should knowfreq 45%

basics

~20 s

You changed the measured population, not the model. The mixed figure is comparable to nothing: not to your earlier run on the public list, and not to anything anyone else reports on it. Report the public list and your own behaviours as two results, each with its own count, and never merge them into one headline.

open as a page

Your team reports a broad multi-dimension trust battery's per-dimension scores each release. After an output-moderation classifier is put in front of the assistant, one dimension improves clearly and another drops. How do you work out what actually happened before you report it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Pull the item-level results for both dimensions and look at what the moderation layer did to each prompt. The usual story: it blocks unsafe completions, so the dimension that rewards refusal rises, while a dimension that penalises unhelpful or over-cautious answers falls on the same behaviour. One change, two scoreboards.

open as a page

Replacing a contaminated public attack corpus means building and maintaining a held-out attack set of your own. What does that cost an organisation over time, and what rules stop the replacement from becoming contaminated too?

level: principalimportance: should knowfreq 36%

basics

~20 s

It costs skilled authoring time, labelled ground truth, and a refresh cadence as the set ages. Keeping it clean is policy: never publish examples, only aggregates; send it only to endpoints under no-training terms; hold a sequestered slice nobody iterates against; and keep the tuning team from seeing it.

open as a page

You lead red-teaming for a product built on a hosted model that already carries published red-team benchmark scores. How do you split a fixed engagement between replaying those published items at your own application boundary and authoring cases only your application can fail - and what do you tell leadership the vendor's published number is still worth?

level: principalimportance: should knowfreq 35%

basics

~20 s

Spend most of the engagement on cases only your application can fail - its system prompt, tools, retrieval and tenant data - and reserve a thin, repeatable slice of published items as a calibration tripwire. Tell leadership the vendor number is a supplier-selection and regression signal about a component, never a claim about the product you ship.

open as a page

You own the rule for how third-party safety-benchmark scores may be cited in your organisation's model-vendor reviews. How do you decide how long such a published number stays citable, and what do you require once it has expired?

level: principalimportance: should knowfreq 28%

basics

~20 s

Tie expiry to events, not just to the calendar: any endpoint or guard change, a dataset or judge revision, or a new attack family invalidates the number. Add a calendar backstop of a quarter or two. Once expired, a published score may inform shortlisting only; anything gating a launch must be reproduced in-house with a recorded manifest.

open as a page

Your organisation publishes a quarterly safety attack-success rate for every model it ships, computed by an automated harm judge, and better judges keep appearing. How do you decide when to change the judge, and what do you owe readers of the historical series?

level: principalimportance: should knowfreq 30%

basics

~20 s

Treat the judge as part of the metric definition. Pin a versioned judge for the published series, archive every raw completion, and when you upgrade, re-score the whole history and publish both series over an overlap window. Never swap silently; state the judge and its measured error rates.

open as a page

Two candidate builds give you an attack-success rate on a red-team suite and a refusal rate on a benign prompt set, and between the builds the two rates move in opposite directions. As the red team, how do you report that pair so nobody is misled, and what claims do you refuse to make from it?

level: principalimportance: should knowfreq 27%

basics

~20 s

Report both rates per build with their prompt-set sizes, labelling rules and run configuration, and state the direction each moved. Refuse to combine them into one index, to declare a winner, and to imply the two rates are in the same units. Add the qualitative slice: which legitimate requests the safer build now declines.

open as a page

Leadership proposes making a score on a public fixed-attack jailbreak leaderboard, such as a JailbreakBench-style suite, the mandatory release gate for every model or prompt update your product ships. What is the case on both sides, and what would you put in the gate instead?

level: principalimportance: should knowfreq 30%

basics

~20 s

For: it is cheap, repeatable and comparable across releases, so a regression is visible. Against: a public frozen suite can be optimised against and says nothing about your own application's attacks. Use it as a non-regression tripwire, and gate release on fresh adversarial runs against your deployed stack.

open as a page

You lead red teaming for a company shipping a code assistant, a bank chat assistant and a medical triage bot. Would you make a standard harmful-behaviour set such as HarmBench the organisation's unit of measurement, keep a per-product behaviour register, or both? Which would you choose, and what is each number allowed to justify?

level: principalimportance: should knowfreq 35%

basics

~20 s

Both, with different jobs. The shared list is a cheap cross-product regression signal and lets you compare candidate models. It cannot represent product risk, so every product also owns a behaviour register drawn from its own threat model, and that register, not the shared figure, gates release.

open as a page

Where does a broad multi-dimension trust battery such as TrustLLM belong in an AI red-team program that also runs adaptive attack tooling and product-specific scenarios, and what claim should its dimension scores never be used to support?

level: principalimportance: should knowfreq 26%

basics

~20 s

Use it as a cheap, repeatable tripwire and as a way to pick where to attack next. Run it early, then on a schedule. Never let a dimension score stand as an assurance claim that the system is safe: it is a fixed, thin, non-adaptive prompt set that knows nothing about your product.

open as a page

Your organisation wants to track one safety benchmark's attack-success rate as a quarterly figure across model upgrades. What has to be frozen for that series to mean anything, and what silently breaks it?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Freeze everything except the model: the pinned item list, the attack method and its budget, attempts per item and the aggregation rule, decoding settings, the target wrapper, and the exact judging procedure. Silent breakers are a refreshed item list, an updated judge, a changed default sampling setting, and the list leaking into training data.

open as a page

You are asked to set the house standard for attempts per behaviour in red-team results that gate a model release. What do you fix, what do you leave to each team, and what does over-fixing cost you?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Fix what makes numbers comparable across releases: the behaviour list, a minimum attempts per behaviour, the stopping rule, the hit rule, and a caption that states all of them. Leave extra deep runs and new attacks free, reported separately. Over-fixing costs discovery: a frozen budget stops anyone probing where the model actually looks weak.

open as a page

showing 31–48 of 48