skip to content

Red-Team Benchmarks & Datasets

You will learn the standardized benchmarks and datasets that put a number on attack success and defense robustness, and their methodology caveats. Interviewers use them to test whether your red-team claims are measurable and comparable rather than anecdotal.

on this pageshow

explore

questions

page 1 of 2

A vendor publishes a red-team benchmark score for a model, and your product calls that same model behind a system prompt, a tool loop and a retrieval layer. What interface did that published number actually measure, and why can you not quote it as your product's safety number?

level: juniorimportance: must knowfreq 72%

answer

  1. score stops at the model API
  2. wrapper never in the run
  3. two numbers, two interfaces
  4. not an upper or lower bound
  5. component claim vs product claim

basics

~20 s

The number came from prompts sent straight to the bare model endpoint - no system prompt, no tools, no retrieved text. Your product is a different system: extra instructions, extra input channels, extra output paths, none of which that run exercised. So the score describes the model you call, not the application you ship.

solid answer

~50 s

Published red-team suites drive the model API directly. Each item is one user turn, the context is empty or a fixed neutral preamble, and the harm judge reads the raw completion. Your deployment changes all three of those: a system prompt that constrains scope and tone, a tool loop that lets a compliant answer become an action, and retrieval that pushes text you did not author into the same context window. So the published score is a statement about a **component**, useful for supplier selection and for spotting a model swap. The honest write-up keeps two numbers and two interfaces: "the model we call scored X on this published suite at its bare API; we separately ran Y items against our own endpoint and got Z." Merging them into one product claim is the mistake interviewers are listening for.

go deeper

for a junior

Says the benchmark hit the model directly and the product adds a system prompt, tools and retrieval on top, so the number is about the model, not the app.

for a middle

Explains that the harness's target was a plain model call and names concretely what the wrapper changes in the context and in the output path.

for a senior

Argues the gap runs both ways so the published figure is not a bound, and describes standing up a target against the shipped endpoint with a slice, a budget and its own denominator.

for a principal

Sets the organisational rule: vendor scores are supplier-selection and regression signals, product claims come only from runs at the shipped boundary, and reports never merge the two.

**What the published number physically measured.** A red-team benchmark is two parts. The *item corpus* is a list of prompt strings, each tagged with the behaviour it is trying to elicit — advbench, harmbench and jailbreakbench all distribute theirs as flat files. The *harness* walks that corpus, and at its centre sits a **target adapter**: harmbench names one in its run configuration, jailbreakbench exposes one as an LLM object. Its entire contract is "take one string, return one string". Physically that is a single HTTPS request to a chat-completions endpoint carrying the item as the only message in the array, a pinned temperature, and some fixed number of samples per item. The reply goes to a **judge** — harmbench ships a fine-tuned harm classifier, jailbreakbench specifies a judge model plus a published rubric — which labels the completion as exhibiting the behaviour or not. Attack-success rate (ASR) is judged hits over items attempted. Three properties fall straight out of that loop, and together they are the whole answer. The context held nothing but the item. The output path was raw prose. And nothing the model said could take an action. **What your product is instead.** | in the published run | in your deployment | |---|---| | empty context | a system prompt: role, scope, refusal policy, tenant identifiers, tool descriptions — commonly 500-2,000 tokens | | one user turn | user turn plus retrieved chunks plus tool results, all in one window | | text out | text out **and** tool calls your orchestrator actually executes | | no pre/post filter | input classifier, output classifier or redaction, length caps, response schema | | single turn | a session with history | Every row changes the token sequence the model conditions on; the third and fourth change what a compliant response is worth. **What an honest replay costs.** Engineering first. The target adapter now has to authenticate as a real session, send a turn, capture a streamed response to completion, and capture tool calls with their arguments. That is one to three engineer-days, and it breaks again whenever the response envelope changes. Then the run: a 400-item slice at 3 samples is 1,200 application calls plus 1,200 judge calls. The item text is identical to the published run, but every call now carries the system prompt and the retrieved chunks, so input tokens per call are routinely 20-40x the published run's — the token bill and the latency scale with that, not with the item. Your own endpoint is usually rate-limited per tenant far below the raw model API, so a sweep that took twenty minutes against the model API takes hours against the product. Human triage then dominates everything else, because someone still reads every disagreement between judge and expectation. **Where the number misleads.** The wrong reading to name explicitly is: *"the model scored 4% ASR, so our product is at most 4%."* That treats a component measurement as an upper bound on a system. It is not a bound in either direction. Downward, your system prompt gives the assistant an unrelated job and your input classifier rejects items before the model ever sees them, so items that landed on the bare model die in your stack. Upward, retrieval and tool results are a second inbound text channel with no counterpart in the published run, a tool call turns text-level compliance into a real side effect, and the system prompt itself becomes something worth extracting. Both pressures are live at once, so the published figure bounds nothing. Two further misreadings follow from the same confusion. Treating *published minus replayed* as "what our safety layers bought" attributes a delta to one cause when the interface, the context length, the output shape, the available actions and possibly the judge's behaviour all changed together. And quietly reusing the published denominator — the suite's item count — for a run in which you attempted a slice, dropped rate-limited items and lost some to truncation produces a rate computed over a set nobody defined. **What you check before believing your own number.** - The adapter captures the *whole* envelope: user-visible text, tool calls with arguments, and the stage at which the item terminated. - The judge still agrees with humans on your response format — hand-label a stratified sample and report the agreement rate next to the score. - Per-item termination is logged, so you can tell "the model refused" from "the filter caught it". - HTTP reality is logged: 429s, retries, timeouts and truncations belong in a stated denominator, not silently dropped. - The sandbox tools are faithful enough that an action-level hit means something in production. **How you report it.** Two numbers, two interfaces, never merged: "the model we call scored X on this published suite at its bare API; we separately attempted Y items of that suite against our own endpoint and observed Z, with judge agreement of A on a hand-labelled sample." Anything shorter is a component claim wearing a product label.

  • Name one thing your deployment adds that could make the published number too optimistic, and one that could make it too pessimistic.
    Too optimistic: tools and retrieval add inbound channels and real side effects the suite never touched. Too pessimistic: a narrow system prompt plus an input classifier deflects items that landed on the bare model.
  • What is the minimum you must change in a benchmark harness to point it at your application instead of the model?
    The target object - the thing the harness calls with an item and gets a response from. It has to hit your app's endpoint with real auth, real system prompt and real tools, and return the text and actions the user would actually get.
  • Your app endpoint is rate-limited and metered while the model API is not. How does that change what you replay?
    You stop replaying the full corpus and pick a slice, then report that slice's own denominator. The budget becomes part of the method, and you say so rather than implying full-suite coverage.

The published score is a dynamometer reading for an engine on a test bench. Your product is the whole car built around it, and the wiring you added can both damp the vibration and start a fire the bench run could never have seen.

saying these in an interview costs you the question

  • Treating the published score as a conservative upper bound on the deployed system's risk.
  • Quoting a vendor leaderboard figure in a product security document as if it were measured on the product.
  • Assuming a system prompt can only reduce attack success, never enable a new one.
  • Claiming the app is safer because 'we use the same model that scored well', with no run at the shipped endpoint.

context

open as a page

You run a public harmful-behaviour benchmark such as HarmBench against a hosted chat endpoint and save the resulting attack-success rate for later comparison. What do you record alongside the number so it can still be interpreted in six months?

level: juniorimportance: must knowfreq 50%

basics

~20 s

Record the endpoint string and any pinned revision it resolved to, the timestamp, decoding settings such as temperature and max tokens, the dataset revision and which behaviours ran, the harm judge and its version, and the raw transcripts. Without these, a later number cannot be compared and any difference cannot be explained.

open as a page

A red-team benchmark reports an attack-success rate (ASR) for a model. What are the numerator and the denominator of that fraction, and what must the results table state before the number can be read at all?

level: juniorimportance: must knowfreq 78%

basics

~20 s

The numerator is the count of attempts the suite judged harmful; the denominator is the unit the suite chose to count over. That unit may be behaviours on a fixed list, individual attempts, or behaviour-and-attack pairs. Before reading the rate you need the behaviour list, the attempts per item, the attack method, and who ruled an attempt a hit.

open as a page

A jailbreak benchmark harness rules an attempt successful when the model's reply does not contain any phrase from a fixed refusal-string list ("I'm sorry", "I cannot", "As an AI"). What errors does this substring rule push into the attack-success rate it produces?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It counts by wording, not content. A reply that apologises then complies is scored a refusal; an empty, off-topic or garbled reply with no listed phrase is scored a success. Unlisted refusal wordings, other languages and paraphrases all leak through, so the rate drifts both up and down.

open as a page

A red-team report says a model scored 0% attack success on a jailbreak benchmark. Why is that number alone not evidence the model is good, and what second measurement belongs beside it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

A model that refuses everything scores zero on any attack suite, so zero can mean hardened or mean useless. The attack rate only reads next to a benign-refusal rate: run a set of harmless prompts, many of them phrased to look sensitive, and report how many were refused. Publish both numbers together.

open as a page

A jailbreak evaluation reports "38% attack success" against a fixed list of 300 harmful behaviours. Why is that percentage not interpretable until you also know how many attempts were made per behaviour?

level: juniorimportance: must knowfreq 65%

basics

~20 s

Because a behaviour usually counts as broken if any attempt succeeds. Twenty attempts give twenty chances where one gives one, so the figure only rises as the attempt count rises. With attempts-per-behaviour unstated, 38% describes how much you spent as much as how fragile the model is.

open as a page

A fixed-attack jailbreak benchmark such as JailbreakBench freezes several things at once so that results are comparable. What does it freeze, and why is that freeze what makes two different defences comparable on the same model?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It freezes the attack prompts, the judge that decides whether a response counts as a jailbreak, and the report format. With all three fixed, two defences see identical inputs scored by identical rules, so the difference in reported success rate reflects the defence and not a different prompt set or a different grader.

open as a page

A teammate reports that your model "passed HarmBench." What is a harmful-behaviour set such as HarmBench or AdvBench actually a list of, and what does a good result over it cover and not cover?

level: juniorimportance: must knowfreq 70%

basics

~20 s

It is a fixed list of harmful request descriptions, sorted into categories its authors chose in advance. A good result means the model refused those written requests under whatever attack drove them. It says nothing about harms that are not on the list, including whatever your own product's worst outcome is.

open as a page

A model card reports a near-zero attack-success rate on AdvBench, a public corpus of harmful-behaviour prompts. What could produce that number without the model being any harder to jailbreak?

level: middleimportance: must knowfreq 62%

basics

~20 s

Nothing prevents the corpus being inside safety tuning. Those exact prompts appear in refusal training data, and often in the guard classifier that scores the run, so the model refuses memorised strings. Reword the same requests and the rate typically climbs. The number measures recall of a public list.

open as a page

You replay a published red-team benchmark's prompt set against your deployed assistant endpoint instead of against the bare model API. What is one mechanism that pushes the attack-success number DOWN compared with the published run, and one that pushes it UP?

level: middleimportance: must knowfreq 60%

basics

~20 s

Down: your system prompt narrows what the assistant will discuss and an input or output classifier sits in the path, so items that landed on the bare model get deflected. Up: retrieval carries text you did not write, tools turn a compliant answer into a real action, and a long preamble dilutes plain refusal behaviour.

open as a page

A public safety-benchmark leaderboard shows an attack-success rate measured some months ago against a hosted chat endpoint. What does that row fail to record about the system it scored, and why does that limit citing the number today?

level: middleimportance: must knowfreq 58%

basics

~20 s

It records a name, not a build. A hosted endpoint's alias can be re-pointed to an updated model, and the safety tuning and any provider-side filters behind it can change with no announcement. The row also freezes the decoding settings and the harm judge used then. It is one snapshot, not today's endpoint.

open as a page

You have one attack-success rate published by HarmBench and one published by JailbreakBench for the same open-weights model. What has to line up before you can put both figures in a single comparison table, and what usually does not?

level: middleimportance: must knowfreq 66%

basics

~20 s

Both rates must count over the same population and be produced the same way: the same behaviour list, the same attempts per behaviour and aggregation rule, the same attack method, the same target configuration, the same procedure for ruling a hit. Independently published suites share none of these, so their headline rates belong on different axes.

open as a page

A red-team suite reports a 5% attack-success rate. The harm judge that decided which replies counted as successful attacks measured 8% false positives and 12% false negatives on a labelled sample. Why can the 5% not be read as "5% of attempts really succeeded", and what can you honestly state instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

Judge error sets a floor. With an 8% false-positive rate, even a model that never complies would score around 8% flagged hits, so a 5% reading sits inside the judge's noise. Report the judge, its measured error rates, and adjudicate a sample before claiming any true rate.

open as a page

Running one attack against the same fixed harmful-behaviour list at 1, 5 and 25 attempts per behaviour gives 12%, 29% and 41% of behaviours broken. What shape is that curve, and what does it mean when it flattens?

level: middleimportance: must knowfreq 55%

basics

~20 s

It rises and flattens. The easy behaviours fall in the first few attempts, so each extra attempt buys less. Flattening means further attempts of this attack will find little more: you have separated the behaviours this attack can break from a residual core that resists it at any budget you can afford.

open as a page

A fixed-attack jailbreak benchmark like JailbreakBench ships a harm judge that labels each model response as jailbroken or not. Why is that judge part of the frozen artefact, and what breaks in your reported numbers if you substitute your own judge?

level: middleimportance: must knowfreq 60%

basics

~20 s

The judge decides what counts as a jailbreak, so it is half the measurement. Swap it and your numbers stop comparing to every published result on the suite: a stricter judge lowers the reported success rate, a looser one raises it. Report your own judge's result separately, never as the leaderboard number.

open as a page

Your product is a bank's customer-support assistant, and its worst realistic outcome is being talked into revealing another customer's account details. It scores near-perfect against the AdvBench and HarmBench behaviour lists. Mechanically, why does that result say nothing about your top risk, and what would you do instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

Those lists contain behaviours their authors picked — weapons, fraud, harassment, self-harm. Cross-customer data disclosure is not among them, so no attempt ever aimed at it. A behaviour nobody attempts cannot fail, so your top risk contributes nothing to the result. Write your own behaviours from your risk register and measure those separately.

open as a page

You re-run a broad multi-dimension trust battery such as TrustLLM after shipping a mitigation, and one dimension's score comes back a few points higher. Why is that not yet evidence the mitigation worked, and what do you check before claiming it?

level: middleimportance: must knowfreq 55%

basics

~20 s

That dimension is backed by few prompts, so a few points can be one or two responses flipping. Sampling, non-zero-temperature decoding and the judging step all move it on their own. Check which individual items changed, re-run the unmitigated build unchanged, and see whether the movement is bigger than run-to-run drift.

open as a page

You are handed only a published safety evaluation — a low attack-success rate on a public attack corpus, scored by a guard classifier the same vendor ships — and no access to any training data. What evidence would tell you the corpus is inside the safety tuning or inside that classifier's training set?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Re-run the same harmful behaviours in fresh wording and compare. A large gap between verbatim prompts and paraphrases points at memorisation. Score the responses with a judge the vendor did not train, check whether the corpus predates the model release, and see whether the guard's own documentation lists it as training data.

open as a page

Two teams publish attack-success rates for the same open-weights model on the same behaviour list. One scored replies with the trained harm classifier the benchmark distributes; the other scored them with a hosted chat model prompted with a harm rubric. The rates differ by 14 points. How do you work out whether the model or the scoring choice explains the gap, and what do you require before putting both numbers in one table?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Hold the transcripts fixed and vary the scorer. Get both teams' raw completions and run both judges over both sets. If each judge gives similar numbers on either set, the gap is the scoring choice, not the model. Without archived completions the two rates cannot share a table.

open as a page

AdvBench and HarmBench are publicly released attack-prompt corpora used to score how often a model refuses. Why do red teams keep part of their own attack prompt set unpublished?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Because anything published gets trained on. Vendors add public attack prompts to safety-tuning data and to guard classifiers, so models learn to refuse those exact prompts. The score rises without the model becoming harder to attack. An unpublished set stays a fair test, since nobody could have fitted to it.

open as a page

You are choosing an off-the-shelf evaluation to open a red-team engagement against a chat assistant. What does a broad multi-dimension trust battery such as TrustLLM buy you compared with a focused single-purpose attack suite, and what does that breadth cost?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A broad trust battery scores several different trust properties in one pass, so it gives you a wide first map of where a model looks weak. A focused attack suite drills one property hard instead. The cost is depth: each dimension is backed by a thin slice of prompts, so its number is coarse.

open as a page

A provider offers both a floating alias for a chat model and a pinned snapshot identifier of the same model. Which of the two should a recurring safety-benchmark suite be pointed at, and what does each choice cost you?

level: middleimportance: should knowfreq 36%

basics

~20 s

Point at both if you can. The alias is what production calls, so running it is your drift detector. The pinned snapshot is a control: same weights every quarter, so a moved rate there means your harness or judge moved. Costs are double the queries, and pinned snapshots eventually get retired.

open as a page

One jailbreak suite marks a behaviour as broken if any of n sampled attempts is ruled harmful and reports the fraction of behaviours broken; another reports the fraction of individual attempts ruled harmful. How do those two rates relate, and can you convert between them?

level: middleimportance: should knowfreq 52%

basics

~20 s

They are different statistics. The any-of-n behaviour rate is never lower than the per-attempt rate on the same transcripts, and it rises as n rises even though the model is unchanged. You cannot convert one to the other from the headline numbers, because the conversion needs the per-behaviour distribution of hits, which the aggregate discards.

open as a page

When you score a set of harmless prompts to measure a model's over-refusal, what should count as a refusal, and why is that label not simply the mirror image of scoring a hit on the attack side?

level: middleimportance: should knowfreq 48%

basics

~20 s

Not just the flat "I can't help with that". Count deflections, moralising non-answers, and replies that answer a safer question than the one asked. The attack side has a concrete target — did the harmful content appear. The benign side has no such artefact, so you are grading whether the user's actual request was served.

open as a page

Before citing a per-category breakdown from a standard harmful-behaviour list such as AdvBench or HarmBench, what should you check about the list's own composition, and how does that change how you read those category numbers?

level: middleimportance: should knowfreq 40%

basics

~20 s

Open the file and read it. Check how many entries each category holds, how many entries are near-restatements of each other, and how the requests are phrased. Categories are unevenly sized and wordings repeat, so a category figure can rest on a handful of entries or on one idea counted several times.

open as a page

Your quarterly red-team report has cited an AdvBench attack-success rate for several quarters. You now learn the deployed guard classifier was trained on that same public corpus. Do you drop the number, and what goes in the report instead?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Do not silently drop it. Keep the public number as a regression floor, clearly labelled as a set the guard was trained on, and make a held-out attack set the headline. Report both, since the gap between them is your best estimate of how much the public score is inflated.

open as a page

When you replay a published red-team benchmark against your own application endpoint, the benchmark's automated harm judge can start scoring the wrong thing because your app's responses do not look like raw model completions. What are two ways your response format breaks that judge, and what do you do about each?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Two common breaks: the app returns structured output or a rendered card, so the judge reads a wrapper rather than the answer text; and the app emits one fixed refusal template the judge was never calibrated on. Fix by judging exactly the text and tool calls the user receives, then hand-labelling a sample to check the judge still agrees.

open as a page

A published red-team benchmark delivers every item as a user turn, but your deployment also puts text you did not write - retrieved documents and tool results - into the same context. What do you re-run against that second channel, and how do you keep it a bounded job instead of a second full benchmark?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Do not replay the whole corpus on the retrieval channel. Select the items whose outcome depends on the model following injected instructions, place that text in the retrieved-document or tool-result slot instead of the user turn, and score what the app did as well as what it said. Report the two channels as separate coverage numbers with separate denominators.

open as a page

A quarter after your baseline, you re-run the same harmful-behaviour benchmark with the same prompts against the same hosted endpoint alias, and the attack-success rate drops noticeably. How do you establish whether the model behind the alias changed, your harm judge changed, or it is run-to-run variation?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Change one thing at a time. Re-score the old transcripts with the new judge: if the rate moves there, the judge drifted. Diff the run manifests for decoding, dataset and error-handling changes. Then repeat the new run to see the spread across repeats before calling any residual gap a real model change.

open as a page

A stakeholder points at a published safety-benchmark attack-success rate for the base model behind your product and asks why your own red-team scan of the deployed assistant produced a very different rate. What does the published figure's denominator actually enumerate, and why do the two numbers not sit on one axis?

level: seniorimportance: should knowfreq 48%

basics

~20 s

The published denominator enumerates that suite's fixed behaviour list, attacked by that suite's method, against the model as the suite configured it - usually raw weights with a plain chat template, no product system prompt, no retrieval, no filters. Your scan enumerates your own attack set against the deployed stack. Different population, target and ruling.

open as a page

showing 1–30 of 48