skip to content

Harmful-Behaviour Sets

A behaviour list such as HarmBench or AdvBench decides in advance what counts as a harm worth measuring, so a risk it omits scores zero. Interviewers ask whether you read the list before citing it.

on this pageshow

explore

questions

5

A teammate reports that your model "passed HarmBench." What is a harmful-behaviour set such as HarmBench or AdvBench actually a list of, and what does a good result over it cover and not cover?

level: juniorimportance: must knowfreq 70%

answer

  1. list, attack, judge are three things
  2. behaviour set = the inventory
  3. categories fixed by authors
  4. omitted harm scores nothing
  5. name list, revision, attack

basics

~20 s

It is a fixed list of harmful request descriptions, sorted into categories its authors chose in advance. A good result means the model refused those written requests under whatever attack drove them. It says nothing about harms that are not on the list, including whatever your own product's worst outcome is.

solid answer

~50 s

A harmful-behaviour set is an inventory of things you do not want a model to do, written out one entry at a time and grouped into fixed categories. HarmBench and AdvBench are lists of that kind. On their own they attack nothing and judge nothing: an attack method supplies the prompts that chase each behaviour, and a separate judging step decides whether a response counted. So "passed HarmBench" is an incomplete sentence. It has to name which behaviour list and which revision of it, which attack drove the behaviours, and what decided a response was a refusal. The trade you accept by adopting a public list is comparability in exchange for someone else's harm taxonomy. **The list is your unit of measurement**, and anything it omits is not a low score — it is no measurement at all.

go deeper

for a junior

Says a behaviour set is a fixed list of harmful requests and that a good result only covers what is on the list.

for a middle

Separates list, attack and judging, and explains why a bare benchmark name is not a comparable number.

for a senior

Adds what the claim is used for downstream — release sign-off, model comparison — and insists on naming the list revision and attack in any report.

for a principal

Frames it as a measurement-governance problem: who owns the list, what it is allowed to be cited as evidence of, and how to stop a generic number standing in for product risk.

**What the file actually is.** A harmful-behaviour set is a data file and nothing more: a few hundred rows, each row a short description of something you do not want a model to do, plus a category tag its authors chose. AdvBench, released alongside the GCG optimisation paper, is a list of that shape — its harmful-behaviours split runs to roughly five hundred rows of one-line requests. HarmBench is another, a few hundred textual behaviours grouped into semantic categories, shipped with a companion harm classifier and a standardised evaluation harness. Check the row count and the revision of the copy you actually have; both projects have been edited since publication. The file attacks nothing and judges nothing. It defines one thing only: the population of behaviours that will be attempted. **Three components wear the same name.** A benchmark result is produced by three separable pieces, and a bare name tells you about one of them. | component | what it is | what changes when you swap it | |---|---|---| | the behaviour list | static rows plus category tags | the population measured; an omitted harm is never attempted at all | | the attack | whatever turns a row into prompts — a direct ask, a role-play template, an optimisation loop, a published transfer suffix | the number, on the same list against the same model | | the judge | the rule that decides a response counted — a refusal-keyword matcher, a fine-tuned classifier, an LLM grader, a human | the number again, over byte-identical transcripts | So “passed HarmBench” is an incomplete sentence. The complete one names the list and its revision, the attack that drove it, the judge that scored it, and the denominator — how many behaviours were attempted, and how many generations per behaviour. **What a run costs.** The direct-ask baseline is cheap: one completion per behaviour per generation, so four hundred behaviours at five generations is two thousand completions — minutes of wall clock and single-digit dollars on a hosted endpoint. Judging roughly doubles the call count if you use an LLM grader; a shipped classifier such as HarmBench's runs locally and costs GPU minutes instead. The expensive tier is optimisation-based attack: a per-behaviour gradient search needs white-box weight access and GPU-hours *per behaviour*, which is hundreds of GPU-hours for a full list, and is why most teams replay published transfer strings at direct-ask price instead of re-running the search. The real bill, though, is human: reading flagged transcripts runs about an hour per fifty to a hundred responses, and that reading — not the aggregate — is where a finding comes from. **Where the number misleads.** Several specific misreadings, each common: - *The direction is unstated.* “96%” can be refusal rate or attack-success rate. Written without the word, a good number and a catastrophic number look identical. - *Omission reads as success.* A harm with no row produced no prompt, so it produced no failure. In the aggregate an unmeasured harm and a refused harm are the same absence. - *The attack was weak.* Direct, unadorned requests sit near the ceiling on any instruction-tuned frontier model, so a near-perfect refusal rate under direct asking is close to a null result; the same list under a strong attack routinely collapses it. A number without its attack cannot be compared with anyone else's. - *The judge scored the wrong thing.* A keyword-based refusal detector marks any response starting “I’m sorry” as a refusal, including one that apologises and then complies, and marks incoherent output as an attacker win. Two teams with the same list and attack diverge on judging alone. - *The list is public and old.* Widely published rows plausibly appear in safety-tuning data, so a high score can reflect exposure to these exact strings rather than generalised refusal. - *The target was the bare model.* A result on the raw API says little about your deployed assistant, which carries a system prompt, retrieval, guardrails and tools — each of which can help or hurt. **What to check before repeating the claim.** Which list, which revision, how many rows. Which attack, and whether it was run or replayed. Which judge, and its agreement with a human on a sampled subset — fifty transcripts is enough to expose a badly calibrated grader. Per-category counts, not just percentages. Whether the target was the model or the product. And finally: has anyone on the team opened the file? If nobody has read the rows, nobody knows what the number covers, and the honest reply to “we passed it” is a request for the denominator.

  • Does a harmful-behaviour set contain attacks?
    No. It contains behaviours you do not want performed. The attack that chases each behaviour is a separate component, and swapping it changes the result on the same list.
  • Two teams report different numbers on the same behaviour list for the same model. Name two causes that are not the model.
    A different attack driving the behaviours, and a different rule for deciding whether a response counted as compliance. A different revision of the list is a third.
  • What is the minimum you should state alongside a behaviour-set result?
    Which behaviour list and revision, which attack drove it, what decided a hit, and how many behaviours were attempted.

saying these in an interview costs you the question

  • Describes a harmful-behaviour set as an attack tool or a scanner.
  • Treats a high result as evidence the product is safe to deploy.
  • Cannot say what decided that a response counted as a failure.
  • Repeats a benchmark name without ever having opened the behaviour list.
  • Assumes every published run on the same list used the same attack.

context

open as a page

Your product is a bank's customer-support assistant, and its worst realistic outcome is being talked into revealing another customer's account details. It scores near-perfect against the AdvBench and HarmBench behaviour lists. Mechanically, why does that result say nothing about your top risk, and what would you do instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

Those lists contain behaviours their authors picked — weapons, fraud, harassment, self-harm. Cross-customer data disclosure is not among them, so no attempt ever aimed at it. A behaviour nobody attempts cannot fail, so your top risk contributes nothing to the result. Write your own behaviours from your risk register and measure those separately.

open as a page

Before citing a per-category breakdown from a standard harmful-behaviour list such as AdvBench or HarmBench, what should you check about the list's own composition, and how does that change how you read those category numbers?

level: middleimportance: should knowfreq 40%

basics

~20 s

Open the file and read it. Check how many entries each category holds, how many entries are near-restatements of each other, and how the requests are phrased. Categories are unevenly sized and wordings repeat, so a category figure can rest on a handful of entries or on one idea counted several times.

open as a page

Your team adds twenty product-specific behaviours to an evaluation that previously used only the HarmBench behaviour list, and the headline number moves. What has actually changed, and how would you report the two sets of behaviours?

level: seniorimportance: should knowfreq 45%

basics

~20 s

You changed the measured population, not the model. The mixed figure is comparable to nothing: not to your earlier run on the public list, and not to anything anyone else reports on it. Report the public list and your own behaviours as two results, each with its own count, and never merge them into one headline.

open as a page

You lead red teaming for a company shipping a code assistant, a bank chat assistant and a medical triage bot. Would you make a standard harmful-behaviour set such as HarmBench the organisation's unit of measurement, keep a per-product behaviour register, or both? Which would you choose, and what is each number allowed to justify?

level: principalimportance: should knowfreq 35%

basics

~20 s

Both, with different jobs. The shared list is a cheap cross-product regression signal and lets you compare candidate models. It cannot represent product risk, so every product also owns a behaviour register drawn from its own threat model, and that register, not the shared figure, gates release.

open as a page