skip to content

Before citing a per-category breakdown from a standard harmful-behaviour list such as AdvBench or HarmBench, what should you check about the list's own composition, and how does that change how you read those category numbers?

level: middleimportance: should knowfreq 40%

answer

  1. read the file, not the name
  2. counts beside percentages
  3. near-duplicate entries reweight
  4. blunt single-turn phrasing
  5. same label, different taxonomy

basics

~20 s

Open the file and read it. Check how many entries each category holds, how many entries are near-restatements of each other, and how the requests are phrased. Categories are unevenly sized and wordings repeat, so a category figure can rest on a handful of entries or on one idea counted several times.

solid answer

~50 s

Three properties of the file decide how much a category number means. **Size.** Categories are rarely balanced. A category with a small number of entries flips on one or two responses, so its percentage is noise dressed as a measurement. Always print counts next to percentages. **Redundancy.** It is a well-known criticism of AdvBench in particular that many entries restate the same behaviour in slightly different words. Where that holds, a theme is effectively counted several times, which inflates both its weight in the aggregate and the apparent consistency of the result. **Phrasing.** If the entries are blunt single-turn requests, you are measuring refusal to a direct ask, not resistance in a realistic conversation. The reading that survives all three: a category number is a statement about those specific strings, not about the harm class its label names.

go deeper

for a junior

Knows to open the list and see how many entries each category has before quoting a percentage.

for a middle

Adds near-duplicate entries and blunt single-turn phrasing as reasons a category figure overstates what was measured.

for a senior

Deduplicates or clusters before aggregating, reports counts with percentages, and treats small categories as pointers to read responses rather than as metrics.

for a principal

Requires that any list adopted as an org measurement come with a written note on its composition and what its category labels are allowed to mean.

## Why this is the whole job This leaf's charter says the **behaviour list is your unit of measurement**, and no measurement can be interpreted without inspecting the instrument. That means opening the file and reading rows — which most people quoting the benchmark have never done. Four properties of the file decide how much any category number is worth. ## Four properties of the file **1. Entries per category.** Categories in published lists are rarely balanced; some hold dozens of rows and some hold a handful. The arithmetic is unforgiving: in a category of six rows, a single response is worth 16.7 percentage points, so a figure that moves from 100% to 83% between runs represents one response changing. **Print counts beside every percentage.** “5/6” invites the right question; “83%” invites a ticket. **2. Near-duplicate rows.** It is a well-known criticism of AdvBench that a substantial share of its rows restate the same behaviour with different wording. Where that holds, the theme is counted several times: its share of the aggregate is inflated, and — more insidiously — the result looks more *consistent* than the evidence warrants, because a model that fails one phrasing fails its near-copies too, producing a tight, confident-looking cluster of failures that is really one observation. **Cluster before you aggregate.** - Crude normalised string similarity finds the obvious families in seconds; - sentence embeddings with a similarity threshold find the paraphrases. Report cluster counts next to row counts. **3. Phrasing and modality.** Are the rows blunt single-turn imperatives? Text-only? Aimed at a bare model with no system prompt? Each restriction bounds the claim. A list of direct asks measures the outermost refusal layer and nothing beneath it, which is why direct-ask results sit near the ceiling on tuned models and carry almost no information. **4. The harm definitions themselves.** Category labels encode the authors' policy and jurisdiction. **Dual-use** material — clinical dosing, offensive-security technique, legal or financial advice — may be harmful under their notion of an assistant and entirely legitimate under yours, or the reverse. Two lists whose categories share a name are two taxonomies written by two sets of authors; only the rows say what is inside. ## What the inspection costs - Sorting by category and printing counts is minutes of scripting. - Reading a sample of twenty rows per category is roughly an hour for a mid-sized list, and it is the highest-return hour in the whole adoption. - Clustering for near-duplicates is half a day the first time, then a script you keep. - Writing the one-page **composition note** — what each label means here, how many distinct ideas each category holds, what modality the rows assume — is another half day, and it is the artefact that keeps the next reader from re-learning it. Set against a six-figure model-evaluation programme, this is free. ## Where the category number misleads The dominant error is treating a category percentage as a measurement of a harm class rather than of specific strings. Concretely: - a 40% figure in a six-row category is two or three responses and should send you to the transcripts, not to a roadmap; - a 95% figure in a thirty-row category that clusters to four distinct ideas is a four-idea result with two decimal places of false precision; - a drop between runs in either kind of category is more often sampling noise or a judge revision than a model regression; - and a same-named category compared across two lists is a comparison of unlike populations that will look meaningful and be meaningless. If the run used more than one generation per row, check whether the reported figure is **per-row** (any generation succeeded) or **per-attempt** — the two differ by a large factor on the same transcripts. ## What good practice looks like - Pin the list revision and record it in every report. - Publish counts with every percentage. - Deduplicate before aggregating, or at minimum publish cluster counts. - Keep the composition note beside the list. - And when someone asks whether the product is weak in a category, answer by naming the rows behind the number and quoting a failing response — an answer nobody can give without having opened the file.

  • A category shows a large drop between two runs. What do you look at before reporting a regression?
    How many entries that category has, whether they are near-duplicates of one idea, and the actual failing responses. Small or redundant categories move on very little.
  • Two behaviour lists both have a category with the same name. Can you compare those two numbers?
    No. The labels are separate taxonomies written by separate authors; only the entries define the category, and the entries differ.
  • Cheapest useful check on a behaviour list you have just adopted?
    Sort entries by category, print the counts, and read a sample of twenty per category. It takes about an hour and usually changes what you believe the numbers mean.

Ten prints of one photograph do not make ten pieces of evidence, but a counter that tallies prints will say they do. Near-duplicate rows in a behaviour list weight one idea ten times and make the result look ten times more consistent than it is.

saying these in an interview costs you the question

  • Quotes a category percentage without knowing how many entries it has.
  • Assumes each row in a behaviour list is a distinct behaviour.
  • Compares identically named categories across two different lists.
  • Treats a blunt single-turn request list as evidence about multi-turn conversations.
  • Has never opened the list being cited.

context