skip to content

You lead red teaming for a company shipping a code assistant, a bank chat assistant and a medical triage bot. Would you make a standard harmful-behaviour set such as HarmBench the organisation's unit of measurement, keep a per-product behaviour register, or both? Which would you choose, and what is each number allowed to justify?

level: principalimportance: should knowfreq 35%

answer

  1. both lists, different authority
  2. shared = regression and model choice
  3. register = ship decision
  4. Goodhart on a shared number
  5. state what the list omits

basics

~20 s

Both, with different jobs. The shared list is a cheap cross-product regression signal and lets you compare candidate models. It cannot represent product risk, so every product also owns a behaviour register drawn from its own threat model, and that register, not the shared figure, gates release.

solid answer

~50 s

A single org-wide list is attractive because it gives one comparable number across products and across model swaps, and one team can maintain it. It fails as a release gate for three reasons: its taxonomy was written without your products, it cannot express tool- or domain-mediated harms, and once a number is the goal, teams optimise toward it and the unmeasured risks stay unmeasured. Per-product registers only is the opposite failure: no comparability, no way to assess a model change, three teams rewriting the same generic harms at three quality levels. So I would run both and separate their authority. **The shared list justifies model-selection and regression claims; the product register justifies ship decisions.** Central team owns the shared list and its revisions; each product owns its register with review. Neither number is ever averaged into the other, and both cite their list revision.

go deeper

for a junior

Recognises that one shared list cannot cover three very different products.

for a middle

Proposes shared list plus per-product behaviours and keeps the two results separate.

for a senior

Assigns authority to each number, requires target configuration and list revisions in reports, and plans register refresh when a product gains tools.

for a principal

Owns the incentive design: what each measurement may justify, how to prevent optimisation to the shared figure, the over-refusal counterweight, and what the bespoke work is worth in analyst time.

**The judgement is about authority, not about which list is better.** Run both, and write down what each number is permitted to prove. Everything else follows from that one decision. **The shared list.** Its single virtue is comparability: the same rows, the same attack, the same judge, across three products and across candidate models. That makes it the right instrument for two questions — “has this model version regressed on generic harms since the last one” and “of these two candidate models, which starts from a better place”. Its cost is that it is product-agnostic by construction, so it can never be evidence that a product is safe to ship. One central team maintains it, one revision is pinned per quarter, and its result is labelled a regression signal on the page where it appears. **The product registers.** Each product's realistic worst outcome is specific: the code assistant proposing a dependency or a snippet that quietly weakens the caller's security posture, the bank assistant resolving the wrong customer identity and disclosing their data, the triage bot giving dangerous advice or failing to escalate a red-flag presentation. None of these is in a generic list. Two of the three are only reachable through the product's own tools and several turns of conversation, and the triage failure needs clinical review to judge at all — no shipped classifier can decide whether an escalation was owed. These behaviours are written locally from the threat model, and a release decision rests on them. **What this costs, concretely.** The shared list is the cheap half: pinning a revision, running it per model candidate and per release, and reading the failures is a few days per quarter for one owner, plus a small compute bill — a few hundred rows at a handful of generations, plus judging, is dollars. The registers are the expensive half and the cost is people: an hour or two to author each behaviour with its target configuration and judging rule, fifteen to twenty behaviours per product, a refresh whenever a product gains a tool or a data source, and for the triage bot a clinician's time to adjudicate outcomes — which is the real budget line and the one that gets cut first. Size it honestly up front: the top handful of risks per product get behaviours, and you say out loud that risks six and below are unmeasured this quarter. **Second-order effects to plan for.** - *Goodhart.* The moment the shared figure appears in a launch deck or a performance review it becomes a target, and the cheapest way to move it is to work on generic harms that were never the product's risk. Keep it explicitly stripped of ship authority. - *Over-refusal.* Scoring only refusal of harmful behaviours rewards refusing, so a triage bot drifts toward declining legitimate clinical questions and the measurement applauds. Each product needs a benign counterpart set — requests that must be answered — reported beside the harmful one, or the trade is invisible. - *Divergence.* Three registers drift in quality within two quarters. A shared entry template, a common report shape and central review of additions and deletions hold them together while the content stays with the team that owns the threat model. **Where the numbers mislead.** A blended company-wide safety score is the classic mistake: it is comparable to nothing external, it hides which population moved, and it lets a strong shared-list result mask a weak register. A per-product register with twenty behaviours produces small counts, so quarter-to-quarter movement is mostly noise — report counts and read transcripts rather than charting a rate. And a rising shared-list number across a model migration can simply mean the new model was tuned on those public rows. **How I would report it, and what I would check.** One page per product: shared-list block (name, revision, rows attempted, per-category counts, attack, judge), register block (revision, behaviour count, per-category counts, target configuration including which tools were live), and one line stating what the shared list does *not* include for this product. The check I would run on any team claiming coverage is to ask the product owner to name the register entry that covers their top risk and show me the transcript of it failing at least once historically. If they cannot, the register is decoration.

  • What one artefact stops a shared behaviour-set result from being read as product coverage?
    A written statement, attached to the number every time it is quoted, naming what that list does not include for that product.
  • How do you keep three product registers from drifting apart in quality?
    A shared template for entries, a common report shape, and central review of additions and deletions — while leaving the content with the product team that owns the threat model.
  • Why does scoring only harmful behaviours distort a medical triage product?
    It rewards refusal, so the product drifts toward refusing legitimate clinical questions. You need a benign counterpart set to see that cost at all.

saying these in an interview costs you the question

  • Picks a single org-wide behaviour list as the release gate for all three products.
  • Proposes one blended company-wide safety score.
  • Ignores over-refusal when the measurement rewards refusal only.
  • Assigns no owner or review process for edits to the behaviour lists.
  • Cannot say what evidence each number is and is not allowed to support.

context