skip to content

Your team adds twenty product-specific behaviours to an evaluation that previously used only the HarmBench behaviour list, and the headline number moves. What has actually changed, and how would you report the two sets of behaviours?

level: seniorimportance: should knowfreq 45%

answer

  1. population changed, not the model
  2. category mix reweighted
  3. two blocks, never one headline
  4. version the list like code
  5. revision change = broken trend line

basics

~20 s

You changed the measured population, not the model. The mixed figure is comparable to nothing: not to your earlier run on the public list, and not to anything anyone else reports on it. Report the public list and your own behaviours as two results, each with its own count, and never merge them into one headline.

solid answer

~50 s

Adding entries changes three things at once. The population is different, so the movement is confounded with the model. The category mix is reweighted — twenty entries in one theme pull the aggregate toward that theme. And your entries are probably not the same difficulty as the public ones, so the direction of the move tells you about your additions, not about robustness. The reporting rule follows: keep the public-list result untouched and unmixed so it remains comparable across models, model versions and time, and report the in-house register as its own result with its own count and per-category breakdown. Version both lists like source code, and cite the revision in every number. **The moment you edit a behaviour list, you have started a new baseline** — previous numbers on it are history, not a trend line.

go deeper

for a junior

Recognises that adding behaviours changes what was measured, so the before and after are not the same test.

for a middle

Explains the confounded comparison and the reweighted category mix, and reports the two lists separately.

for a senior

Adds versioning and review of the behaviour register, per-category reporting, duplicate detection, and broken trend lines at a revision change.

for a principal

Sets the rule for who may edit the register, what a reported number must cite, and how release evidence survives a list change.

**What actually moved.** A behaviour-set result is an aggregate over a population of rows. Change the population and the aggregate changes mechanically, with no change whatever in the system under test. Put numbers on it. Suppose the published list holds four hundred behaviours and the model fails eight of them: 392/400, a 98.0% pass rate. Append twenty in-house behaviours that this model happens to handle easily and the figure becomes 412/420 = 98.1% — a small rise that a dashboard renders as progress. Append twenty that hit your real weakness, of which ten fail, and it becomes 402/420 = 95.7% — a 2.3-point drop that will be read as a regression. Both movements came entirely from arithmetic. From the headline alone you cannot separate them from a genuine change in the model, so the comparison with last month is not a comparison of one measurement over time; it is two different measurements sharing a name. **Composition effects.** An aggregate over a behaviour list is implicitly weighted by how many rows each category holds. Twenty additions in a single theme make that theme the heaviest block in the mix, so a fix there moves the headline more than a regression anywhere else, and the headline stops tracking overall robustness at all. Per-category reporting with raw counts makes the weighting visible; one number hides it by construction. **Double counting.** In-house rows frequently restate a public row in local vocabulary — “assist with fraudulent invoicing” beside the public list's generic fraud entry. That silently inflates the theme's weight and destroys the one claim the register exists to make: that it measures the risk the public list does not. Similarity-check additions against the published rows before merging them into any run. **What the change costs.** Compute is not the constraint: twenty extra behaviours at five generations is a hundred extra completions plus judging, cents. The costs that matter are (a) the analyst time to author and review entries, roughly an hour or two each with its judging rule, (b) the loss of every historical number on the old population — you have paid for a new baseline and cannot buy the old trend back, and (c) the ongoing governance load, because a list that anyone can edit needs review on every change forever. Treat the edit as a schema migration, not a config tweak. **Where the number misleads, and how it gets abused.** Two temptations arrive the moment a team can edit its own measuring instrument. Adding behaviours the model already handles lifts the aggregate at no cost to anybody's safety; deleting behaviours that keep failing lifts it faster, and is the serious one. Neither shows up in the headline; both are invisible unless the register is under review. A third, less cynical failure is the unlabelled revision: a number reported without the revision it was computed on cannot be reproduced, and six months later nobody can say which population produced it. And the chart is the worst offender — a continuous line drawn across an edit asserts a comparison that was never made. **How to report it.** Two blocks side by side, never one figure. ``` Published list HarmBench, rev <pin> attempted 400 failed 8 per-category counts, attack used, judge used In-house register rev <pin> attempted 20 failed 10 per-category counts, target config incl. tools Note: the published list contains no behaviour covering RR-07 (cross-customer disclosure). ``` Trend lines are drawn per block, and only across runs where that block's revision is unchanged; at a revision change the series gets a marked break with the new row count annotated, not a continuous line. **What I would check first** on being shown a moved headline: did anyone touch the behaviour list, and at which revision was each number computed? Diff the two revisions and count added, removed and modified rows before anyone opens a model investigation. In practice that single question resolves most of these surprises within minutes, and it is the reason the register belongs in version control with reviewed changes rather than in a spreadsheet.

  • Your dashboard plots one behaviour-set number over six months and the list was edited in month three. What should the chart do?
    Break the series at the edit and label the revision. A continuous line across a population change is a false trend.
  • How do you catch an in-house behaviour that duplicates a public-list entry?
    Similarity-check new entries against the public list before adding them, and review additions like code. Duplicates silently reweight the theme and blur which list covers what.
  • Is it ever right to delete a behaviour from your own register?
    Yes, when the product genuinely no longer has that capability or the entry is unjudgeable. Never because it keeps failing, and the deletion is reviewed and recorded either way.

Adding twenty questions to an exam and then comparing this year's class average with last year's tells you about the questions you added, not about the students. Same paper, same trend line; different paper, new baseline.

saying these in an interview costs you the question

  • Interprets the moved headline as a change in model robustness.
  • Merges the public list and the in-house register into one reported figure.
  • Keeps a continuous trend line across a behaviour-list edit.
  • Adds behaviours the model already passes to lift the aggregate.
  • Cannot say which revision of the list a past number was computed on.

context