skip to content

Two candidate builds give you an attack-success rate on a red-team suite and a refusal rate on a benign prompt set, and between the builds the two rates move in opposite directions. As the red team, how do you report that pair so nobody is misled, and what claims do you refuse to make from it?

level: principalimportance: should knowfreq 27%

answer

  1. no blended safety index
  2. different denominators, different rubrics
  3. slice breakdown + transcript exemplars
  4. state what was held constant
  5. name the tradeoff, don't price it

basics

~20 s

Report both rates per build with their prompt-set sizes, labelling rules and run configuration, and state the direction each moved. Refuse to combine them into one index, to declare a winner, and to imply the two rates are in the same units. Add the qualitative slice: which legitimate requests the safer build now declines.

solid answer

~50 s

The pair is two measurements on different prompt sets under different labelling procedures, so equal-sized moves on each side are not comparable quantities, and any weighting that trades one against the other encodes a product value judgment the red team does not own. So the deliverable is: per build, both rates, each with its prompt-set size, its labelling rubric, the decoding configuration and system prompt used; the delta on each side; and the benign slice breakdown showing *which* legitimate requests the safer build lost. Add a handful of real declined benign transcripts — worth more to a decision than a percentage. What I refuse: a single blended safety score, a recommendation phrased as "build B is safer" without "and declines this class of legitimate request", and any claim that the pair generalises to traffic the prompt sets do not resemble. Naming the tradeoff precisely is the job; pricing it is the product owner's.

go deeper

for a junior

Should report both rates side by side rather than only the attack number.

for a middle

Adds the prompt-set sizes, labelling rules and held-constant run configuration, and knows not to average the two into one score.

for a senior

Adds the slice breakdown and transcript exemplars, and states the population caveat that neither rate forecasts production incidence.

for a principal

Frames the refusal to blend as an accountability boundary — weighting the two rates is a product judgment with an owner — and designs the report so the refusal cost cannot be dropped from the decision.

**What is actually on the table.** Two builds, four numbers: an attack-success rate and a benign-refusal rate for each. The attack rate counts attempts that produced a targeted harmful behaviour, over attempts made. The benign-refusal rate counts harmless prompts the model declined, over prompts sent. Between the builds they move in opposite directions, which is the ordinary case for a safety intervention, and the entire question is how to hand that over without the recipient collapsing it into a ranking. **Why not one number.** A blended index needs a weight, and the weight *is* the decision: how much user friction one avoided elicitation is worth. That depends on who the users are, what the deployment surface is, what a hit would actually cost, and what the regulatory and reputational exposure looks like — none of which is contained in either rate. Publishing a blend launders a product value judgment into arithmetic, and it makes two builds look ordered when they are only different. Worse, whoever picked the weight is usually the analyst, which quietly moves the tolerance decision out of the accountable seat. **Why the deltas are not commensurable.** They differ on every axis that would make a subtraction meaningful: | | attack-success rate | benign-refusal rate | |---|---|---| | denominator | attack attempts made | benign prompts sent | | labelling | presence of targeted content | whether the request was served | | label objectivity | near-objective, an artefact to point at | graded judgment, rubric-dependent | | typical set size | a few hundred behaviours, often multiple attempts each | a few hundred prompts | "Down four points here, up four points there" is a coincidence of scale, not a wash. A four-point attack move on 400 behaviours is about sixteen behaviours that no longer elicit; a four-point benign move on 300 prompts is about twelve legitimate requests now declined. Whether those trade evenly is a question about the product, not about the arithmetic. **What the report contains.** Per build: both rates, each with its set size and its named labelling rubric. The serving configuration held constant across builds — system prompt, decoding settings, maximum output length — stated explicitly, because a difference in any of them makes the comparison meaningless. The benign breakdown by slice, so a surface-keyed intervention is visible rather than pooled away. Transcript exemplars from both sides, especially five or six real declined benign requests: decision-makers reason far better from concrete refusals than from a delta. And a population caveat, because both prompt sets are convenience samples assembled by the team, not a model of real traffic. **What it costs to do it this way.** The extra work over shipping one number is real but small: the slice breakdown is a grouping over labels you already have, and pulling exemplars is an hour of reading. The honest version of this deliverable costs perhaps half an engineer-day more than the dishonest one, and the dishonest one is shorter, cleaner and easier to present — which is exactly why it wins by default unless someone insists. The larger cost is political, paid by the person who declines to name a winner in a room that wants one. **Where the numbers mislead.** The blended-score trap is the first. The second is base rate: neither rate forecasts production incidence, because production traffic is not drawn from either prompt set. If attack-shaped traffic is a small fraction of real requests while sensitive-looking legitimate requests are common, a build that trades a few attack points for a few refusal points can be far worse in production than the symmetric-looking table suggests — and the report cannot tell you that, which is precisely why it must say so rather than imply otherwise. The third is composition: if one build's benign decline is concentrated in a single slice that happens to be a core product use case, a pooled rate that moved two points understates a change that breaks a feature. The fourth is presentation order — an attack improvement in the headline with the refusal cost in an appendix is a real distortion even when every number in the document is correct. **What you refuse to claim.** That one build is "safer" full stop. That the pair licenses a threshold or a ship decision. That a rate measured on this suite transfers to attacks or user requests the suite does not contain. And you refuse to let a build be selected on the attack column with the benign column reduced to a footnote: if the organisation wants to buy safety with refusals, that should be a stated and owned choice rather than a side effect of which column made the slide. **The organisational point.** The recurring failure here is not statistical, it is about who holds the pen. A red team that publishes one ranked number has silently made the product's tolerance decision for it. Publishing the pair keeps that decision where accountability for users sits, and it preserves the red team's credibility — which is spent in full the first time a build declared "safer" ships and users report it refusing ordinary work.

  • Leadership insists on one number for a slide. What do you give them?
    Both rates as a labelled pair with the slice breakdown one click away, plus two or three real declined-benign transcripts. If a single figure is unavoidable, it is theirs to define with a stated weight and their name on it, not mine to invent.
  • What single addition most improves how this pair is used in practice?
    Concrete exemplars of the legitimate requests the safer build now declines. People reason about five real refusals far better than about a percentage, and it moves the discussion to which refusals are actually acceptable.

The two rates are prices quoted in different currencies. Adding them requires an exchange rate, and that rate — how much user friction one blocked harmful answer is worth — is a business decision neither measurement contains.

saying these in an interview costs you the question

  • Producing a single weighted safety score and calling one build the winner.
  • Presenting the attack-rate improvement in the headline with the refusal cost in an appendix.
  • Treating an equal-sized move on each side as a wash, despite different denominators and rubrics.
  • Implying the suite's rates predict incidence in production traffic.
  • Comparing builds that also differ in system prompt or decoding configuration.

context