A safety-tuned build drops your attack-success rate on the same jailbreak suite from 22% to 6%. What do you measure, and how do you design the comparison, to show that this is real hardening rather than the model simply refusing more of everything?
answer
- benign set is the control
- two-by-two: attack down, benign flat
- slice plain vs sensitive-looking
- hold system prompt + decoding fixed
- held-out attacks vs tuned-on ones
basics
~20 sRe-run the frozen benign prompt set against both builds and compare refusal rates. Real hardening leaves the benign rate flat while the attack rate falls; a blanket refusal shift moves both together. Hold the prompt sets, system prompt, decoding settings and labelling rubric identical, and change only the build.
solid answer
~50 sThe attack rate alone cannot distinguish the two hypotheses, because both of them lower it. You need the benign side as the control. Design it as a paired comparison: one frozen attack set, one frozen benign set, both builds, identical system prompt, identical decoding configuration, identical labelling rubric. Then read the two-by-two: - attack down, benign flat → hardening; - attack down, benign up by a comparable margin → a refusal shift, not a defence; - attack down, benign up slightly → a real trade you now have to describe in both units. Strengthen it by slicing the benign set into a plain-harmless half and a sensitive-looking half. A topic- or keyword-driven intervention shows up as a large move on the sensitive-looking slice with the plain half untouched — a signature that a single pooled benign number hides. Also sanity-check that the surviving 6% are not simply the attacks whose prompts changed shape between runs.
go deeper
Should at least say the harmless prompt set must be re-run against both builds and the refusal rates compared.
Lays out the paired design with frozen prompt sets and identical decoding, and reads the attack/benign two-by-two correctly.
Adds the benign slice split that exposes a topic-filter signature, blind mixed labelling, and the alternative explanations — truncation, system-prompt change, empties scored as non-hits, attacks present in the tuning data.
Makes the honest reporting of a risen benign rate a non-negotiable part of the deliverable and defines up front which held-out slices exist so the result cannot be tuned toward.
**Why the attack number cannot answer this.** "Learned to resist these attacks" and "became more reluctant in general" are observationally identical on a one-tailed instrument. Every extra refusal, whatever produced it, subtracts from the attack-success rate. The two hypotheses only disagree somewhere the attack suite never looks — on requests that should be answered — so the benign prompt set is not a nicety here, it is the control arm of the experiment. **The comparison design.** Four runs, one variable. 1. Freeze both prompt sets before the tuning run, so neither can be adjusted once the result is visible. 2. Run both builds under the same system prompt, the same decoding configuration and the same maximum output length. A temperature or system-prompt difference alone moves both rates. 3. Label in one mixed, source-blind batch under one rubric, so the labeller cannot tell which build produced which response. 4. Report per build: attack rate and benign-refusal rate, each with its set size, plus both deltas. **Slices carry the diagnosis.** A pooled benign number can hide the mechanism, so split it into a *plain harmless* slice — everyday requests with no sensitive surface — and a *sensitive-looking but legitimate* slice: clinical, security-education, fiction-violence, self-harm-support, dual-use technical. A filter that keys on topic or vocabulary barely touches the first slice and hammers the second. A genuine improvement in judging intent moves neither much. That shape is the evidence; the pooled percentage is only a summary of it. **What it costs.** Multiply everything by four. An attack suite of roughly 400 behaviours at five generations each is about 2,000 calls per build; a 300-prompt benign set is another 300; two builds makes roughly 4,600 model calls, plus a judge call per attack response if the suite scores with a model, which roughly doubles the call count. On a hosted mid-priced endpoint that is tens of dollars, not hundreds — the money is never the constraint. Wall-clock is set by rate limits: at 60 requests per minute a 9,000-call sweep is a couple of hours, and a self-hosted target with a single GPU can be considerably slower. The binding cost is again human: four labelled sets, of which the two benign ones need careful reading, is most of a day per comparison, and a blind mixed batch means you cannot reuse last week's labels. Budget one engineer-day per paired comparison and you will be about right; budget one afternoon and the benign side is what gets dropped. **Where the number misleads.** Start with the headline itself: 22% to 6% is a 73% relative reduction, which reads as transformative, while the absolute move is 16 points — on a 400-behaviour suite, about 64 behaviours that used to elicit and now do not. Relative framing on a small base is the standard way a modest change is oversold. Then the alternative explanations, each of which produces this drop with no change in safety judgment at all: a shorter maximum output length truncating completions before the targeted content arrives; a changed system prompt suppressing the persona the attacks leaned on; a retry or timeout path returning empty strings that the judge scores as non-hits; and contamination, where the attack prompts were themselves in the tuning corpus, so the drop is memorised on this list and evaporates on a held-out one. Finally, sampling error: with about 400 behaviours the standard error on a rate near 20% is roughly two points, so the pooled 16-point move is comfortably real, but a five-point move on a 60-prompt benign slice is not, and slice-level deltas need intervals before they are read as signal. **What you check before calling it hardening.** Read the two-by-two: attack down with benign flat is hardening; attack down with benign up by a comparable margin is a refusal shift wearing a defence's clothes; attack down with benign up slightly is a real trade you now have to describe in both units. Then rule the alternatives out one at a time — diff the serving configuration between builds, count errored and empty responses on both sides, and score a held-out attack slice that was kept out of every tuning corpus. Verify that the surviving 6% are the same behaviours the old build also failed on rather than a different set, because a rotating failure set is a sign of variance rather than progress. And check labeller agreement on the mixed sample before you trust the benign delta at all. **What you write.** Both builds, both rates, both deltas, the benign slice breakdown, the held-out attack result, and an explicit list of what was held constant. If the benign rate rose, it belongs in the headline, not a footnote — a defence bought entirely with refusals is a product decision somebody else has to make knowingly, and the red team's job is to make sure the price tag is next to the number when they make it.
- Attack rate fell and the benign refusal rate rose by a similar margin. What do you write?That the build has not been shown to be more robust — it is more refusal-prone, and the attack-rate move is consistent with that alone. Report both deltas together and decline to call it hardening.
- How do you tell a memorised drop from a generalised one?Hold out a slice of attack behaviours from any tuning corpus and score it separately. A drop concentrated on the tuned-on slice, with the held-out slice flat, is memorisation on that list, not a defence.
- Both rates were labelled by different people on different days. Does that matter?Yes. Benign-side labelling is a judgment call, so labeller drift alone can move the rate. Re-label a mixed, source-blind sample from both builds under one rubric and check agreement before trusting the delta.
A goalkeeper who concedes fewer goals may have improved, or may have had the pitch shrunk. You find out by checking whether the team still scores — the attacking half of the scoreline is the control.
saying these in an interview costs you the question
- Declaring hardening from the attack-rate drop alone.
- Comparing builds that differ in system prompt, max output length or decoding settings as well as weights.
- Only reporting a pooled benign number, so a topic-filter signature is averaged away.
- Evaluating on attack prompts that were in the tuning data and treating the drop as generalisation.
- Burying a risen benign-refusal rate in an appendix while headlining the attack-rate drop.