Another team publishes "2% attack success" on the same public harmful-behaviour list you used, where you measured 30%, and their write-up never states attempts per behaviour. What can you legitimately conclude, and what do you do to make the two comparable?
answer
- unlabelled rate = lower bound
- match n, hit rule, attack, target config
- reproduce, don't argue
- publish your rate at n=1 too
- unreproducible figure = citation not measurement
basics
~20 sAlmost nothing about the gap. Their 2% may be one attempt per behaviour against your twenty. Treat it as a lower bound, ask for attempts per behaviour, the stopping rule and what ruled a hit, and re-measure yourself at a declared attempt budget before putting the two numbers in one table.
solid answer
~50 sThe honest conclusion is that the two figures are not on the same axis. Because best-of-n aggregation is monotone in the attempt budget, an unlabelled 2% is only a statement that *at whatever budget they bought*, 2% of behaviours fell — it is a lower bound on what a better-funded run would report. Three things must match before a comparison means anything: the attempt budget per behaviour, what ruled an attempt a hit, and the attack strategy. A stricter hit rule alone can move a rate by tens of points with n identical, so "they used fewer attempts" is a hypothesis you check, not a conclusion. The practical move is to reproduce rather than argue: run their configuration as far as it is documented, at your own declared n, on the same list, with your own hit rule, and publish both your number and theirs with budgets attached. Where their setup is undocumented, say so — an unreproducible figure is a citation, not a measurement.
go deeper
Notices the missing attempt count and does not treat the two percentages as directly comparable.
Names the axes that must match — attempts, hit rule, attack — and asks for the missing parameters before drawing any conclusion.
Reproduces at declared settings, presents their likely n alongside yours, and includes the target's surrounding configuration as a difference that changes the system being measured.
Fixes the reporting convention so external and internal figures arrive with budgets attached, rather than adjudicating each dispute by hand.
**Read their number charitably first.** Under any-hit aggregation the reported rate is monotone in the attempt budget, so an unlabelled figure carries exactly one honest reading: *at whatever budget they bought, that fraction of behaviours fell.* It is a lower bound on what the same run would report with more attempts. Their 2% and your 30% are therefore not contradictory. They are entirely compatible with one model, one list, and two budgets — 2% at n = 1 and 30% at n = 20 are ordinary readings of the same system. Announcing a fifteenfold robustness advantage on that evidence is precisely the error being probed. **The axes that must match before a comparison means anything.** | axis | why it moves the rate | |---|---| | attempts per behaviour | any-hit aggregation is monotone in n; this is the axis their write-up omitted | | the hit rule | a refusal-string check, a trained classifier and a judge model disagree sharply on borderline completions; a swap can move the rate by tens of points with n untouched | | the attack | a single-turn template and an adaptive multi-turn strategy are different experiments even on identical behaviours | | the target's surrounding configuration | system prompt, decoding temperature, and any input or output filter in front of the model; a bare endpoint and the shipped stack are different systems | | the list version | public lists get revised and subsetted; "the same list" is a claim to verify, not assume | The order matters when you investigate. Attempts is the most likely explanation and the cheapest to obtain, but "they used fewer attempts" is a hypothesis you test, not a conclusion you publish. A refusal-string hit rule that counts any non-refusal as success will typically report *higher* than a judge model that demands the completion actually be harmful and on-topic — so a difference in the other direction can also be a hit-rule artefact. **What it costs to settle.** Asking is nearly free and usually works: attempts per behaviour, stopping rule, hit rule, attack, target configuration is a five-line reply, and omitting them is overwhelmingly a reporting-convention failure rather than manipulation. Reproducing is not free. Re-running a documented configuration on a few-hundred-behaviour list at n = 20 is thousands of target generations plus the matching judge calls, hours of wall clock against a rate-limited endpoint, and a day or two of engineer time to wire up their harness and reconcile their hit rule with yours. Budget it deliberately, because the value of the exercise is a comparable table, not the satisfaction of winning an argument. **The move.** Ask for the missing parameters. Failing that, reproduce as far as their setup is documented, at your own declared n, on the verified list version, with your own hit rule, and publish a table whose every row carries list and version, attempts per behaviour, hit rule, attack and target configuration. Because you kept per-attempt outcomes, also present *your* rate at n = 1 — the value they most likely used — which makes the comparison as favourable to them as your evidence permits and removes the suspicion that you picked a budget that flattered you. Where their configuration is undocumented, say so plainly: an unreproducible figure is a citation, not a measurement, and belongs in prose rather than in the comparison table. **Where this misleads if you get it wrong.** Copying their number into your table under your column headings silently asserts it was measured under your conditions, and every downstream reader will treat it that way. Averaging the two produces a figure describing no run at all. Concluding "the model changed" imports a cause the data cannot distinguish from four other causes. And treating the omission as bad faith is both usually wrong and strategically bad — the durable fix is a caption standard both write-ups adopt, not an adjudication of this one dispute. **What you would check.** That the list you both used is the same version and the same subset. That their stopping rule is not early-stop-on-first-hit, which changes what per-attempt statistics you can recover from their logs. That the hit rules are stated well enough to be re-implemented. That their target was measured bare or behind a filter, and which one yours was. And whether their write-up reports any figure at more than one n — if it does, the slope tells you where on the saturation curve their single headline number sits.
- They reply that they used the same attempt budget as you. What is the next axis you check?What ruled an attempt a hit. A refusal-string check, a trained classifier and a judge model give materially different rates on identical transcripts, with n unchanged.
- What is the single line you would add to both write-ups to make this dispute impossible next time?A caption that states behaviour list and size, attempts per behaviour, the hit rule, the attack strategy and the target's surrounding configuration.
saying these in an interview costs you the question
- Concluding the other system is fifteen times more robust.
- Assuming the difference must be the attempt budget without checking the hit rule and attack.
- Copying their number into your table under your own column headings.
- Comparing a bare-endpoint measurement with one taken behind an input or output filter.
- Treating the omission of n as deliberate manipulation rather than a missing convention.