A stakeholder points at a published safety-benchmark attack-success rate for the base model behind your product and asks why your own red-team scan of the deployed assistant produced a very different rate. What does the published figure's denominator actually enumerate, and why do the two numbers not sit on one axis?
answer
- benchmark measures the model, scan measures the system
- system prompt, retrieval, tools, filters
- their rubric vs your triage
- raw hits vs deduplicated findings
- wrapper on versus wrapper off
basics
~20 sThe published denominator enumerates that suite's fixed behaviour list, attacked by that suite's method, against the model as the suite configured it - usually raw weights with a plain chat template, no product system prompt, no retrieval, no filters. Your scan enumerates your own attack set against the deployed stack. Different population, target and ruling.
solid answer
~60 sName the three gaps. **Population.** The suite's denominator is its own curated behaviour list, built to be broadly comparable across models. Your scan's denominator is whatever your team chose to probe on your product's surface — its tools, its data, its user flows. Neither list is a sample of the other. **Target.** A benchmark almost always measures the model in isolation. Your deployment adds a system prompt, possibly retrieval and tools, input and output classifiers, rate limits and a refusal-tuned product persona. Those can push the rate down; retrieved content and tool access can also open paths a bare model never had, pushing it up. **Ruling.** The suite's judging procedure decides harmful compliance against its own rubric. Your triage decides what counts as a reportable issue for your product, which is a business definition, not the same rubric. So the honest answer is that the published number is a property of the model-as-tested and belongs in a model-selection conversation; your scan is a property of the deployed system and belongs in a remediation conversation.
go deeper
Should at least say the benchmark tested the model while the scan tested the whole product.
Should list the population, target and ruling differences and decline to put the two on one axis.
Should decompose the chain hop by hop, argue the direction of the gap is not predictable, flag the raw-hit versus deduplicated-finding unit mismatch, and propose the wrapper-on/wrapper-off experiment.
Should separate the two series as organisational metrics with different owners and decisions, and stop either being quoted as a proxy for the other.
The stakeholder is really asking which number to believe. The answer is that the two measure different objects, and the useful reply names the decision each one supports. ### What the published denominator enumerates A safety-benchmark ASR enumerates that suite's own curated behaviour list — a fixed set of harmful request descriptions, built to be broadly comparable across models — attacked by that suite's fixed method at a fixed budget, against the model configured as the suite configures it. In almost every case that means the raw weights or a bare hosted endpoint with a plain chat template: no product system prompt, no retrieval corpus, no tools, no input or output classifier, no rate limits, no session state. The ruling is the suite's own harm rubric, applied by its own classifier or judge model. Your scan enumerates a different denominator entirely: whatever probes your team chose for *your* product's surface, driven by whatever methods your engagement ran, against the deployed stack, ruled by your triage definition of a reportable finding. ### Decomposing the gap hop by hop 1. **Item set** — their behaviour list to your probe set: different categories, difficulty and size, and neither is a sample of the other. 2. **Attack** — their fixed method and budget to an engagement's adaptive, often multi-turn probing. 3. **Target surface** — bare weights to a product: system prompt, tools and actions, retrieval corpus, memory, session state. 4. **Defences in path** — none, to input and output classifiers, refusal tuning and rate limiting. 5. **Ruling** — their harm rubric to your business definition of "reportable". 6. **Unit of the number** — their per-behaviour or per-attempt *rate* to your report's count of triaged, deduplicated *findings*. Step 6 alone ends the comparison: a finding is a class of failure that may cover dozens of raw hits, so "12 findings" and "18%" are not the same kind of quantity at different scales — they are different kinds of quantity. ### Direction is not predictable The tempting assumption is that the deployed stack must score at or below the published rate because defences were added. Often true for the suite's own behaviours. But a product also carries surfaces the suite never touched: content pulled into context from documents or the web, tool invocation with real side effects, multi-turn state that accumulates. Those can open paths a prompt-only attack on bare weights could not reach, so the deployed number is not bounded above by the published one, and treating the published figure as a ceiling is how a product ships believing it inherited a safety property it never had. ### What the two measurements cost They are not the same size of job. A benchmark sweep is machine time: list size times attempts times methods generations plus judging, hours of wall clock, a bill in tens to hundreds of dollars against a metered API — cheap enough to repeat per model and per release. An engagement against a deployed product is engineer-weeks: probes written for your surface, human triage and deduplication of raw hits into findings, reproduction of each finding, and inference against your own stack including every guard call in the path, which both raises per-attempt cost and slows the sweep because production rate limits are tighter than a lab's. Human triage, not inference, is the dominant line item, and that is precisely why the output is a finding count rather than a rate. ### What you would actually do If the question is **model selection**, run the same suite yourself against each candidate under one configuration and one judge, and compare those. If the question is **product risk**, run your own probe set against the deployed stack and track it across releases. Keep the two series separate, labelled, and owned by different decisions, and never let one be quoted as a proxy for the other. If someone genuinely needs to know what the guardrails bought, there is a clean experiment: run *your* probe set against the same model with the product wrapper enabled and then disabled, judged identically both times, at the same attempts per item. Only the wrapper varies, so the difference is attributable to it. That is a real number; the cross-axis subtraction is not. Also record raw hits alongside deduplicated findings in the engagement, so that at least one of your own quantities is a rate and can be trended.
- The stakeholder insists on one comparable number. What experiment do you offer instead?Run your own probe set against the same model with the product wrapper enabled and disabled, under one ruling procedure. That difference is attributable to the wrapper, unlike a cross-suite comparison.
- Why can the deployed system score higher than the bare model on your probe set?Because the product adds surfaces the bare model lacks: content pulled into context, tools and actions, multi-turn state. Those can create paths that no prompt-only attack on raw weights could reach.
- What unit mismatch hides inside comparing a scan to a benchmark rate?The scan's report counts triaged, deduplicated findings while the benchmark counts behaviours or attempts. One finding may correspond to many raw hits, so the two are not the same kind of quantity.
A crash-test star rating for a bare chassis and a fleet's real-world accident rate for the finished car are both about safety, but subtracting one from the other tells you nothing about the airbags. To price the airbags you crash the same car twice, with them and without.
saying these in an interview costs you the question
- Claiming the deployed system must score at or below the published rate because defences were added
- Treating a published benchmark rate as a baseline for a product risk metric
- Comparing a count of deduplicated report items with a per-attempt benchmark rate
- Ignoring retrieval and tool surfaces that the benchmark never exercised
- Explaining the gap purely by model version and stopping there