A vendor publishes a red-team benchmark score for a model, and your product calls that same model behind a system prompt, a tool loop and a retrieval layer. What interface did that published number actually measure, and why can you not quote it as your product's safety number?
answer
- score stops at the model API
- wrapper never in the run
- two numbers, two interfaces
- not an upper or lower bound
- component claim vs product claim
basics
~20 sThe number came from prompts sent straight to the bare model endpoint - no system prompt, no tools, no retrieved text. Your product is a different system: extra instructions, extra input channels, extra output paths, none of which that run exercised. So the score describes the model you call, not the application you ship.
solid answer
~50 sPublished red-team suites drive the model API directly. Each item is one user turn, the context is empty or a fixed neutral preamble, and the harm judge reads the raw completion. Your deployment changes all three of those: a system prompt that constrains scope and tone, a tool loop that lets a compliant answer become an action, and retrieval that pushes text you did not author into the same context window. So the published score is a statement about a **component**, useful for supplier selection and for spotting a model swap. The honest write-up keeps two numbers and two interfaces: "the model we call scored X on this published suite at its bare API; we separately ran Y items against our own endpoint and got Z." Merging them into one product claim is the mistake interviewers are listening for.
go deeper
Says the benchmark hit the model directly and the product adds a system prompt, tools and retrieval on top, so the number is about the model, not the app.
Explains that the harness's target was a plain model call and names concretely what the wrapper changes in the context and in the output path.
Argues the gap runs both ways so the published figure is not a bound, and describes standing up a target against the shipped endpoint with a slice, a budget and its own denominator.
Sets the organisational rule: vendor scores are supplier-selection and regression signals, product claims come only from runs at the shipped boundary, and reports never merge the two.
**What the published number physically measured.** A red-team benchmark is two parts. The *item corpus* is a list of prompt strings, each tagged with the behaviour it is trying to elicit — advbench, harmbench and jailbreakbench all distribute theirs as flat files. The *harness* walks that corpus, and at its centre sits a **target adapter**: harmbench names one in its run configuration, jailbreakbench exposes one as an LLM object. Its entire contract is "take one string, return one string". Physically that is a single HTTPS request to a chat-completions endpoint carrying the item as the only message in the array, a pinned temperature, and some fixed number of samples per item. The reply goes to a **judge** — harmbench ships a fine-tuned harm classifier, jailbreakbench specifies a judge model plus a published rubric — which labels the completion as exhibiting the behaviour or not. Attack-success rate (ASR) is judged hits over items attempted. Three properties fall straight out of that loop, and together they are the whole answer. The context held nothing but the item. The output path was raw prose. And nothing the model said could take an action. **What your product is instead.** | in the published run | in your deployment | |---|---| | empty context | a system prompt: role, scope, refusal policy, tenant identifiers, tool descriptions — commonly 500-2,000 tokens | | one user turn | user turn plus retrieved chunks plus tool results, all in one window | | text out | text out **and** tool calls your orchestrator actually executes | | no pre/post filter | input classifier, output classifier or redaction, length caps, response schema | | single turn | a session with history | Every row changes the token sequence the model conditions on; the third and fourth change what a compliant response is worth. **What an honest replay costs.** Engineering first. The target adapter now has to authenticate as a real session, send a turn, capture a streamed response to completion, and capture tool calls with their arguments. That is one to three engineer-days, and it breaks again whenever the response envelope changes. Then the run: a 400-item slice at 3 samples is 1,200 application calls plus 1,200 judge calls. The item text is identical to the published run, but every call now carries the system prompt and the retrieved chunks, so input tokens per call are routinely 20-40x the published run's — the token bill and the latency scale with that, not with the item. Your own endpoint is usually rate-limited per tenant far below the raw model API, so a sweep that took twenty minutes against the model API takes hours against the product. Human triage then dominates everything else, because someone still reads every disagreement between judge and expectation. **Where the number misleads.** The wrong reading to name explicitly is: *"the model scored 4% ASR, so our product is at most 4%."* That treats a component measurement as an upper bound on a system. It is not a bound in either direction. Downward, your system prompt gives the assistant an unrelated job and your input classifier rejects items before the model ever sees them, so items that landed on the bare model die in your stack. Upward, retrieval and tool results are a second inbound text channel with no counterpart in the published run, a tool call turns text-level compliance into a real side effect, and the system prompt itself becomes something worth extracting. Both pressures are live at once, so the published figure bounds nothing. Two further misreadings follow from the same confusion. Treating *published minus replayed* as "what our safety layers bought" attributes a delta to one cause when the interface, the context length, the output shape, the available actions and possibly the judge's behaviour all changed together. And quietly reusing the published denominator — the suite's item count — for a run in which you attempted a slice, dropped rate-limited items and lost some to truncation produces a rate computed over a set nobody defined. **What you check before believing your own number.** - The adapter captures the *whole* envelope: user-visible text, tool calls with arguments, and the stage at which the item terminated. - The judge still agrees with humans on your response format — hand-label a stratified sample and report the agreement rate next to the score. - Per-item termination is logged, so you can tell "the model refused" from "the filter caught it". - HTTP reality is logged: 429s, retries, timeouts and truncations belong in a stated denominator, not silently dropped. - The sandbox tools are faithful enough that an action-level hit means something in production. **How you report it.** Two numbers, two interfaces, never merged: "the model we call scored X on this published suite at its bare API; we separately attempted Y items of that suite against our own endpoint and observed Z, with judge agreement of A on a hand-labelled sample." Anything shorter is a component claim wearing a product label.
- Name one thing your deployment adds that could make the published number too optimistic, and one that could make it too pessimistic.Too optimistic: tools and retrieval add inbound channels and real side effects the suite never touched. Too pessimistic: a narrow system prompt plus an input classifier deflects items that landed on the bare model.
- What is the minimum you must change in a benchmark harness to point it at your application instead of the model?The target object - the thing the harness calls with an item and gets a response from. It has to hit your app's endpoint with real auth, real system prompt and real tools, and return the text and actions the user would actually get.
- Your app endpoint is rate-limited and metered while the model API is not. How does that change what you replay?You stop replaying the full corpus and pick a slice, then report that slice's own denominator. The budget becomes part of the method, and you say so rather than implying full-suite coverage.
The published score is a dynamometer reading for an engine on a test bench. Your product is the whole car built around it, and the wiring you added can both damp the vibration and start a fire the bench run could never have seen.
saying these in an interview costs you the question
- Treating the published score as a conservative upper bound on the deployed system's risk.
- Quoting a vendor leaderboard figure in a product security document as if it were measured on the product.
- Assuming a system prompt can only reduce attack success, never enable a new one.
- Claiming the app is safer because 'we use the same model that scored well', with no run at the shipped endpoint.