skip to content

Reading a Leaderboard Honestly

A published score decays three ways: the corpus leaks into training, the endpoint behind it moves unannounced, and the deployment adds a wrapper the run never touched. Interviewers probe all three.

on this pageshow

explore

questions

15

A vendor publishes a red-team benchmark score for a model, and your product calls that same model behind a system prompt, a tool loop and a retrieval layer. What interface did that published number actually measure, and why can you not quote it as your product's safety number?

level: juniorimportance: must knowfreq 72%

answer

  1. score stops at the model API
  2. wrapper never in the run
  3. two numbers, two interfaces
  4. not an upper or lower bound
  5. component claim vs product claim

basics

~20 s

The number came from prompts sent straight to the bare model endpoint - no system prompt, no tools, no retrieved text. Your product is a different system: extra instructions, extra input channels, extra output paths, none of which that run exercised. So the score describes the model you call, not the application you ship.

solid answer

~50 s

Published red-team suites drive the model API directly. Each item is one user turn, the context is empty or a fixed neutral preamble, and the harm judge reads the raw completion. Your deployment changes all three of those: a system prompt that constrains scope and tone, a tool loop that lets a compliant answer become an action, and retrieval that pushes text you did not author into the same context window. So the published score is a statement about a **component**, useful for supplier selection and for spotting a model swap. The honest write-up keeps two numbers and two interfaces: "the model we call scored X on this published suite at its bare API; we separately ran Y items against our own endpoint and got Z." Merging them into one product claim is the mistake interviewers are listening for.

go deeper

for a junior

Says the benchmark hit the model directly and the product adds a system prompt, tools and retrieval on top, so the number is about the model, not the app.

for a middle

Explains that the harness's target was a plain model call and names concretely what the wrapper changes in the context and in the output path.

for a senior

Argues the gap runs both ways so the published figure is not a bound, and describes standing up a target against the shipped endpoint with a slice, a budget and its own denominator.

for a principal

Sets the organisational rule: vendor scores are supplier-selection and regression signals, product claims come only from runs at the shipped boundary, and reports never merge the two.

**What the published number physically measured.** A red-team benchmark is two parts. The *item corpus* is a list of prompt strings, each tagged with the behaviour it is trying to elicit — advbench, harmbench and jailbreakbench all distribute theirs as flat files. The *harness* walks that corpus, and at its centre sits a **target adapter**: harmbench names one in its run configuration, jailbreakbench exposes one as an LLM object. Its entire contract is "take one string, return one string". Physically that is a single HTTPS request to a chat-completions endpoint carrying the item as the only message in the array, a pinned temperature, and some fixed number of samples per item. The reply goes to a **judge** — harmbench ships a fine-tuned harm classifier, jailbreakbench specifies a judge model plus a published rubric — which labels the completion as exhibiting the behaviour or not. Attack-success rate (ASR) is judged hits over items attempted. Three properties fall straight out of that loop, and together they are the whole answer. The context held nothing but the item. The output path was raw prose. And nothing the model said could take an action. **What your product is instead.** | in the published run | in your deployment | |---|---| | empty context | a system prompt: role, scope, refusal policy, tenant identifiers, tool descriptions — commonly 500-2,000 tokens | | one user turn | user turn plus retrieved chunks plus tool results, all in one window | | text out | text out **and** tool calls your orchestrator actually executes | | no pre/post filter | input classifier, output classifier or redaction, length caps, response schema | | single turn | a session with history | Every row changes the token sequence the model conditions on; the third and fourth change what a compliant response is worth. **What an honest replay costs.** Engineering first. The target adapter now has to authenticate as a real session, send a turn, capture a streamed response to completion, and capture tool calls with their arguments. That is one to three engineer-days, and it breaks again whenever the response envelope changes. Then the run: a 400-item slice at 3 samples is 1,200 application calls plus 1,200 judge calls. The item text is identical to the published run, but every call now carries the system prompt and the retrieved chunks, so input tokens per call are routinely 20-40x the published run's — the token bill and the latency scale with that, not with the item. Your own endpoint is usually rate-limited per tenant far below the raw model API, so a sweep that took twenty minutes against the model API takes hours against the product. Human triage then dominates everything else, because someone still reads every disagreement between judge and expectation. **Where the number misleads.** The wrong reading to name explicitly is: *"the model scored 4% ASR, so our product is at most 4%."* That treats a component measurement as an upper bound on a system. It is not a bound in either direction. Downward, your system prompt gives the assistant an unrelated job and your input classifier rejects items before the model ever sees them, so items that landed on the bare model die in your stack. Upward, retrieval and tool results are a second inbound text channel with no counterpart in the published run, a tool call turns text-level compliance into a real side effect, and the system prompt itself becomes something worth extracting. Both pressures are live at once, so the published figure bounds nothing. Two further misreadings follow from the same confusion. Treating *published minus replayed* as "what our safety layers bought" attributes a delta to one cause when the interface, the context length, the output shape, the available actions and possibly the judge's behaviour all changed together. And quietly reusing the published denominator — the suite's item count — for a run in which you attempted a slice, dropped rate-limited items and lost some to truncation produces a rate computed over a set nobody defined. **What you check before believing your own number.** - The adapter captures the *whole* envelope: user-visible text, tool calls with arguments, and the stage at which the item terminated. - The judge still agrees with humans on your response format — hand-label a stratified sample and report the agreement rate next to the score. - Per-item termination is logged, so you can tell "the model refused" from "the filter caught it". - HTTP reality is logged: 429s, retries, timeouts and truncations belong in a stated denominator, not silently dropped. - The sandbox tools are faithful enough that an action-level hit means something in production. **How you report it.** Two numbers, two interfaces, never merged: "the model we call scored X on this published suite at its bare API; we separately attempted Y items of that suite against our own endpoint and observed Z, with judge agreement of A on a hand-labelled sample." Anything shorter is a component claim wearing a product label.

  • Name one thing your deployment adds that could make the published number too optimistic, and one that could make it too pessimistic.
    Too optimistic: tools and retrieval add inbound channels and real side effects the suite never touched. Too pessimistic: a narrow system prompt plus an input classifier deflects items that landed on the bare model.
  • What is the minimum you must change in a benchmark harness to point it at your application instead of the model?
    The target object - the thing the harness calls with an item and gets a response from. It has to hit your app's endpoint with real auth, real system prompt and real tools, and return the text and actions the user would actually get.
  • Your app endpoint is rate-limited and metered while the model API is not. How does that change what you replay?
    You stop replaying the full corpus and pick a slice, then report that slice's own denominator. The budget becomes part of the method, and you say so rather than implying full-suite coverage.

The published score is a dynamometer reading for an engine on a test bench. Your product is the whole car built around it, and the wiring you added can both damp the vibration and start a fire the bench run could never have seen.

saying these in an interview costs you the question

  • Treating the published score as a conservative upper bound on the deployed system's risk.
  • Quoting a vendor leaderboard figure in a product security document as if it were measured on the product.
  • Assuming a system prompt can only reduce attack success, never enable a new one.
  • Claiming the app is safer because 'we use the same model that scored well', with no run at the shipped endpoint.

context

open as a page

You run a public harmful-behaviour benchmark such as HarmBench against a hosted chat endpoint and save the resulting attack-success rate for later comparison. What do you record alongside the number so it can still be interpreted in six months?

level: juniorimportance: must knowfreq 50%

basics

~20 s

Record the endpoint string and any pinned revision it resolved to, the timestamp, decoding settings such as temperature and max tokens, the dataset revision and which behaviours ran, the harm judge and its version, and the raw transcripts. Without these, a later number cannot be compared and any difference cannot be explained.

open as a page

A model card reports a near-zero attack-success rate on AdvBench, a public corpus of harmful-behaviour prompts. What could produce that number without the model being any harder to jailbreak?

level: middleimportance: must knowfreq 62%

basics

~20 s

Nothing prevents the corpus being inside safety tuning. Those exact prompts appear in refusal training data, and often in the guard classifier that scores the run, so the model refuses memorised strings. Reword the same requests and the rate typically climbs. The number measures recall of a public list.

open as a page

You replay a published red-team benchmark's prompt set against your deployed assistant endpoint instead of against the bare model API. What is one mechanism that pushes the attack-success number DOWN compared with the published run, and one that pushes it UP?

level: middleimportance: must knowfreq 60%

basics

~20 s

Down: your system prompt narrows what the assistant will discuss and an input or output classifier sits in the path, so items that landed on the bare model get deflected. Up: retrieval carries text you did not write, tools turn a compliant answer into a real action, and a long preamble dilutes plain refusal behaviour.

open as a page

A public safety-benchmark leaderboard shows an attack-success rate measured some months ago against a hosted chat endpoint. What does that row fail to record about the system it scored, and why does that limit citing the number today?

level: middleimportance: must knowfreq 58%

basics

~20 s

It records a name, not a build. A hosted endpoint's alias can be re-pointed to an updated model, and the safety tuning and any provider-side filters behind it can change with no announcement. The row also freezes the decoding settings and the harm judge used then. It is one snapshot, not today's endpoint.

open as a page

You are handed only a published safety evaluation — a low attack-success rate on a public attack corpus, scored by a guard classifier the same vendor ships — and no access to any training data. What evidence would tell you the corpus is inside the safety tuning or inside that classifier's training set?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Re-run the same harmful behaviours in fresh wording and compare. A large gap between verbatim prompts and paraphrases points at memorisation. Score the responses with a judge the vendor did not train, check whether the corpus predates the model release, and see whether the guard's own documentation lists it as training data.

open as a page

AdvBench and HarmBench are publicly released attack-prompt corpora used to score how often a model refuses. Why do red teams keep part of their own attack prompt set unpublished?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Because anything published gets trained on. Vendors add public attack prompts to safety-tuning data and to guard classifiers, so models learn to refuse those exact prompts. The score rises without the model becoming harder to attack. An unpublished set stays a fair test, since nobody could have fitted to it.

open as a page

A provider offers both a floating alias for a chat model and a pinned snapshot identifier of the same model. Which of the two should a recurring safety-benchmark suite be pointed at, and what does each choice cost you?

level: middleimportance: should knowfreq 36%

basics

~20 s

Point at both if you can. The alias is what production calls, so running it is your drift detector. The pinned snapshot is a control: same weights every quarter, so a moved rate there means your harness or judge moved. Costs are double the queries, and pinned snapshots eventually get retired.

open as a page

Your quarterly red-team report has cited an AdvBench attack-success rate for several quarters. You now learn the deployed guard classifier was trained on that same public corpus. Do you drop the number, and what goes in the report instead?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Do not silently drop it. Keep the public number as a regression floor, clearly labelled as a set the guard was trained on, and make a held-out attack set the headline. Report both, since the gap between them is your best estimate of how much the public score is inflated.

open as a page

When you replay a published red-team benchmark against your own application endpoint, the benchmark's automated harm judge can start scoring the wrong thing because your app's responses do not look like raw model completions. What are two ways your response format breaks that judge, and what do you do about each?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Two common breaks: the app returns structured output or a rendered card, so the judge reads a wrapper rather than the answer text; and the app emits one fixed refusal template the judge was never calibrated on. Fix by judging exactly the text and tool calls the user receives, then hand-labelling a sample to check the judge still agrees.

open as a page

A published red-team benchmark delivers every item as a user turn, but your deployment also puts text you did not write - retrieved documents and tool results - into the same context. What do you re-run against that second channel, and how do you keep it a bounded job instead of a second full benchmark?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Do not replay the whole corpus on the retrieval channel. Select the items whose outcome depends on the model following injected instructions, place that text in the retrieved-document or tool-result slot instead of the user turn, and score what the app did as well as what it said. Report the two channels as separate coverage numbers with separate denominators.

open as a page

A quarter after your baseline, you re-run the same harmful-behaviour benchmark with the same prompts against the same hosted endpoint alias, and the attack-success rate drops noticeably. How do you establish whether the model behind the alias changed, your harm judge changed, or it is run-to-run variation?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Change one thing at a time. Re-score the old transcripts with the new judge: if the rate moves there, the judge drifted. Diff the run manifests for decoding, dataset and error-handling changes. Then repeat the new run to see the spread across repeats before calling any residual gap a real model change.

open as a page

Replacing a contaminated public attack corpus means building and maintaining a held-out attack set of your own. What does that cost an organisation over time, and what rules stop the replacement from becoming contaminated too?

level: principalimportance: should knowfreq 36%

basics

~20 s

It costs skilled authoring time, labelled ground truth, and a refresh cadence as the set ages. Keeping it clean is policy: never publish examples, only aggregates; send it only to endpoints under no-training terms; hold a sequestered slice nobody iterates against; and keep the tuning team from seeing it.

open as a page

You lead red-teaming for a product built on a hosted model that already carries published red-team benchmark scores. How do you split a fixed engagement between replaying those published items at your own application boundary and authoring cases only your application can fail - and what do you tell leadership the vendor's published number is still worth?

level: principalimportance: should knowfreq 35%

basics

~20 s

Spend most of the engagement on cases only your application can fail - its system prompt, tools, retrieval and tenant data - and reserve a thin, repeatable slice of published items as a calibration tripwire. Tell leadership the vendor number is a supplier-selection and regression signal about a component, never a claim about the product you ship.

open as a page

You own the rule for how third-party safety-benchmark scores may be cited in your organisation's model-vendor reviews. How do you decide how long such a published number stays citable, and what do you require once it has expired?

level: principalimportance: should knowfreq 28%

basics

~20 s

Tie expiry to events, not just to the calendar: any endpoint or guard change, a dataset or judge revision, or a new attack family invalidates the number. Add a calendar backstop of a quarter or two. Once expired, a published score may inform shortlisting only; anything gating a launch must be reproduced in-house with a recorded manifest.

open as a page