skip to content

A public safety-benchmark leaderboard shows an attack-success rate measured some months ago against a hosted chat endpoint. What does that row fail to record about the system it scored, and why does that limit citing the number today?

level: middleimportance: must knowfreq 58%

answer

  1. alias is a label, not a build
  2. snapshot of a pipeline, not a model
  3. judge has versions too
  4. decoding settings move the rate
  5. shortlist signal, not a control

basics

~20 s

It records a name, not a build. A hosted endpoint's alias can be re-pointed to an updated model, and the safety tuning and any provider-side filters behind it can change with no announcement. The row also freezes the decoding settings and the harm judge used then. It is one snapshot, not today's endpoint.

solid answer

~50 s

A leaderboard row is a **measurement of a moment**, and almost everything that made the measurement reproducible sits outside the row. What is usually missing: - **Which build answered.** A product alias is a routing label; the weights, the serving stack and the system-side filtering behind it can all be replaced without a version bump you can see. - **How it was sampled.** Temperature, top-p, max tokens, retries and refusal handling all move an attack-success rate. - **Who decided a hit.** The benchmark's harm judge is itself a model or classifier that gets retrained; the same transcripts re-judged later score differently. - **Which subset ran.** Behaviour lists and prompt sets get revised between dataset revisions. So the honest reading is: *on that date, that harness, judged that way, that alias produced this rate.* Directional signal for a shortlist, not a control you can claim for the endpoint you would call today. If the number matters to a decision, re-run it yourself and record the manifest.

go deeper

for a junior

Says the number is old and the model may have been updated since, and that they would re-run rather than cite it.

for a middle

Names the specific things a row omits: served build behind the alias, decoding settings, harm judge identity, dataset revision.

for a senior

Separates model change from provider-side filtering and from judge change, and explains why an attack-success rate cannot distinguish them on its own.

for a principal

Frames it as a policy question: what published numbers may be cited for, what must be reproduced in-house, and what evidence a vendor must supply.

### What the row actually asserts Read literally, a leaderboard row says "this model is N% jailbreakable". What it can honestly assert is much narrower: *on this date, this harness, driving this endpoint name, with this harm judge, over this behaviour list, produced this attack-success rate.* Attack-success rate (ASR) is a fraction. The numerator is the attempts some automated judge marked as harmful compliance; the denominator is the attempts the harness decided to count. Both halves are chosen by whoever ran the suite, not by the model. That is why the row is a measurement of a pipeline, not a property of weights. ### The four layers that move underneath, none of them named in the row **1. The served build.** A hosted chat endpoint is addressed by a product-level name — a routing label, not a build identifier. Behind it a provider can replace the weights, add or retune safety post-training, and enable or tighten serving-side filtering (an input or output classifier that sits in front of the model and can block a request before the model ever sees it). None of those require the name to change, and an aggregate ASR cannot tell "the model now refuses" from "a filter caught the prompt upstream" — both push the number the same way. **2. The harness.** Prompt templating, whether a system prompt was attached at all, temperature and top-p, `max_tokens`, attempts per behaviour, retry-on-error policy, and how truncated or empty responses are counted. Also the success definition: "a behaviour counts as a success if any of five attempts succeeded" and "the fraction of individual attempts that succeeded" are computed from identical transcripts and can differ by a factor of several. **3. The judge.** Public harmful-behaviour benchmarks score attempts automatically. HarmBench ships its own fine-tuned classifier as the judge; JailbreakBench publishes its jailbreak artifacts precisely so that others can re-score them. A judge is a model with revisions, a threshold, and its own false-positive profile. Re-scoring the same transcripts under a stricter judge months later moves the rate with the endpoint untouched. **4. The dataset revision.** Behaviour sets get corrected, extended and pruned. A rate over a different denominator is not a comparison. Note that "coverage" here means behaviours present in the benchmark, not attack surface reached in your application. ### What it cost to produce, and why the cost changes your reading Do the arithmetic on the run you are being asked to trust. Two hundred behaviours at five attempts each is 1,000 completions from the target endpoint plus roughly 1,000 judge calls. On a mid-priced hosted endpoint that is tens of dollars of spend and, at modest concurrency with rate limits, an hour or two of wall clock. The expensive part is human: triaging the flagged transcripts to see which "successes" are real is a half-day to a day of an analyst's time, and that cost does not shrink when you re-run. That arithmetic is the operative point. Refreshing a benchmark number is *cheap* relative to the meeting spent arguing about whether last quarter's row still holds. If the number matters enough to cite in a decision, it is almost always cheaper to reproduce it than to defend the stale one. ### Where the number misleads - **Attribution.** A low ASR is read as a safe model. It may be a provider-side filter, a mis-templated prompt the endpoint refused for the wrong reason, or a lenient judge that never fired. - **Cross-row comparison.** Two rows for the same endpoint from different groups are usually two different experiments. Differences in harness and judge routinely exceed differences between models. - **Direction bias.** A number that fell is accepted; a number that rose gets investigated. An unexplained improvement is just as often a harness bug. - **Silent denominators.** Dropped timeouts and unrecorded attempts-per-behaviour make two rates look like drift when only the bookkeeping changed. ### What you would check before citing it Measurement date; the exact endpoint string and any served-revision identifier the provider returned; judge identity, revision and threshold; dataset revision and the behaviour count; attempts per behaviour and the success definition; and whether transcripts were published. Transcripts are the one artefact that partially rescues an old run — with them you can re-judge under your own scorer and separate judge drift from everything else. Without them, the number and that judge can never be prised apart, and the row is shortlisting signal at best.

  • The endpoint provider says nothing changed. Does that settle it?
    No. Absence of a changelog entry is not evidence of no change, and provider-side filtering often sits outside the model changelog entirely. Re-running is cheaper than arguing.
  • Two leaderboard rows disagree wildly on the same endpoint. What is the most likely cause?
    Different harnesses and different harm judges, not different models. Compare templating, decoding settings, judge identity and the behaviour subset before concluding anything about the endpoint.
  • What is the one artefact that makes an old run partially rescuable?
    Published transcripts. With the raw attempts and responses you can re-judge them under your own scorer and separate judge drift from model drift, though you still cannot recover the model as it was.

It is like quoting a restaurant's hygiene score from two years ago: the kitchen may have a new chef, and the grade was awarded by an inspector whose standards have since been rewritten. Nothing on the certificate tells you which of the two changed.

saying these in an interview costs you the question

  • Treats the leaderboard rate as a current property of the model.
  • Assumes the same endpoint name means the same weights months later.
  • Ignores that the harm judge is itself a versioned component.
  • Compares two rows from different harnesses as if they were the same experiment.
  • Cites a published number to sign off a deployment without reproducing it.

context