You lead red-teaming for a product built on a hosted model that already carries published red-team benchmark scores. How do you split a fixed engagement between replaying those published items at your own application boundary and authoring cases only your application can fail - and what do you tell leadership the vendor's published number is still worth?
answer
- hours go to authored app cases
- thin frozen replay slice
- tripwire for substrate change
- component evidence vs product claim
- name the interface in every claim
basics
~20 sSpend most of the engagement on cases only your application can fail - its system prompt, tools, retrieval and tenant data - and reserve a thin, repeatable slice of published items as a calibration tripwire. Tell leadership the vendor number is a supplier-selection and regression signal about a component, never a claim about the product you ship.
solid answer
~60 sThe published corpus has already been run by the vendor, and its items were public long enough to be optimised against, so re-running all of it at your boundary buys little per hour spent. What a small replayed slice does buy is comparability: a fixed, cheap set you re-run on every model swap or config change, whose movement tells you the substrate changed under you. The bulk of the hours go where nothing published exists - the specific system prompt, the specific tool set and its blast radius, the retrieval corpus and its trust boundary, and whatever tenant or role separation the product promises. Those are the failures no leaderboard can have covered, and they are also the ones that map to a real consequence. To leadership, the framing is a supply chain one: the vendor score is evidence about a component you buy, useful for choosing between suppliers and for noticing regressions. Product assurance comes only from runs at the shipped boundary, and any external document must quote the second, never the first.
go deeper
Says most effort should go to testing the actual product rather than re-running someone else's benchmark.
Justifies the split by yield - published items are already run and public, application surfaces are untested - and keeps a small replay for comparison.
Sizes the frozen slice against a metered endpoint and cadence, and separates harness cost from triage cost when planning hours.
Turns it into policy: component evidence versus product claim, every external number names its interface, and the replay slice is frozen so it stays a tripwire.
**The allocation argument, in the units the budget is actually spent in.** An engagement's hours go to three places: harness work, machine time, and human triage. Machine time is the smallest of the three at this scale — a few thousand calls is a rounding error against a security engineer's week — so any split argued on token cost is arguing about the wrong line item. Replaying published items is *cheap on harness* (the corpus and a judge already exist), *cheap on triage* (reference labels exist, so disagreements are few and quick), and **low yield**: the vendor already ran it, the items have been public long enough to have been optimised against and to have leaked into safety tuning and guard training data, and — the point of this leaf — the corpus says nothing about the wrapper you actually ship. Authoring application-specific cases is *expensive on harness* (each case needs a fixture, a trigger and an oracle) and *expensive on triage* (no reference labels; a human decides), but it is **high yield**, because each finding concerns a surface nobody outside your company has tested and each maps to a consequence you can price. A defensible split therefore leans heavily toward authored work, with a small fixed replay slice that survives from engagement to engagement. **Why keep any replay at all.** Three concrete reasons, and none of them is coverage. 1. **Calibration.** A corpus with known behaviour is how you find out that your target adapter drops tool calls or that your judge cannot read your response format. You want to discover that on items whose answers are known, not on your own novel cases. 2. **Tripwire.** A frozen, cheap slice re-run every release turns a silent substrate change — a model version swap, an endpoint routing change, a system-prompt edit by another team — into a visible movement on a set that did not otherwise change. 3. **A common language with the vendor.** When you escalate, a shared published corpus is the artefact they can reproduce. **Sizing the slice, with the numbers stated.** Small enough to re-run on a metered, production-like endpoint at your release cadence: on the order of 100 items at 3 samples is 300 application calls and 300 judge calls, minutes of machine time, and — because the labels are known — under an hour of triage in a normal week. Stratify it across the categories that matter to your product, then **freeze it**: pin the item list with a manifest hash, and record the model version, the system-prompt version and the filter configuration alongside each result. **What you tell leadership, precisely.** | number | what it is evidence for | where it may appear | |---|---|---| | the vendor's published score | a purchased component, measured at the bare model API | supplier comparison, regression alerting | | your frozen replay slice | comparability and change detection | internal release reports | | your authored application cases | product risk at the shipped boundary | external assurance, with denominators and judge agreement | **Where the number misleads, and what it costs when it does.** The characteristic organisational failure is substitution: a good vendor score is free, quotable and arrives with a chart, so it drifts into a customer-facing document and quietly discharges the obligation to test the product. The cost of that is not a bad number, it is an untested surface. Three more traps. **Slice growth** — engineers add interesting items run over run until movement in the figure reflects the changing item set rather than the substrate, and the tripwire has become a random number generator. **Triage treated as free** in planning, which is how engagements end with a large unread queue of tool hits and no triaged findings. And **the unlabelled interface**: any safety figure quoted without naming the interface it was measured at invites the reader to assume it was the product. **What you check.** That every safety number in an external document names its interface. That the replay slice's manifest hash is unchanged since the last run, and that model version, prompt version and filter config are recorded with each result. That authored-case hours actually exceeded replay hours, rather than being planned that way and then inverted because the replay was easier to schedule. And that reported product claims carry their own denominator and their judge's agreement rate — because once that rule exists, the budget argument settles itself: the only numbers anyone may quote are the ones you have to pay to produce.
- Why keep any published items in the engagement at all if the yield is low?They calibrate your harness and judge against known behaviour and act as a tripwire: unexplained movement on a frozen slice signals that the model or endpoint changed under you.
- What written rule keeps a vendor score from leaking into product assurance?Every safety number in an external document must name the interface it was measured at, and only numbers measured at the shipped boundary may be quoted as product claims.
Keep the replay slice frozen the way a lab keeps a control specimen: it is useful precisely because it never changes, so when its reading moves you know the thing being measured moved and not the ruler.
saying these in an interview costs you the question
- Spending the engagement reproducing a public corpus because it is easy to schedule.
- Letting a vendor leaderboard score appear in a customer-facing assurance document.
- Growing the replay slice run by run until it is no longer comparable.
- Treating triage effort as free when planning how much to run.