Where does a broad multi-dimension trust battery such as TrustLLM belong in an AI red-team program that also runs adaptive attack tooling and product-specific scenarios, and what claim should its dimension scores never be used to support?
answer
- three tiers, three evidentiary weights
- triage and tripwire only
- never an assurance claim
- score becomes a target, refusals rise
- keep private content untunable
basics
~20 sUse it as a cheap, repeatable tripwire and as a way to pick where to attack next. Run it early, then on a schedule. Never let a dimension score stand as an assurance claim that the system is safe: it is a fixed, thin, non-adaptive prompt set that knows nothing about your product.
solid answer
~50 sGive it two jobs and refuse it a third. **Job one: triage.** Run it early to find which trust properties look weakest, so the expensive adaptive work goes where it pays. **Job two: regression tripwire.** Run it on a schedule and on major model or prompt changes; because the content is fixed, a large move is worth investigating even though a small move is not. **The job it must never have: assurance.** A dimension score cannot support "the assistant is safe on this dimension", because the prompt set is thin, static and authored for a generic model rather than your product, and because a real attacker adapts while the battery does not. Assurance has to come from scenarios written against your own threat model and from adaptive attempts by people. At program level that means budget and reporting rules: the battery is the cheapest tier and gates nothing on its own; a release decision cites product-specific evaluations, with battery trend attached as context.
go deeper
Says the battery is a starting point and that deeper testing is needed before trusting a model.
Separates triage and regression-tripwire use from assurance use, and can say why thin, fixed prompt sets do not support the latter.
Designs the tiering: battery output drives where adaptive attack effort goes, product scenarios gate the release, and every reported figure carries its item count and noise floor.
Owns the incentive problem — the easiest number becomes the target — and answers with reporting rules, budget allocation across tiers, and privately held evaluation content.
**The portfolio view.** A mature program runs three tiers that differ in cost, in what they can be adapted to, and — the part that gets forgotten — in evidentiary weight. | tier | what it is | rough cost per cycle | what it may support | |---|---|---|---| | broad battery (TrustLLM-style) | fixed prompts, several scored dimensions | tens of dollars of tokens, hours of wall clock, ~zero human time once scheduled | orientation; large-regression alarm | | focused suite / adaptive tooling | one property attacked hard; HarmBench- or JailbreakBench-style behaviour sets, plus a person or a search loop that iterates | comparable tokens but days of skilled attention per campaign | a claim about a specific failure mode | | product scenarios | your threat model, your data flows, your consequences | the largest human cost; not comparable across companies | a release decision | **Where the battery belongs.** Two jobs, and it must be refused a third. *Triage*: run it early so the expensive adaptive work goes where the profile is worst. *Regression tripwire*: run it on a schedule and on major model or prompt changes, where the fixed content is an asset — the same content re-run is at least self-consistent, so a large move is worth a look even though a small one is not. **The job it must never have: assurance.** Four independent reasons, each sufficient on its own. Its items are fixed and single-shot, so no score speaks to an attacker who iterates against the refusal he just received. Its per-dimension item counts are thin once you divide a shared budget across six dimensions and then across sub-tasks, so a few points is inside the instrument's own jitter. Its dimension definitions and its scorers were authored for models in general — if the harm your product actually owns is not one of those dimensions, it is scored nowhere and reads as absence of a problem. And public fixed content decays: it is optimised against and, over time, absorbed into training data, which lifts scores without changing behaviour. **The organisational failure to design against.** Battery numbers are the cheapest to produce, which is exactly why they migrate into slide decks and then into commitments. Once a dimension score is a target, teams tune to it, and the cheapest way to move most trust dimensions is to refuse more — which raises the number, degrades the product, and leaves exploitability where it was. The fix is structural rather than exhortative: - state in the reporting template which tier each figure came from and which decisions that tier may support; - never publish a dimension score without its item count and its measured run-to-run spread; - hold back a private scenario set the product teams never see, refreshed periodically, so at least one set cannot be tuned against; - gate releases on tier three only, with tier-one trend attached as context. **Where the number misleads at program level.** Two specific misreadings. First, *trend over time on public content is confounded with contamination*: a battery line that improves quarter over quarter may be tracking the corpus leaking into training sets rather than your program working, and you cannot distinguish those from the score alone — which is precisely what the private set is for. Second, *a green profile is read as coverage*: dimensions the battery does not have simply do not appear, so a model can be uniformly strong across every column and still be wide open on the risk your product owns. A battery has no way to signal its own blind spots, and an all-green report is the shape in which that silence is most convincing. **How I would size the spend.** The battery should be a scheduled job costing a small standing token budget and effectively no recurring human attention — if a person is babysitting it, the instruments are inverted. The majority of the human budget goes to tiers two and three, allocated using the battery's triage output and the standing threat model. A useful sanity check on any program: ask what fraction of red-team engineer-days last quarter went into re-running fixed content. If it is not close to nothing, the program is buying breadth it has already paid for and starving the only two tiers that can substantiate a claim or stop a release.
- Leadership wants a single trust number tracked quarterly. What do you offer instead?A small dashboard: battery dimension trends with their noise floors as context, plus the product-scenario pass rate and open red-team findings as the figures that actually carry weight. One number invites tuning to it.
- What is the cheapest structural defence against evaluation content being tuned against?Hold back a private scenario set the product teams never see, refreshed periodically. Public battery content can be optimised against; content nobody can train on stays informative.
saying these in an interview costs you the question
- Lets a battery dimension score gate a release.
- Treats a broad battery run as a substitute for a threat model.
- Has no answer for what happens when teams optimise against the published dimension scores.
- Spends most of the human red-team budget re-running fixed content.
- Reports battery numbers without item counts or run-to-run spread.