skip to content

You are choosing an off-the-shelf evaluation to open a red-team engagement against a chat assistant. What does a broad multi-dimension trust battery such as TrustLLM buy you compared with a focused single-purpose attack suite, and what does that breadth cost?

level: juniorimportance: should knowfreq 45%

answer

  1. survey instrument vs drill
  2. breadth divides the prompt budget
  3. one number per dimension
  4. fixed prompts, non-adaptive
  5. triage targets, don't make claims

basics

~20 s

A broad trust battery scores several different trust properties in one pass, so it gives you a wide first map of where a model looks weak. A focused attack suite drills one property hard instead. The cost is depth: each dimension is backed by a thin slice of prompts, so its number is coarse.

solid answer

~50 s

Think of it as a survey instrument versus a drill. A broad battery such as TrustLLM runs a set of separately-scored trust dimensions in a single sweep and hands back one number per dimension. That is genuinely useful at the start of an engagement: it is cheap, it is fast, and it tells you which areas deserve real attack work. What you pay for that is resolution. Total prompt budget is split across every dimension, so any one dimension rests on a small prompt set and a single scoring pass. That is enough to notice a model that is badly broken somewhere, and not enough to characterise how it breaks, under which attack strategy, or at what strength. It is also static content: it measures fixed prompts, not an adaptive attacker. The practical pattern is to use the battery to *choose targets* and a focused suite or a live attack tool to *make claims*.

go deeper

for a junior

Says the battery covers several trust properties at once while a focused suite goes deep on one, and that the broad numbers are rough.

for a middle

Explains that the prompt budget is split across dimensions, so each dimension rests on few items, and that battery content is static rather than adaptive.

for a senior

Adds the operational split — battery to prioritise, focused work to substantiate — and names what to inspect: items per dimension, the judging method, refusal handling, run-to-run stability.

for a principal

Frames it as a portfolio decision about where evaluation budget goes, and is explicit that neither instrument covers product-specific risk, which the team must author itself.

The two instruments answer different questions, and picking the wrong one wastes the expensive part of an engagement. **What each thing is.** A *broad trust battery* is a fixed corpus of prompts grouped into named **dimensions** — TrustLLM's are truthfulness, safety, fairness, robustness, privacy and machine ethics. Each dimension is not one test; it is a stack of sub-tasks assembled from existing datasets, and each sub-task carries its own **scorer**: sometimes exact match against a labelled answer, sometimes a refusal detector (a rule or small classifier deciding "did the model decline?"), sometimes another model asked to judge the response against a rubric. The dimension figure you see is an aggregate over those heterogeneous sub-tasks. A *focused attack suite* — HarmBench, AdvBench, JailbreakBench — fixes one question instead: given a list of target behaviours and one or more attack methods, what fraction of behaviours were elicited? That fraction is the **attack success rate (ASR)**, and its denominator is behaviours attempted, not prompts sent. **What the battery buys.** Orientation. One pass gives you a profile across several axes, computed the same way for every model you point it at, so it is usable for two jobs: deciding where to spend attack effort, and re-running later as a regression tripwire on content that cannot argue back because it never changes. **What a pass costs.** A full battery is tens of thousands of generations — order 10^4, spread across the sub-tasks. At a few hundred tokens in and a few hundred out per item, that is roughly ten to twenty million tokens per model per pass, and the model-judged sub-tasks add a second call for every item they cover, so budget close to double on those. On a mid-priced metered endpoint that lands in the tens of dollars per model; at modest concurrency it is single-digit hours of wall clock. The real first-run cost is engineer time — writing the adapter to your endpoint, wiring judge credentials, getting item-level output to persist — call it a day, and near zero afterwards because the job is schedulable. Compare that to an adaptive campaign, where a person iterates against refusals: the same token spend buys you one dimension's worth of ground and costs days of attention. That asymmetry is exactly why batteries get over-used. **Where the number misleads.** Four readings are wrong, and each is common. - *The aggregate hides its denominators.* "Safety: 0.83" is a mean over sub-tasks with very different item counts and difficulty. One saturated sub-task with many items can carry a dimension while the sub-task you actually care about sits at chance. - *Refusal is credited as correct.* On most safety-flavoured sub-tasks, declining is the scored-right answer. A model tuned to refuse aggressively climbs that column and sinks on the truthfulness or over-refusal side. Read one column alone and you will promote the least useful model in the set. - *Rank order is not a difference.* Batteries publish point estimates with no interval. Two models a couple of points apart on a thin dimension are not distinguishable by that run. - *Fixed public content ages.* Once a suite is public it is trained near, tuned against, and in some cases absorbed outright. The score rises while behaviour does not, and the battery quietly stops measuring the thing it was built for. Underneath all four: the prompts are static and single-shot. A battery item does not respond to the refusal it just got; an attacker does. So a good dimension score is evidence about fixed strings, never evidence of robustness. **What I would check before quoting a figure.** How many items sit behind the specific dimension I care about, and how they split across sub-tasks. Which scorer decided each sub-task, and whether it is a model — if it is, which model and which version. Whether refusals count as passes there. Whether the run persists per-item results at all; if it only emits aggregates, the number is unauditable. And the cheapest check of the four: run the same model twice unchanged and look at the spread, because that spread is the resolution limit on everything you will later claim. **The working split.** Battery to prioritise. Focused suite or an adaptive tool driven by a person to substantiate a claim about a specific failure mode. Your own product-scenario set to decide anything about shipping. Reversing that order — letting the cheapest instrument carry the heaviest claim — is the failure this choice exists to avoid.

  • You must report one number to a stakeholder from a broad trust battery run. What do you attach to it?
    The per-dimension breakdown, the item count behind each dimension, how responses were judged, and the fact that the content is fixed and non-adaptive. A bare aggregate hides all four.
  • When is a broad battery clearly the wrong tool?
    When you already know which risk you care about. Then the whole budget should go into that one area with an adaptive attack tool, not be spread across dimensions you are not going to act on.

saying these in an interview costs you the question

  • Treats a broad battery as a substitute for product-specific evaluation.
  • Cannot say why a per-dimension score is less trustworthy than an overall score built from the same prompts.
  • Believes the battery measures robustness against an adaptive attacker.
  • Compares two models on a battery without checking that dimension composition and judging were identical.

context