skip to content

Broad Trust Batteries

A battery like TrustLLM scores several trust dimensions in one run and buys that breadth with a thin slice of prompts behind each figure. Interviewers ask what a broad score can actually support.

on this pageshow

explore

questions

4

You re-run a broad multi-dimension trust battery such as TrustLLM after shipping a mitigation, and one dimension's score comes back a few points higher. Why is that not yet evidence the mitigation worked, and what do you check before claiming it?

level: middleimportance: must knowfreq 55%

answer

  1. few items = large per-item weight
  2. re-run baseline for a noise floor
  3. diff flipped items, not averages
  4. judge and sampling drift too
  5. refusal counted as safe

basics

~20 s

That dimension is backed by few prompts, so a few points can be one or two responses flipping. Sampling, non-zero-temperature decoding and the judging step all move it on their own. Check which individual items changed, re-run the unmitigated build unchanged, and see whether the movement is bigger than run-to-run drift.

solid answer

~60 s

A per-dimension figure in a broad battery is an average over a small item set, so its per-item weight is large: with a thin slice behind the dimension, a handful of responses changing verdict is enough to move the headline several points. Three sources do that without any help from your mitigation — sampling variation in generation, drift in whatever judged the responses, and any change in the model or serving stack between the two runs. So before claiming anything I want three things. **Item-level diffs**: which specific prompts flipped, and did they flip in the direction the mitigation would predict, or somewhere unrelated? **A control**: re-run the *pre-mitigation* build under identical settings and see how much the score moves with no change at all — that is your noise floor. **A mechanism**: can I point at why those specific items should have changed? Movement with no plausible mechanism is noise wearing a result's clothes. If the delta is inside the noise floor, the honest report is "no measurable change at this resolution".

go deeper

for a junior

Recognises the score is an average over a small set of prompts and that a small move might just be a couple of responses changing.

for a middle

Names the concrete drivers — sampling, the judging step, refusal accounting, changes in the serving stack — and asks for item-level results.

for a senior

Insists on a measured noise floor from repeated unchanged baseline runs, and on a mechanism linking the flipped items to the mitigation before reporting a win.

for a principal

Sets the team norm: no dimension delta is reported without its item count, its control spread, and the flipped-item list, so nobody downstream over-reads a single figure.

The failure mode is reading a low-resolution instrument as though it had high resolution. Everything follows from **per-item weight**: how much of the headline one prompt's verdict is worth. **The arithmetic.** A dimension score is a pass rate — passes divided by items scored. With n items behind it, each item is worth 100/n percentage points. At n = 40 that is 2.5 points per prompt, so a 5-point move is two prompts changing verdict. A broad battery like TrustLLM divides its budget across six dimensions and, inside each, across sub-tasks, so the effective n for the thing you actually shipped a fix against is often small even when the whole run was tens of thousands of prompts. Before interpreting any delta, compute 100/n for that sub-task. If your delta is one or two item-widths, you have not measured anything yet. **What flips items other than your mitigation.** - *Decoding.* If generation is sampled (temperature above zero, top-p below one) rather than greedy, borderline prompts land differently between runs by construction. Same model, same prompt, different verdict. - *The scorer.* Whatever decides pass or fail has its own error rate. A model-based judge drifts with its own version and with prompt-order effects; a keyword refusal detector flips on phrasing that has nothing to do with safety. - *Refusal accounting.* Most trust dimensions score a refusal as the safe outcome. A mitigation that only raises refusal rate lifts the column without making the model any less exploitable — and usually costs you points on a helpfulness or over-refusal dimension at the same time. - *The stack around the model.* A hosted endpoint silently updated, a system-prompt edit, a changed default `max_tokens` that truncates long answers into failures, a filter added in the response path. None of these are the mitigation, all of them move the score. **The check that turns a delta into a claim.** Three artefacts, in this order. 1. **A measured noise floor.** Re-run the *unchanged* baseline through the same dimension two or three times with identical settings and record the spread. That is the instrument's own jitter, and any later delta inside it is not a result. You do not need the whole battery for this — re-run only the dimension in question, which is a fraction of a full pass and therefore hours and small money rather than a day and real money. A noise floor is the cheapest thing in this whole procedure and the one most often skipped. 2. **Item-level diffs.** The useful artefact is the list of prompts whose verdict changed, not the average. Get both directions: improved and regressed. 3. **A mechanism.** Can you point at why *those specific* items should have moved? If the flipped items are the ones your change targets, you have a story. If they are scattered across unrelated prompts and roughly balanced in both directions, you have drift wearing a result's clothes. **Where the number misleads even after it survives.** A real improvement on this dimension is an improvement on *this fixed prompt set under this scorer*. It does not generalise to a rephrasing of the same request, because nothing in the run was adaptive. And if the dimension credits refusal, the honest statement is "the model declines more of these specific prompts", which is a different sentence from "the model is harder to exploit". Watch the paired dimension: if safety rose and a helpfulness or truthfulness column fell by a comparable amount on the same items, you moved along a tradeoff rather than up. **Cost of doing it properly.** Roughly triple the token and wall-clock cost of the single post-fix run: baseline repeats plus the post-change run, on one dimension rather than six. In engineer time it is an afternoon, mostly spent making per-item output persist so the diffs are possible at all. That is the price of the difference between a claim and an anecdote, and it is cheap against the cost of announcing a fix that later "regresses" when the number drifts back down on its own. **What to report.** The delta, the noise floor beside it, the item count behind the dimension, and the flipped-item list. A dimension score reported bare — no denominator, no control — invites everyone downstream to over-read it, and you will spend the next release explaining a regression that never happened.

  • How do you establish the noise floor in practice?
    Re-run the unchanged build through the same battery a few times with identical settings and record the spread of each dimension. Any later delta smaller than that spread is not a result.
  • The dimension score rose but exploitability in manual testing is unchanged. What is the most likely explanation?
    The mitigation increased refusals on borderline prompts. The battery credits refusal as the safe outcome, so the score rises while the underlying failure path is still reachable by a rephrased attack.
  • What would make you trust a per-dimension delta more without changing the battery?
    A pre-registered prediction of which items should flip, plus deterministic decoding and a frozen, versioned judging step so the only thing that changed between runs is the mitigation.

A 40-item dimension is a 40-seat parliament: each seat is 2.5 points of the headline, so two prompts changing their vote swings the score five points with nobody's opinion having actually changed. Measure how much the chamber wobbles on its own before you credit an election result.

saying these in an interview costs you the question

  • Reports the delta as a win with no baseline re-run.
  • Never looks at item-level results, only the aggregate.
  • Assumes the battery is deterministic without checking sampling settings.
  • Does not notice that a higher refusal rate can raise a trust dimension without reducing exploitability.
  • Reaches for a p-value on the aggregate while ignoring that the judge itself may have drifted.

context

open as a page

You are choosing an off-the-shelf evaluation to open a red-team engagement against a chat assistant. What does a broad multi-dimension trust battery such as TrustLLM buy you compared with a focused single-purpose attack suite, and what does that breadth cost?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A broad trust battery scores several different trust properties in one pass, so it gives you a wide first map of where a model looks weak. A focused attack suite drills one property hard instead. The cost is depth: each dimension is backed by a thin slice of prompts, so its number is coarse.

open as a page

Your team reports a broad multi-dimension trust battery's per-dimension scores each release. After an output-moderation classifier is put in front of the assistant, one dimension improves clearly and another drops. How do you work out what actually happened before you report it?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Pull the item-level results for both dimensions and look at what the moderation layer did to each prompt. The usual story: it blocks unsafe completions, so the dimension that rewards refusal rises, while a dimension that penalises unhelpful or over-cautious answers falls on the same behaviour. One change, two scoreboards.

open as a page

Where does a broad multi-dimension trust battery such as TrustLLM belong in an AI red-team program that also runs adaptive attack tooling and product-specific scenarios, and what claim should its dimension scores never be used to support?

level: principalimportance: should knowfreq 26%

basics

~20 s

Use it as a cheap, repeatable tripwire and as a way to pick where to attack next. Run it early, then on a schedule. Never let a dimension score stand as an assurance claim that the system is safe: it is a fixed, thin, non-adaptive prompt set that knows nothing about your product.

open as a page