skip to content

Your team reports a broad multi-dimension trust battery's per-dimension scores each release. After an output-moderation classifier is put in front of the assistant, one dimension improves clearly and another drops. How do you work out what actually happened before you report it?

level: seniorimportance: should knowfreq 35%

answer

  1. one intervention, two columns
  2. refusal credited here, penalised there
  3. attribute flips to the filter
  4. model underneath is unchanged
  5. report the tradeoff, not two deltas

basics

~20 s

Pull the item-level results for both dimensions and look at what the moderation layer did to each prompt. The usual story: it blocks unsafe completions, so the dimension that rewards refusal rises, while a dimension that penalises unhelpful or over-cautious answers falls on the same behaviour. One change, two scoreboards.

solid answer

~50 s

Do not report the two movements as two independent findings until you have checked whether they are one behaviour seen from two directions. A filter in front of the model converts some completions into blocks or canned refusals. A dimension whose scoring treats a refusal as the correct outcome gains; a dimension that scores whether the model answered truthfully, helpfully or without unjustified withholding loses on exactly the same responses. So the diagnosis is item-level. For each dimension, list the prompts whose verdict changed and check whether the response was intercepted by the moderation layer. If the same intervention explains both lists, you have a **safety-versus-over-refusal tradeoff**, and reporting it as "one dimension up, one down" misleads the reader. The other checks worth running: was anything else deployed in the same window, is the drop bigger than the dimension's known run-to-run spread, and does the moderation layer see the whole response or only part of it?

go deeper

for a junior

Notices the two movements might share a cause and asks to see which prompts changed.

for a middle

Explains that a refusal is scored as correct by one dimension and as a failure by another, so a filter moves both, and checks item-level verdicts.

for a senior

Attributes each flipped item to the moderation layer, rules out same-window confounds, compares both deltas with the measured noise floor, and reports the tradeoff explicitly rather than two independent numbers.

for a principal

Sets the reporting contract — filter changes are annotated on every affected dimension — and makes the point that filtering changes observations, not the underlying model, so assurance claims cannot rest on it.

**Start from the premise that one shared cause beats two coincidences.** An output-moderation classifier sits in the response path: the model generates, the classifier inspects the completion, and the pipeline either passes it through, replaces it with a canned refusal, or truncates it. That is a single intervention that rewrites the *text* of many responses. A broad battery computes several dimensions from those same texts. It would be surprising if the intervention showed up in exactly one column. **Why one change moves two columns in opposite directions.** The dimensions do not share a scoring convention. A safety-flavoured sub-task treats a refusal or block as the correct outcome — the scorer looks for withholding and awards a pass. A truthfulness, helpfulness or over-refusal sub-task treats the same withheld response as a failure, because the prompt was benign and a substantive answer was required. The classifier does not know which sub-task a prompt came from; it fires on surface features of the completion. So on a benign prompt whose answer happens to look risky, one column gains a point and another loses one, from the same intercepted response. **The diagnosis, in order.** 1. *Item diffs for both dimensions.* Pull per-prompt verdicts from both runs and build two lists: items that improved, items that regressed. 2. *Attribute each flip to the layer.* For each flipped item, was the response blocked, replaced or truncated by the classifier? Your moderation layer should be logging its decision and category per call; if it is not, that is the first thing to fix, because without it this whole analysis is guesswork. 3. *Look for the mirror.* If improved items are ones where withholding is the scored-correct behaviour, and regressed items are ones where a substantive answer was expected, you have one tradeoff, not two effects — and that is a better finding than either score alone. 4. *Rule out confounds.* Anything else that shipped in the same window — a system-prompt edit, a hosted model version bump, changed decoding settings — produces the same pattern. 5. *Compare both moves against the noise floor.* A dimension backed by a thin item set moves several points on its own. Whichever movement sits inside the spread of repeated unchanged runs is not a movement. 6. *Check what the classifier can see.* If it inspects only the final assistant message while the system streams partial output, reasons in a separate channel, or emits tool calls, the score reflects a partial view of what the deployed system actually put on the wire. **Where the number misleads.** This is the paragraph that matters. After the filter, the battery is no longer measuring the model — it is measuring the **pipeline**, and the two have different security properties. The model underneath is byte-for-byte unchanged: every unsafe completion it produced before, it still produces, and the filter's decision is a second classifier with its own false-negative rate standing between that completion and the user. A rephrased request that produces a completion the classifier does not recognise reaches the original behaviour untouched. So a rise in a safety dimension after adding a filter licenses exactly one sentence — "the deployed path emitted fewer scored-unsafe responses on this fixed prompt set" — and not the sentence everyone will hear, which is "the model got safer". The block rate is also a confound in its own right: if the filter intercepts a large share of responses, several dimensions are now partly measuring the classifier's behaviour rather than the model's, and cross-model comparisons through that path are no longer like-for-like. **What it costs to measure properly.** Run the battery twice — once against the raw model endpoint, once against the filtered path — and report both. That doubles tokens and wall clock and adds a per-call moderation charge on the filtered run, which for a full battery pass is real but small money; the standing cost is keeping a raw endpoint reachable in a non-production environment, which is a plumbing and access decision more than a budget one. The gap between the two runs is the filter's contribution, measured rather than inferred, and it is the only way to keep reporting on the model separable from reporting on the guard. **What to report.** Not "dimension A +6, dimension B −5". Report: the classifier intercepted N responses; M of those were prompts where withholding is scored correct and K were prompts where a substantive answer was expected; the net is a deliberate move along a safety/utility tradeoff, prompts attached, both deltas set against the measured noise floor. Then state what it does not establish — filtering changes what the battery observes, not what the model does.

  • The stakeholder asks whether the model is now safer. What do you say?
    That the deployed system emits fewer unsafe responses on this fixed prompt set, and that the model behind the filter is unchanged, so anything that evades the filter reaches the same behaviour as before.
  • How would you separate the filter's effect from the model's in future runs?
    Run the battery twice — once against the raw model endpoint and once against the filtered path — and report both. The gap between them is the filter's contribution, measured rather than inferred.

The filter is a press officer standing between the speaker and the microphone. The transcripts get cleaner, the speaker's opinions do not change, and anyone who gets a question past the press officer hears exactly what was always there.

saying these in an interview costs you the question

  • Reports the two dimension movements as unrelated findings.
  • Celebrates the improved dimension and quietly omits the regression.
  • Never checks which responses the moderation layer actually intercepted.
  • Concludes the model became safer when only the response path changed.
  • Ignores other changes deployed in the same window.

context