A quarter after your baseline, you re-run the same harmful-behaviour benchmark with the same prompts against the same hosted endpoint alias, and the attack-success rate drops noticeably. How do you establish whether the model behind the alias changed, your harm judge changed, or it is run-to-run variation?
answer
- re-judge old transcripts first
- diff the two manifests
- repeat run for the noise band
- residual equals served system
- investigate improvements too
basics
~20 sChange one thing at a time. Re-score the old transcripts with the new judge: if the rate moves there, the judge drifted. Diff the run manifests for decoding, dataset and error-handling changes. Then repeat the new run to see the spread across repeats before calling any residual gap a real model change.
solid answer
~60 sTreat it as an attribution problem with three suspects and isolate each. 1. **Judge drift.** Re-score the *stored baseline transcripts* with the current harm judge. Any movement here is judge-only, with the model held literally constant. This is the cheapest test and it explains a surprising share of apparent drift. 2. **Harness drift.** Diff the two run manifests: decoding settings, prompt template, system prompt hash, dataset revision, behaviour subset, retry and error accounting. A denominator change or a silently dropped error class looks exactly like a safety improvement. 3. **Sampling variation.** Repeat the new run two or three times with different seeds and look at the spread of the rates. If the quarter-over-quarter gap sits inside that spread, you have no evidence of change. Whatever gap survives all three is attributable to the served system — and note that still bundles *model weights* with *provider-side filtering*, which an end-to-end rate cannot separate. Reading refusal wording and error codes in the transcripts often tells you which: a filter tends to produce a uniform, templated block, a retuned model produces varied in-character refusals.
go deeper
Notices the two runs may not be comparable and asks whether anything else changed before believing the drop.
Diffs the manifests and repeats the run to gauge variation, and knows the judge is a suspect alongside the model.
Re-scores stored baseline transcripts to isolate judge drift, uses a pinned revision as a control where available, and reads transcripts to separate a filter from retuning.
Institutionalises it: transcripts retained, judge pinned per baseline, control runs budgeted, and improvements investigated with the same rigour as regressions.
### The rule Never compare two aggregate numbers when you can hold components fixed and compare one at a time. A moved attack-success rate has three ordinary suspects — the judge, your harness, and sampling noise — and only after all three are excluded does the fourth, the served system, become the explanation. The steps below are ordered by cost, cheapest first, which is also roughly the order of how often each turns out to be the culprit. ### Step 1 — re-judge the past (cheapest, most productive) Take the **stored baseline transcripts** and score them with the current harm judge. The prompts and responses are frozen text, so the model is held literally constant; any difference between this re-scored figure and the original baseline rate is judge drift, isolated exactly. Cost: judge calls only, no target-endpoint calls — for a thousand stored attempts that is minutes of compute and a few dollars, an order of magnitude cheaper than re-running the suite. This is the single strongest reason to retain transcripts, and it explains a surprising share of apparent drift, because harm judges are retrained and their thresholds are tuned. ### Step 2 — diff the manifests Compare the two run manifests field by field: dataset revision, behaviour subset size, temperature, `max_tokens`, attempts per behaviour, the success rule, prompt-template hash, system-prompt hash, and error accounting. The frequent culprit is the last one. If the earlier run dropped timed-out attempts from the denominator and the later run counted them as non-hits, the rate falls for a purely clerical reason and looks exactly like a safety improvement. Cost: near zero — this is a text diff, and it should be done before anyone looks at the rates at all. ### Step 3 — measure your own noise Repeat the current run two or three times under identical settings and look at the spread of the resulting rates. At temperature zero the spread is small but not zero — batching, load balancing across serving replicas and non-deterministic kernels all leak in — and at higher temperature it can be wide enough to swallow the quarter-over-quarter gap entirely. You want the *observed* spread as a local comparison band, not a textbook confidence interval. Cost: N times the full suite. Three repeats of a 200-behaviour, 5-attempt suite is roughly 3,000 target completions plus 3,000 judge calls, so this is the step where money starts to matter; a subset of behaviours is a legitimate economy if the subset is fixed in advance. ### Step 4 — attribute the residual, and know what it still bundles Whatever gap survives the first three steps belongs to the served system. That phrase bundles two different things an end-to-end rate cannot separate: retrained or re-tuned **model weights**, and **provider-side filtering** — an input or output classifier in front of the model that can block a request before the model sees it. Read transcripts rather than the aggregate to split them. A filter tends to emit a uniform, templated block message, a distinctive error code or a truncated empty completion; a retuned model produces varied, in-character refusals that engage the substance of the request. Asking the provider whether the alias was re-pointed is worth doing, but plan for silence. ### Step 5 — the control run, if the budget allows If the provider still serves a pinned snapshot matching your baseline date, run the current harness and current judge against that too. Same instrument, old weights: any gap between it and the alias is the served-model change, cleanly attributed. It costs one extra suite's worth of queries and is the closest thing to a controlled experiment available on a hosted endpoint. ### Where the number misleads Three specific traps. **Directional bias**: a fallen rate is welcomed and never investigated, so harness bugs that suppress successes — a mis-templated prompt the endpoint refuses for the wrong reason — are recorded as safety wins. **False precision**: a change from 14% to 11% sounds decisive until three repeats show a spread of four points. **Attribution by adjacency**: the provider happened to announce an update that month, so the movement is credited to it, when the judge revision landed the same week. Write the residual up as "the served system changed, with judge, harness and sampling excluded", and say plainly that the aggregate cannot distinguish model tuning from a filter.
- You kept no transcripts from the baseline. What is left to you?Only forward-looking work: pin the judge now, re-run the current suite for a fresh baseline with transcripts, and treat the old number as uncomparable rather than pretending the comparison holds.
- The gap survives every check. What do you write in the report?That the served system behind the alias changed between the two dates, with judge, harness and sampling ruled out, and that the aggregate cannot say whether it was model tuning or a provider-side filter.
Re-scoring last quarter's saved transcripts with today's judge is like re-weighing yesterday's parcels on this morning's scale: the parcels cannot have changed, so any difference you see belongs to the scale.
saying these in an interview costs you the question
- Reports the drop as a safety improvement without isolating the judge.
- Never repeats the run, so has no sense of its own noise.
- Assumes an unchanged endpoint name means unchanged behaviour.
- Investigates only regressions and waves through improvements.
- Conflates model retuning with a provider-side filter because the aggregate cannot tell them apart.