Your organisation publishes a quarterly safety attack-success rate for every model it ships, computed by an automated harm judge, and better judges keep appearing. How do you decide when to change the judge, and what do you owe readers of the historical series?
answer
- judge is part of the metric definition
- archive completions from quarter one
- re-score history, overlap window
- hosted judge is not pinnable
- separate ownership of the scorer
basics
~20 sTreat the judge as part of the metric definition. Pin a versioned judge for the published series, archive every raw completion, and when you upgrade, re-score the whole history and publish both series over an overlap window. Never swap silently; state the judge and its measured error rates.
solid answer
~60 sYou are producing two goods that pull against each other: a **comparable time series** and a **currently accurate estimate**. Changing the judge improves the second and breaks the first, and the break is invisible — the number moves with no model change, and readers attribute it to the product. The policy that resolves it: the judge is part of the metric's definition, versioned like a schema. Every quarter archives raw completions, prompts, decoding settings and the judge artifact. When you adopt a better judge, you **re-score the entire archive** with it and publish the corrected series alongside the old one for at least one overlap period, with the delta attributed to the scoring change. That is only possible if completions were retained from the first quarter — a retention and storage decision that has to be made before the series starts, and it is the constraint that actually bites. A hosted judge you do not control is not a pinnable artifact; it can move under you between quarters, which argues for a locally-run scorer as the series judge and the better hosted one for exploratory work.
go deeper
Should see that changing the scorer changes the number without the model changing, so the swap has to be disclosed.
Proposes pinning a judge version and re-scoring past runs, and notes that this requires the old completions to have been kept.
Adds the overlap window, the decomposition of a quarter's movement into judge effect and model effect, published error rates, and a minimum resolvable difference.
Treats the judge as a versioned part of the metric definition with retention, drift-canary and ownership policy attached, and can say when comparability should lose to fidelity and when it should not.
**Frame the decision correctly.** A quarterly published attack-success rate is a measurement instrument with two obligations that pull against each other: **comparability** across quarters, and **fidelity** to the current best understanding of what a successful attack is. Upgrading the harm judge buys fidelity and spends comparability, and unlike a model change the spend is invisible from outside — the number moves, nothing in the release notes says the ruler changed, and readers attribute the movement to the product. **The policy I would run.** 1. *The judge is part of the metric definition.* The series is named for its scorer version, exactly as a schema is. Changing the scorer is a metric revision with a version bump and a changelog entry, never an implementation detail someone lands in a scoring script. 2. *Archive from quarter one.* Prompts, full untruncated completions, decoding parameters, seeds, attempts per behaviour and the reduction rule, retained for the life of the series. This is the decision that determines whether an upgrade is ever possible, and it has to be made before the first publication — you cannot retroactively keep transcripts you threw away. The storage is trivial (kilobytes per attempt, single-digit gigabytes per year of quarterly runs); the retention *policy* is the hard part, especially where transcripts contain harmful content that other policies want deleted. Settle that conflict deliberately, in writing, up front. 3. *Re-score the archive on upgrade.* New judge, entire history, both series published over at least one overlap period, with the quarter's movement explicitly decomposed into judge effect and model effect. 4. *Publish the instrument, not just the number.* Judge identity and version, measured false-positive and false-negative rates, the size and provenance of the adjudicated sample behind them, and a minimum resolvable difference below which quarter-over-quarter movement is not narrated as change. 5. *Prefer a pinnable scorer for the headline series.* A hosted model behind an endpoint can be updated without notice, which is a silent drift channel in the metric. Run a locally hosted classifier or a snapshot-pinned model for the series; use the better hosted judge for exploration and for the periodic audit. 6. *Drift canary.* A frozen, human-adjudicated transcript set re-scored every cycle. If its labels move on unchanged inputs, the instrument moved, and you learn that before the headline does. **What the policy costs.** Steady state is cheap: the canary is a few hundred transcripts per cycle, and the archive is negligible storage. The bill lands on an upgrade, because re-scoring the whole history scales with quarters times transcripts — eight quarters at twenty thousand transcripts is a hundred and sixty thousand scoring calls, which is a real GPU day or a real invoice, plus the engineer time to re-run pipelines whose code has since moved on. That cost is the reason organisations skip step 3, and skipping it is precisely what turns the series into fiction. Budget the re-score into the upgrade decision rather than discovering it afterwards. **Where the number misleads.** The headline failure is the silent swap: the scorer is upgraded between quarters, the rate drops three points, and the release notes read it as a safety improvement. The seductive shortcut — apply a fixed offset to past quarters so the line joins up — is worse than doing nothing, because the two judges disagree unevenly across behaviours, attack families and models, so a single offset fabricates a correction no data supports. Equally misleading in the other direction: an unchanged judge with a demonstrated blind spot on an attack family that is now in scope reports a flat, comparable, and entirely uninformative series. **Governance.** Whoever owns the judge owns the metric. If the team optimising the models also owns and tunes the scorer, the number is captured; separate the ownership, or at minimum require that judge changes are reviewed and dated by someone outside model development. **When I would and would not change.** I would refuse a mid-series swap for a marginal accuracy gain when the archive cannot be re-scored — the comparability loss exceeds the benefit. I would change immediately, and take the discontinuity, when the incumbent judge has a demonstrated systematic blind spot on an in-scope attack family, because a series that cannot see a real risk class is worse than a break in the line. Either way the decision is published with its reasoning, so a reader can tell an instrument event from a product event.
- What single engineering decision most determines whether you can ever upgrade the judge?Retaining raw completions and their generation settings from the first quarter. Re-scoring history is only possible if the transcripts still exist; that is a storage and retention commitment made up front.
- How do you detect that a hosted judge has drifted between quarters?Re-score a frozen, human-adjudicated canary set every cycle and track its labels. Movement on unchanged transcripts is instrument drift, not model change.
- Who should own the judge, and why does it matter?Not the team whose models the metric scores. If the measured party controls the scorer, the metric is captured; judge changes should be reviewed and dated independently.
Swapping the judge mid-series is like re-cutting the ruler between two measurements of a growing child. The new ruler is more accurate and the two marks are still uncomparable until you go back and re-measure the old one.
saying these in an interview costs you the question
- Swaps the judge between quarters and reports the resulting movement as a safety improvement.
- Keeps a demonstrably blind judge forever in the name of comparability, with no re-scoring path.
- Anchors a multi-year series on a hosted endpoint that can change without notice.
- Has no archived completions, so the history can never be re-scored under a new definition.
- Lets the team whose models are being measured own and tune the scorer.