skip to content

Your team publishes garak's per-probe calibration scores in a quarterly report. Between two quarters the tested model was not changed, yet several of those relative scores moved. What could explain the movement, and how would you make a quarter-over-quarter comparison of these numbers trustworthy?

level: principalimportance: should knowfreq 25%

answer

  1. relative score has a third input
  2. reference bundle ships with the tool
  3. pin release, bundle, probes, attempts
  4. re-baseline on upgrade, mark the break
  5. trend absolute, annotate with relative

basics

~20 s

A relative score is a position against reference data that ships with the tool, so it moves when that reference data or the probe's prompts change, even with your model untouched. Sampling noise and a silently updated hosted endpoint also move it. Pin the tool and reference bundle per series, and trend absolute scores.

solid answer

~50 s

Three families of cause. **The reference moved:** the calibration data and the probe prompt sets are shipped with the tool and change across releases, so upgrading between quarters redefines the yardstick — if the reference population improves, your position falls with no change on your side. **The measurement moved:** attempts are sampled from a non-deterministic target, so any row with a small attempt count wanders on noise; check whether the interval was even printed before you call the movement real. **The target moved anyway:** behind a hosted endpoint, the served model, its system prompt or its safety filtering can change without a version you control. Make it trustworthy by pinning the run: fix the tool release and the reference bundle for the series, record both beside every published number, hold the probe selection and the attempts-per-probe constant, and re-baseline the earlier quarter whenever the bundle changes so both points share one reference. Trend the absolute score as the primary series and keep the relative position as commentary.

go deeper

for a junior

Recognises that a relative score depends on something other than their own model.

for a middle

Names the reference data and probe content as inputs that change with the tool, and checks sample size before believing movement.

for a senior

Pins and records the whole configuration, re-baselines on upgrade, and uses a re-run of the old configuration to separate real regressions from harness drift.

for a principal

Decides what the published number is for, sets the pinning and re-baselining policy, and accepts the extra run per upgrade as the cost of a trend anyone can rely on.

### Why a relative number is the wrong default thing to trend An absolute score has two inputs: your model and the probe. A calibration score has three: your model, the probe, and the reference population it is scored against — and that third input is owned by the tool and changes on the tool's release schedule, not yours. Trending a quantity whose yardstick moves silently gives you a chart where your posture appears to change while nothing about your system did. The dangerous case is the inverse: a genuine regression on your side, cancelled in the chart by the reference population getting worse at the same time, so the line stays flat and nobody looks. ### Everything that has to be held still Tool release. The calibration data bundled with it. The probe selection you passed to `garak --probes`. The prompt content those probe classes carry, which is edited between releases as new attack literature lands. `garak --generations`. The decoding settings on the target. The system prompt and any guard, retrieval layer or middleware sitting between the harness and the model. Each of these is an input to the printed number, and any one of them moving makes two runs non-comparable. The practical form of this is that the series has a **pinned configuration stored as an artefact**, not a command line somebody retypes each quarter. garak helps here: the run's configuration is recorded in the report file itself alongside the results, so keeping the whole report — not a screenshot of the rate column — preserves the evidence you need to argue that two points are comparable. ### The three families of cause when a number moves **The reference moved.** Upgrading between quarters redefines the population. If the reference models improve and yours stands still, your relative position falls with nothing changed on your side, and the report reads like a regression. **The measurement moved.** Attempts are sampled from a non-deterministic target. Any row with a small attempt count wanders between runs, which is why the first question about any movement is whether the interval printed at all and whether the shift exceeds it. **The target moved anyway.** Behind a hosted endpoint the served weights, a system prompt, or a safety filter can change without a version string you control. "We did not change the model" is a statement about your deployment repository, not about what answered the calls. ### What it costs to do properly The honest handling of an upgrade is to re-run the previous baseline configuration under the new release and publish both series with a marked discontinuity, rather than joining points across the change. That is one extra full sweep per upgrade — if the quarterly run is twenty thousand calls and six hours, the upgrade quarter costs double, in both money and runner time. That is the entire price of a trend anyone can rely on, and it is small next to the cost of a chart that quietly hides a regression. The cheapest discriminator between real movement and harness drift is the same operation in miniature: re-run last quarter's exact pinned configuration today. If the old number reproduces, the change is real; if it does not, the harness moved. ### Where the trend still misleads after all that Backfilling is not available. A stored report is not recomputed against newer reference data, and you cannot re-score old attempts against a new probe prompt set because those prompts were never sent. Nor can you correct the old point by the average shift seen on other probes; the shift is per-probe and not a constant. New probes join the series as new rows starting at the quarter they appeared, never backfilled to look like history. And the failure mode at the other extreme is real too: refusing to upgrade so the yardstick can never move costs you the new probes that are the reason for running the tool at all. A frozen scanner measures last year's attack surface with perfect consistency. ### Decide what the number is for If the quarterly figure is a regression detector for your own deployment, the absolute score under a pinned configuration is the series and the relative position is annotation. If it is a market-position claim, the relative score is the point — but then both quarters must be recomputed under one bundle every time you publish, which means re-running the older configuration, not re-reading its old report.

  • You must upgrade the tool to get new probes. How do you keep the quarterly series honest?
    Re-run the previous baseline configuration under the new release, publish the discontinuity explicitly, and continue the series from the re-baselined point. New probes join as new rows starting from that quarter, not backfilled.
  • What is the cheapest way to tell a real regression from harness drift?
    Re-run last quarter's pinned configuration today. If the old number reproduces, the movement is real; if it does not, something in the harness or the served target changed and the comparison was never like-for-like.

saying these in an interview costs you the question

  • Joining relative scores across a tool upgrade as one continuous trend line.
  • Treating any movement as a model change without checking the interval or attempt count.
  • Assuming a hosted endpoint served the identical model both quarters.
  • No record of which tool release and reference data produced a published number.
  • Refusing to upgrade the tool at all in order to protect the chart.

context