skip to content

Version Drift

A leaderboard row is a photograph of a model that may have been updated since, and nothing in the report says so. Interviewers ask when you would re-run the suite rather than cite last quarter's row.

on this pageshow

explore

questions

5

You run a public harmful-behaviour benchmark such as HarmBench against a hosted chat endpoint and save the resulting attack-success rate for later comparison. What do you record alongside the number so it can still be interpreted in six months?

level: juniorimportance: must knowfreq 50%

answer

  1. run manifest next to the rate
  2. endpoint plus resolved revision
  3. decoding settings and error accounting
  4. judge identity and threshold
  5. keep transcripts to re-judge later

basics

~20 s

Record the endpoint string and any pinned revision it resolved to, the timestamp, decoding settings such as temperature and max tokens, the dataset revision and which behaviours ran, the harm judge and its version, and the raw transcripts. Without these, a later number cannot be compared and any difference cannot be explained.

solid answer

~50 s

Save a **run manifest** next to the rate, because the rate alone is uninterpretable later. Minimum contents: - **What was tested:** endpoint string, any pinned revision identifier the provider exposes, region or deployment if relevant, and whether a system prompt was attached. - **How it was driven:** decoding settings, max tokens, retry and timeout policy, concurrency, and how errors, empty responses and truncations were counted. - **What was asked:** benchmark dataset revision and the exact behaviour subset, since a rate over a different denominator is not a comparison. - **Who scored it:** the harm judge's identity and revision, and its threshold. - **The evidence:** raw prompts and responses. Transcripts are the highest-value item. With them you can re-judge an old run under today's scorer, which separates judge drift from model drift. Without them, the run can never be re-examined and only the aggregate survives — and the aggregate is exactly the part that cannot explain itself.

go deeper

for a junior

Names timestamp, endpoint, decoding settings, dataset revision and judge, and knows the raw outputs should be kept.

for a middle

Adds error and retry accounting, the system prompt hash, and explains that the denominator must match for two rates to be comparable.

for a senior

Has the harness emit the manifest automatically, captures any served-revision metadata the provider returns, and sets a deliberate retention policy for sensitive transcripts.

for a principal

Makes the manifest a required artefact for any safety number cited in a decision, so comparisons across teams and quarters are auditable rather than anecdotal.

### Why the rate alone is uninterpretable later A benchmark result is a compound claim: *this endpoint, driven this way, judged this way, on this date, over this behaviour set, produced this rate.* Every clause you fail to write down is a clause the next reader has to guess, and guessing is how two runs that were never comparable end up on the same trend line. The fix is unglamorous: emit a **run manifest** from the harness alongside every number, and keep the raw transcripts. ### What the manifest must contain ```yaml run_id: 2026-Q3-safety-suite-014 started_at: <iso-8601 timestamp> target: endpoint: <provider endpoint string> pinned_revision: <snapshot id, or 'alias-floating'> served_revision_seen: <id returned in response metadata, if any> system_prompt_sha256: <hash, or 'none'> sampling: temperature: 0.0 max_tokens: 512 attempts_per_behaviour: 5 success_rule: any_attempt_counts_behaviour on_error: retry_3_then_count_as_non_hit dataset: name: <benchmark name> revision: <commit or release id> behaviours_run: 210 judge: name: <harm judge> revision: <id> threshold: 0.5 harness: template_sha256: <hash of the prompt-building code> results: attack_success_rate: 0.14 attempts: 1050 errors: 3 transcripts: <restricted evidence store path> ``` Four fields people skip and later regret. **The served revision, not just the alias.** An endpoint name is a routing label; the weights and the serving-side filtering behind it can be replaced without it changing. If the provider echoes a model or snapshot identifier in response metadata, capture it — that is the only in-band evidence of what actually answered. **Error accounting.** A run where forty attempts timed out and were silently dropped reports a rate over a smaller denominator. Next quarter, with timeouts counted as non-hits instead, the rate falls with no change in behaviour whatsoever. Say explicitly which convention you used. **The success rule and attempts per behaviour.** "A behaviour counts as compromised if any of five attempts landed" and "the fraction of individual attempts that landed" are different statistics computed from the same transcripts, and the first is systematically larger. **Hashes of the system prompt and the templating code.** A harness change then shows up as a one-line manifest diff instead of as a mysterious movement in the rate. ### What it costs The manifest itself is free — a few dozen lines emitted by the harness, no extra model calls. The real costs are storage and handling of the transcripts. A suite of a couple of hundred behaviours at five attempts produces on the order of a thousand prompt/response pairs, a few megabytes of text: trivial to store, but it is a corpus of attempted harmful elicitations, so it belongs in the same restricted evidence store as the rest of the engagement, with an access list and a deliberately chosen retention period. Budget the review time too — someone must confirm the transcripts are complete and correctly redacted before they are archived. Keep the manifest even after the transcripts expire: a manifest with no transcripts still tells you whether the next run is comparable, which is most of the value. ### Where an unmanifested number misleads The failure is always the same shape: a rate moves, and the movement is read as the endpoint getting safer or riskier. With no manifest, that reading cannot be defended, because a changed judge revision, a revised behaviour set, a different temperature, a new retry policy, or a different success rule each produce the identical symptom. Worse, the aggregate is the one artefact that cannot explain itself, and it is the only one that survives into a slide deck. A number with no manifest is not a weak measurement; it is an uninterpretable one, and treating it as weak-but-usable is how it gets cited anyway. ### What you would check Before archiving a run: does the manifest name the judge revision and threshold; does it distinguish the alias from any served revision; does it state the error convention and the success rule; are the transcripts stored and readable; does the behaviour count match the dataset revision. Before comparing two runs: diff the two manifests first and only then look at the rates. Most apparent drift is found in that diff in under a minute.

  • Why is the harm judge's revision as important as the model's?
    Because the judge decides which attempts count. A stricter or retrained judge moves the rate with no change in the endpoint at all, and only a recorded judge revision lets you rule that out.
  • Which single artefact most improves a run's future value?
    The transcripts. They let you re-score an old run under a current judge, which is the only way to separate judge drift from model drift after the fact.

saying these in an interview costs you the question

  • Saves only the aggregate rate and a model name.
  • Does not record the harm judge or its threshold.
  • Drops errored attempts silently, changing the denominator between runs.
  • Keeps no transcripts, making the run impossible to re-judge.
  • Records the alias only, never the revision the provider actually served.

context

open as a page

A public safety-benchmark leaderboard shows an attack-success rate measured some months ago against a hosted chat endpoint. What does that row fail to record about the system it scored, and why does that limit citing the number today?

level: middleimportance: must knowfreq 58%

basics

~20 s

It records a name, not a build. A hosted endpoint's alias can be re-pointed to an updated model, and the safety tuning and any provider-side filters behind it can change with no announcement. The row also freezes the decoding settings and the harm judge used then. It is one snapshot, not today's endpoint.

open as a page

A provider offers both a floating alias for a chat model and a pinned snapshot identifier of the same model. Which of the two should a recurring safety-benchmark suite be pointed at, and what does each choice cost you?

level: middleimportance: should knowfreq 36%

basics

~20 s

Point at both if you can. The alias is what production calls, so running it is your drift detector. The pinned snapshot is a control: same weights every quarter, so a moved rate there means your harness or judge moved. Costs are double the queries, and pinned snapshots eventually get retired.

open as a page

A quarter after your baseline, you re-run the same harmful-behaviour benchmark with the same prompts against the same hosted endpoint alias, and the attack-success rate drops noticeably. How do you establish whether the model behind the alias changed, your harm judge changed, or it is run-to-run variation?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Change one thing at a time. Re-score the old transcripts with the new judge: if the rate moves there, the judge drifted. Diff the run manifests for decoding, dataset and error-handling changes. Then repeat the new run to see the spread across repeats before calling any residual gap a real model change.

open as a page

You own the rule for how third-party safety-benchmark scores may be cited in your organisation's model-vendor reviews. How do you decide how long such a published number stays citable, and what do you require once it has expired?

level: principalimportance: should knowfreq 28%

basics

~20 s

Tie expiry to events, not just to the calendar: any endpoint or guard change, a dataset or judge revision, or a new attack family invalidates the number. Add a calendar backstop of a quarter or two. Once expired, a published score may inform shortlisting only; anything gating a launch must be reproduced in-house with a recorded manifest.

open as a page