skip to content

You own the rule for how third-party safety-benchmark scores may be cited in your organisation's model-vendor reviews. How do you decide how long such a published number stays citable, and what do you require once it has expired?

level: principalimportance: should knowfreq 28%

answer

  1. two clocks: events and a backstop
  2. tier by what it may decide
  3. third-party never gates a launch
  4. name an owner per trigger
  5. mark series discontinuities

basics

~20 s

Tie expiry to events, not just to the calendar: any endpoint or guard change, a dataset or judge revision, or a new attack family invalidates the number. Add a calendar backstop of a quarter or two. Once expired, a published score may inform shortlisting only; anything gating a launch must be reproduced in-house with a recorded manifest.

solid answer

~1 min

Two clocks, and the earlier one wins. **Event clock — the real one.** A published number expires when anything it silently assumed changes: the vendor updates the model or its serving-side filtering, the benchmark revises its behaviour set or its harm judge, or a materially new attack family appears that the suite never contained. Any of these can move the rate without a date passing. **Calendar backstop.** Because you will not hear about most vendor-side changes, add a flat expiry — a quarter for anything gating a decision, longer for background context. The backstop exists precisely because the event clock depends on disclosure you do not get. **Tiering what a number may be used for** matters more than the exact interval. A fresh third-party score is fine for narrowing a vendor shortlist at any age; nothing published by a third party should ever gate a launch, because it measured a bare endpoint with someone else's harness. Once expired, the number drops to context-only and the decision needs an in-house re-run with a stored manifest and transcripts. Budget honestly: this policy only holds if re-running is cheap enough to actually happen, so size the recurring suite to your query budget rather than writing a cadence nobody funds.

go deeper

for a junior

Says an old published number should be re-checked and is not evidence about today's endpoint.

for a middle

Proposes a cadence and names concrete triggers such as a vendor update or a benchmark revision.

for a senior

Ties expiry to events plus a backstop, and separates what a published number may inform from what requires an in-house run with a manifest.

for a principal

Owns the whole policy: use tiers, named trigger owners, budget-feasible cadence, recorded series discontinuities, and explicit disclosure asks of vendors.

### Set what the number may decide, and the interval follows The policy question is not "how many months is a published safety score good for" but "what is this number allowed to decide". Fix that first and the expiry rule almost writes itself, because the tolerance for staleness is a function of the stakes, not of the benchmark. **Tier the uses.** | use | what may be cited | why | |---|---|---| | context and shortlisting | any published score, labelled with its measurement date, harness and judge | cheap, low stakes, and honest about what it is | | comparing vendors | only numbers from the same benchmark revision, same judge, comparable harnesses | in practice this means you produce them yourself; two published rows are two different experiments | | gating a launch or renewal | an in-house run against your integration, within the current cycle, with a manifest and transcripts | the published run drove a bare endpoint; yours has a system prompt, tools and retrieval | The third row is the load-bearing one. No third-party number, however recent, clears a launch gate — not because it is untrustworthy, but because it measured a different system. ### Two clocks, and the earlier one wins **The event clock is the real one.** A published number expires the moment something it silently assumed changes: the vendor updates the model or its serving-side filtering; the benchmark ships a revised behaviour set or a new harm judge; your own system prompt, retrieval corpus or tool surface changes materially; or a technique lands publicly that the suite never contained. Any of these moves the true rate without a date passing. Name an owner for each trigger — vendor-comms watcher, benchmark-release watcher, the platform owner for your own integration changes — or none of them will ever fire. **The calendar backstop exists because the event clock depends on disclosure you do not get.** Most vendor-side changes are never announced to you at all, and provider-side filtering usually sits outside the model changelog. So add a flat expiry: a quarter for anything feeding a decision, longer for background context. Shorten it for vendors who will not answer questions about change. ### What the policy costs, and the failure mode of writing one you cannot fund A quarterly full re-run is not free. A 200-behaviour suite at five attempts is roughly 1,000 target completions plus 1,000 judge calls per vendor per cycle — tens of dollars of endpoint spend, which is negligible — and then half a day to a day of analyst time triaging flagged transcripts, which is not. Multiply by the number of vendors under review and by four cycles a year and the human line item is what decides whether the policy is real. A cadence the budget cannot fund quietly becomes fiction, and fiction is worse than a longer honest interval, because people keep citing the stale number while believing the policy protects them. Two levers when the budget binds. **Cut coverage, not cadence**: keep a stable, fixed subset plus a control running on schedule, and reserve full coverage for triggered events and the annual review. A small suite run reliably beats a large one run once. **Automate the manifest and the diff**, so the recurring cost is triage rather than setup. ### Where the numbers mislead, and what the policy must prevent The characteristic failure is **citation laundering**: a number that was true on its measurement date is quoted in a review, then copied into the next quarter's deck with the date dropped, and by the third quarter it has become a standing fact about the vendor. Prevent it structurally — require the measurement date, harness and judge to travel with the number in the template itself, so a bare figure cannot be pasted. The second failure is the **smoothed series**. When the benchmark revises its behaviour set or you change judges, the series breaks. Draw one line through the break and a harness change becomes a claimed safety improvement two quarters later. Mark discontinuities in the tracker. ### What you would check, and what to demand of vendors Before a number enters a review: does it carry a date, a harness description and a judge identity; is its tier recorded; is an owner named for each trigger that would expire it. Of the vendor: the measurement date and harness behind any figure they quote, whether the endpoint has changed since, and whether they will notify you of safety-relevant updates. The answers are often unsatisfying — and "this vendor cannot tell us when the model changes" is itself a finding worth writing down, because it sets a floor on how stale your evidence can silently become.

  • A vendor refuses to say whether the endpoint changed since a quoted measurement. How does that affect the policy?
    It collapses the event clock to nothing, so only the calendar backstop protects you. Shorten the backstop for that vendor and record the opacity itself as a finding in the review.
  • You cannot fund quarterly full re-runs. What do you cut first?
    Coverage, not cadence. Keep a stable subset plus a control running on schedule and reserve full coverage for triggered events and the annual review; a small reliable series is more informative than an occasional large one.

saying these in an interview costs you the question

  • Picks an interval with no event triggers, so silent vendor changes are never caught.
  • Lets a third-party published number clear a deployment decision.
  • Writes a cadence the query and triage budget cannot fund.
  • Draws one trend line across a judge or dataset revision.
  • Cites a number without carrying its measurement date and harness forward.

context