skip to content

A control in your organisation's AI risk register is marked satisfied on the evidence of an AI red-team assessment against a hosted third-party model your team does not control. How long is that verdict good for, and what do you write into the entry so it expires honestly?

level: seniorimportance: should knowfreq 44%

answer

  1. point-in-time evidence, moving target
  2. three clocks: provider, application, technique
  3. pinned vs floating model alias
  4. expiry triggers, not just a date
  5. canary set detects silent drift

basics

~20 s

It is good until the thing you tested changes. With a hosted model the provider can change it without telling you, so pin the entry to what you exercised: the model identifier, the system prompt and guard configuration, the date, and re-test triggers — provider change, prompt or tool change, or a fixed interval.

solid answer

~50 s

The honest answer is that the evidence is point-in-time and the target is not yours. Three clocks run against it. The **provider clock**: a hosted endpoint can be updated underneath you, and behaviour on refusal boundaries is exactly what moves. The **application clock**: your own system prompt, tool permissions and retrieval corpus change every sprint, and each of those changes the reachable surface. The **suite clock**: attack techniques that did not exist at assessment time were, by definition, not tested. So write an expiry with triggers, not just a date. Record the model identifier and endpoint as configured, the application configuration, the date, and the events that invalidate the verdict. Add a standing review interval as a backstop, because you will not always learn about a provider-side change. The register entry should read as "satisfied as of, under this configuration, re-test on these events", never as an unqualified satisfied.

go deeper

for a junior

Says the result is point-in-time and needs a date, and that a re-test should happen if the model or the app changes.

for a middle

Separates provider-side change from application-side change and records the configuration the verdict was measured under.

for a senior

Writes explicit expiry triggers plus a backstop interval, distinguishes pinned from floating model identifiers, and runs a canary set to detect silent provider drift.

for a principal

Owns the process: who is accountable for the trigger, how application change reaches the register holder, and how re-test capacity is budgeted across a portfolio of AI systems.

### Why the usual control-evidence habit breaks here Traditional assurance assumes a mostly static system: you tested the authorisation check, the code did not change, the evidence stands until the next release. A hosted model breaks that on the vendor's schedule rather than yours, and refusal behaviour at the boundary is precisely the property that moves when a provider updates a model or its safety stack. A governance process built for the static case will carry a stale pass for a year and never notice. ### Three clocks run against the entry *The provider clock.* A hosted endpoint can be updated underneath you. Whether it can depends on one field: the model identifier. A **pinned** identifier names a fixed snapshot the provider commits to serving unchanged; a **floating** alias resolves to whatever the provider currently considers latest. An entry naming a floating alias expires continuously, because the subject of the sentence changes without the sentence changing. *The application clock.* Your own system prompt, tool and function permissions, and retrieval corpus change every sprint. Each of these alters the reachable surface: adding a tool changes what a successful injection can *do*, which can turn a nuisance finding into a material one without any model change at all. *The technique clock.* Attack families that did not exist at assessment time were, by definition, not exercised. This is the weakest of the three in practice but the one auditors ask about. ### What to write into the entry - **Target identity as configured**: provider, model identifier or deployment alias, and explicitly whether that identifier is pinned or floating. - **Application context**: system prompt version, tool and function permissions, retrieval sources, and any input or output guard in front, named by role and configuration. - **What was exercised**: attack families covered, and — as its own line — families not covered. - **Verdict basis**: the measurement, the threshold, and who owns the threshold. - **Expiry triggers**: a detected or announced provider behaviour change; a material change to the prompt, tool set or retrieval corpus; a newly relevant attack family; and a fixed interval as a backstop, because you will not always learn about the first. The entry should read as "satisfied as of *date*, under *this* configuration, re-test on *these* events", never as an unqualified satisfied. ### Detecting the provider clock, and what it costs You cannot rely on being told. The cheap mitigation is a **canary set**: a small, fixed collection of cases with known previous outcomes — some that were blocked, some that were allowed — replayed on a schedule, alerting when an outcome flips. Sizing is the whole question. Fifty cases replayed twice a day is a hundred calls a day, roughly three thousand a month: negligible money and negligible wall clock. The real cost is human — somebody must triage the alerts, and an untriaged canary is worse than none because its silence is later cited as evidence. ### Where the canary's number misleads A canary compared as a single run against a single run will lie to you. A sampled endpoint at non-zero temperature flips individual outcomes on its own, so one changed case is noise, not drift. The fix is mechanical: replay each canary case several times per check and treat the *proportion* of the set that changed against its historical spread as the signal, with an alerting threshold you set from a baseline period rather than from the first flip. Teams that skip this get one of two failures — alert fatigue, so the canary is muted, or a threshold set so high that the drift it exists to catch never trips it. Either way the register keeps its stale pass, and the only difference is which explanation you give afterwards. The second misleading reading is the date field itself. An entry with a review date and no triggers looks maintained. It says only that somebody will look again eventually; it makes no claim that the evidence is still about the same system, and a reader who sees a recent date will assume it does. ### The organisational half Somebody must own the trigger, and it is rarely the red team. The entry needs a named owner and a route by which the application team's sprint change becomes visible to whoever holds the register — otherwise the release that adds a new tool silently invalidates evidence nobody revisits. What separates a strong answer here from an average one is treating expiry as a process with an accountable owner and a detection mechanism, not as a date field somebody fills in.

  • The provider changes the model behind a floating alias and nothing in your register mentions it. What is the first thing you fix?
    Pin to a versioned identifier where possible, and make the alias itself a recorded field of the entry so the mismatch is visible. Then add scheduled canary checks so the next change is detected rather than assumed.
  • The application team adds a new tool the assistant can call. Does that invalidate the entry?
    Usually yes for anything touching the reachable surface — new tool access changes what a successful injection can do. It should be an explicit re-test trigger, scoped to the affected findings rather than a full re-run.

It is a roadworthiness certificate for a car the manufacturer can rebuild in your driveway overnight without telling you. The certificate is honest about the day it was issued and says nothing about the vehicle you are driving now, which is why the entry needs triggers rather than only a date.

saying these in an interview costs you the question

  • Treating an AI control verdict as valid until the next annual audit.
  • Recording a floating model alias as if it were a fixed target.
  • No mention of the application-side configuration the result depends on.
  • Assuming the provider will notify you about behaviour changes.
  • Making re-test a red-team decision with no accountable register owner.

context