The team behind an LLM product edits its system prompt most weeks, the safety-classifier service upgrades on its own schedule, and the provider can swap the model behind the endpoint at any time. A full red-team suite rerun costs real metered spend and analyst time. How would you design the assurance cadence?
answer
- tier ladder with escalation
- trigger on config events, not the calendar
- pin the model, reduce the change rate
- results void on fingerprint change
- bill churn to the change owner
basics
~20 sMake the trigger mechanical rather than calendar-based: subscribe to the config and deploy pipeline, and tier the suite. A tiny high-severity tier runs on every prompt or guard change, a mid tier per release train, the full suite on a model swap or quarterly. Any tier hit escalates to the next.
solid answer
~50 sFour design moves. **Tier the suite by cost and severity.** A canary and highest-severity tier costing minutes and pocket change, a representative mid tier, and the full suite. Each change class earns a tier, and any hit in a cheap tier escalates upward. This is what makes weekly change survivable. **Trigger mechanically.** Wire triggers to the events themselves — prompt and tool-definition config commits, guard version changes, and your own model-swap detection — instead of a calendar that is wrong the day after it runs. **Reduce the change rate at the source.** Negotiate pinned dated model identifiers and a notice window. Turning a silent swap into a scheduled event is cheaper than any amount of rerunning. **Be honest about scoping.** A one-line instruction edit can move behaviour anywhere, so whatever subset you rerun is a sampled bet and should be described that way rather than as coverage. Price it so the change-owner's budget funds the per-change tier.
go deeper
Suggests rerunning after big changes and on a regular schedule.
Proposes a small fast subset per change plus a full run on a longer clock, and explains why the endpoint being unchanged is not evidence of stability.
Designs the tier ladder with escalation rules, wires triggers to config and guard version events, and adds swap detection for the unannounced case.
Attacks the change rate itself through pinning and pipeline discipline, funds tiers where the churn originates, makes stale results visibly void, and reports the fraction of the suite currently fresh instead of implying coverage.
The structural problem is that the change rate of the system exceeds the assurance rate of the programme. Weekly prompt edits, an independently released classifier and a provider that can repoint the served model at will together produce more invalidating events than any rerun policy can answer. No cadence fixes that on its own, so you work three levers at once: reduce the number of changes that demand a response, make the response cheap, and make staleness visible when neither worked. ### Reduce the change rate at source Pinning a dated, immutable identifier in the request's `model` field converts the worst class — the unannounced swap — into a planned migration with a window you can test in. Requiring that system-prompt and tool-definition edits ship through the same pipeline as code, with a hash recorded on merge, converts the second class into an event with an artefact you can subscribe to. Both are organisational asks rather than tooling, and they are the highest-leverage moves available to whoever leads this. Expect pushback on both: pinning gives up free provider improvements and creates a migration owner, and routing prompt edits through review slows the team whose whole advantage was that prompts are not code. ### Make the response cheap: a tier ladder Publish the ladder with its costs, so the product team can see what its change rate buys: | Tier | Content | Trigger | Rough cost | |---|---|---|---| | 0 | Benign canaries + swap detection | Continuous | ~20 requests/day | | 1 | Highest-severity attack prompts, modest trials | Every prompt or guard change | minutes, low hundreds of requests | | 2 | Stratified sample across the suite | Per release train, or any Tier 1 movement | low thousands of requests + triage | | 3 | Full suite at high trial counts | Model swap, Tier 2 movement, or the quarterly clock | tens of thousands of requests + days of analyst time | Escalation is the mechanism that makes the cheap tiers legitimate. A Tier 1 pass is not a safety claim; it is a tripwire whose only jobs are to fire and to buy the funding for Tier 2. Triggers wire to events — config commits carrying prompt or tool-definition changes, guard version changes, and your own swap detection — rather than to a calendar that is wrong the day after it runs. ### Make staleness visible Every filed result carries the configuration fingerprint it was measured under, and the reporting surface marks a result whose fingerprint no longer matches production as **void**, not as passing. This is also the programme's political protection: when an incident eventually lands on a configuration nobody retested, the record shows the result was already marked stale and what a refresh would have cost. ### Fund it where the churn originates Put Tier 0 and Tier 1 spend on the product team's budget line. A team that ships prompt edits weekly and pays for the resulting runs consolidates its edits; a team whose reruns are billed to security has no reason to. This is the only feedback loop in the design that acts on the change rate itself. ### Where the numbers mislead Three readings to guard against. First, a green cheap tier read as "the change was safe" — it covers a handful of behaviours at low trial counts, and a one-line instruction edit can move behaviour anywhere in the surface, so any tier-scoped rerun is a sampled bet and should be written up as one. Second, a freshness percentage read as coverage: "82% fresh" means 82% of the suite's last results still match the live fingerprint, not that 82% of the attack surface is covered — the suite itself is a sample of an unbounded space, and that denominator was never full coverage to begin with. Third, the quarterly report that lists last quarter's results without dates, implying the system has been in that state since. Publish the freshness figure as a standing metric precisely because it falls on its own as the product changes; that decay is the pressure you want visible to the people creating it. ### Accept and state the residual On any realistic budget most of the surface is untested most of the time. Say the number rather than letting the format imply otherwise, and record which change classes you decided not to answer at all — that is a risk acceptance, and it should have a named owner outside the security team.
- Why not simply run the full suite nightly and stop worrying about triggers?Because it spends the metered budget and analyst triage time continuously for a system that mostly did not change, and the noise floor of nightly rate fluctuation swamps the signal you actually want.
- What one metric would you put on the programme dashboard for this?The share of the suite whose last result matches the live configuration fingerprint — freshness. It goes down by itself as the product changes, which is exactly the pressure you want visible.
saying these in an interview costs you the question
- A pure calendar cadence with no change-triggered path.
- Claiming a rerun can be scoped precisely to the behaviour a system-prompt edit touched.
- Presenting a cheap per-change tier as evidence the system is still safe rather than as a tripwire.
- Absorbing all rerun cost into the security budget, leaving the change rate with no owner or feedback.