The model behind a client's application is a hosted third-party service the client cannot pin or freeze. How do you write scope and findings validity so the report still means something three months later?
answer
- measurement of a moving system
- stamp identifier, config, dates, rates
- validity window plus named triggers
- recommendations the client can actually apply
- regression set, not a one-shot report
basics
~20 sRecord the exact model identifier, endpoint and dates tested, and state that findings describe that snapshot. Write a validity window and named re-test triggers into the scope: a provider model update, a system-prompt change, a new tool or corpus. Push the client toward durable application-side controls rather than one-off prompt patches.
solid answer
~60 sA finding against a hosted model is a measurement of a moving system, so it needs an expiry the same way a scan of a patched host does — but the thing that moves is not under the client's control. Three moves make the report survive: 1. **Stamp the measurement.** Model identifier, deployment or region, any routing rules, sampling settings, the instruction layer, the retrieval index, the tool list, the moderation components, and the dates. Without this the report is unfalsifiable later. 2. **Declare validity explicitly.** State a window and the events that end it early: the provider updates or retires the model behind that identifier, the client changes the instruction layer, a tool or data source is added, or a moderation component is reconfigured. 3. **Aim recommendations at what the client controls.** A model-side behaviour they cannot fix becomes an application-side control — output handling, tool authorisation, confirmation on irreversible actions — plus a signal that catches regression. The consequence is that a one-shot annual test is the wrong product shape here, and the scope conversation should say so.
go deeper
Knows the report should say which model and date were tested.
Captures the full configuration stamp and understands that a hosted model can change without the client acting.
Writes a validity window with named re-test triggers, stores a reproducer set with attempt counts and rates, and re-aims recommendations at application-side controls the client owns.
Reshapes the commercial engagement — continuous or triggered re-testing with a maintained regression set instead of a point-in-time report — and makes the absence of any provider change-notification path a finding in its own right.
## The asymmetry that breaks the normal finding lifecycle In a conventional engagement the tested system changes only when the client changes it, so a finding stays true until they act on it, and a re-test is a clean before-and-after. Behind a hosted model neither half holds. - The vendor can update, re-route or retire a deployment on its own schedule; - the server-side safety stack sitting in front of the weights is updated independently of the weights; - and the system is **probabilistic**, so the same input can produce a refusal and a compliance on consecutive calls with nothing changed at all. Two consequences follow, and they point opposite ways: a fixed finding can silently return, and a genuine finding can become irreproducible in a way that makes your report look like it was wrong. ## What actually moves — check which of these you are pinned to - A model identifier in a client's config is often an ***alias*** that the provider re-points to a newer snapshot, rather than a pinned snapshot string; those behave completely differently across three months. - **Sampling settings** — temperature, top-p, seed if the provider offers one — change the rate without changing the system. - **Routing** between a cheap and an expensive tier can be traffic-dependent. - And the **moderation layer** in front of the model is a separate product with its own release train. ## Stamping the measurement Record the identifier and everything around it: deployment or region, routing rules, sampling parameters, the instruction layer (hashed if you may not store it), the retrieval index, the tool list, the moderation components and their enabled state, and the dates. Then store a **reproducer set**: the attempt inputs, the observed outputs, and timestamps, so a later disagreement is settled by re-running rather than by memory. Crucially, where the finding is statistical — and on a probabilistic system nearly all of them are — record the *attempt count and the number of successes*, not "it worked". A finding written as "3 successes in 30 attempts" is falsifiable. A finding written as "the model complied" is not. ## Where the number misleads, and this is the paragraph that matters Three months on, someone re-runs your reproducer ten times, sees ten refusals, and reports the issue fixed. That inference is unsound. If the true success rate is one in ten, the probability that ten independent attempts all fail is 0.9 to the tenth power, roughly 0.35 — so a clean ten-attempt re-test happens by chance about a third of the time against an entirely unmitigated system. - To distinguish a fix from sampling you need the **original attempt count**, at minimum, and preferably more; - and you need the **original sampling settings**, because raising temperature alone will move the rate. - Worse, if the re-test used a **different judge model or a different scoring prompt**, the rate moved for a reason that has nothing to do with the target at all. A **partial mitigation**, meanwhile, has exactly the shape people misread as a fix: the rate drops but does not reach zero, and a small re-test cannot see the difference. ## What it costs Stamping adds perhaps an hour per environment. The reproducer set adds a day to pin and script. Re-running it on a schedule costs the same token bill as a small slice of the original campaign, repeated — cheap per run, and it is that cheapness which makes periodic re-testing viable where a full annual engagement is not. Against that, the cost of an unstamped report is that it cannot be defended, cannot be re-tested, and quietly becomes an audit artefact asserting a safety posture nobody has measured for a year. ## Validity and recommendation altitude Write the window into the scope, with the events that end it early: - the provider updates or retires the deployment behind that identifier, - the client edits the instruction layer, - a tool or data source is added, - or a moderation component is reconfigured. Name who watches for those triggers — usually the client — and what they do when one fires. Then re-aim the recommendations. The client cannot patch the model, so "the model should refuse this" is not actionable. - Constrain what the output is permitted to do, - put authorisation at the tool boundary, - require confirmation before irreversible actions, - keep untrusted retrieved content out of the instruction position, - and add a detection that fires if the behaviour returns. Those survive a model swap; a prompt patch does not. ## What you check before signing - Whether the identifier is an alias or a pinned snapshot. - Whether the client has any notification path from the provider about model changes. If they do not, that is your first finding — because it means no trigger will ever be observed, and the validity window you wrote is decorative.
- Three months on, a finding no longer reproduces. What do you conclude?Not much on its own. Re-run the recorded reproducer at the original attempt count and compare rates; a drop is consistent with a provider-side change, a sampling difference, or a client-side edit. Only a configuration diff tells you which, which is why the stamp matters.
- What does a cheap regression set look like for this kind of client?A small pinned collection of the attempt inputs behind the accepted findings, with recorded expected outcomes and attempt counts, runnable on a schedule and on any provider or configuration change, reporting a rate per finding rather than pass or fail.
Concluding a finding is fixed because ten re-tries all refused is like concluding a die has no six after ten rolls: with a one-in-ten behaviour, a clean run of ten happens by chance about a third of the time.
saying these in an interview costs you the question
- Writes findings with no model identifier, configuration or date, then defends them when they no longer reproduce.
- Recommends 'the model should refuse this' to a client who cannot change the model.
- Treats a point-in-time report as valid indefinitely for audit purposes.
- Records a probabilistic finding as a binary success with no attempt count or rate.
- Has no view on who notices when the provider changes the deployment.