skip to content

Your AI red-team findings were filed against a hosted chat application whose endpoint URL and request contract never change. Which changes underneath that unchanged endpoint can silently invalidate those findings, and why does nothing alert you?

level: juniorimportance: must knowfreq 70%

answer

  1. stable URL, moving model
  2. model swap, guard bump, prompt edit
  3. no build number to diff
  4. results carry a config fingerprint
  5. drift cuts both ways

basics

~20 s

Three: the provider swaps the model behind the same endpoint alias, the safety classifier in front of it is upgraded, or the product team edits the system prompt. None changes the URL or the contract, so no deploy alert or failing test fires. Your filed pass and fail results simply stop describing the live system.

solid answer

~50 s

The AI stack has three change surfaces that move behaviour without moving any interface: - **A provider model swap.** A floating model alias can be repointed at a new build. Refusal behaviour and instruction-following shift, so an attack you filed as blocked may now land. - **A guardrail version bump.** The input or output classifier in front of the model is a separate service on its own release train; new thresholds change what reaches the model. - **A system-prompt or context edit.** A one-line instruction change is a config edit, often not even in the application repo, and it can reopen a whole class of behaviour. None of these produce a signal in your tooling: the harness sends the same requests to the same address and gets a well-formed response back. Findings age silently, and drift runs both ways — a fixed finding regresses, or an open one closes with nobody having remediated it.

go deeper

for a junior

Names the three silent change surfaces — model behind the alias, guardrail version, system prompt — and says why no alert fires.

for a middle

Adds that drift runs both ways, and that a result is only meaningful next to the configuration it was measured against.

for a senior

Turns it into practice: fingerprint every filed result, void it on fingerprint change, subscribe to the config pipeline, run canaries for the changes nobody announces.

for a principal

Frames it as an ownership problem — three change surfaces owned by three teams, none of whom currently owe the security programme a notification — and negotiates pinning and notice.

### The unit under test is a configuration, not an address In ordinary application security a finding is attached to a build. There is a commit hash, an image tag, a release number, and rerunning the test against that identifier tests the same artefact. An LLM feature breaks that assumption at the root. The part you can address — hostname, path, and the JSON request body with its `messages` array — is deliberately frozen, because the whole value of the interface is that callers never have to change. The behaviour you actually measured lives behind that frozen surface, in three components that version independently, on three release trains, owned by three teams, none of whom currently owes your programme a notification. **The served model.** The request body's `model` field usually carries a *floating alias* — a family name the provider is free to repoint at a newer build whenever it likes. When they do, your harness sends the identical string to the identical URL and gets an identically shaped response back. A model update re-trains the refusal boundary along with everything else, so an attack you filed as blocked can begin to land, and one you filed as open can stop working. Providers also ship same-name weight updates, where even the version string that does exist never moves. **The guardrail layer.** The input classifier, output classifier or rule framework sitting in front of the model is a separate service with its own version and its own thresholds. A decision threshold moved from 0.8 to 0.6 changes which of your attack prompts reach the model at all; a fail-open-to-fail-closed change decides what happens when that service times out. Neither touches the application's build number. **The prompt and its context.** The system prompt, the tool descriptions handed to the model, and the retrieval corpus behind it are *data*. They are commonly edited through a config console or a content store that sits outside the code-review gate, leaving no artefact and no hash. One added sentence can reopen a class of behaviour that took a week to close. ### Why nothing fires Every monitor pointed at this system watches availability and contract: HTTP status, schema validity, latency, error rate. A swap keeps all four flat by design. Your own smoke test passes, because it asserts that a well-formed response comes back, not that the response still *means* the same thing. The product team's regression tests assert on structure, so they stay green too. The absence of an alert here is not evidence — it is the expected output of instruments that cannot see this class of change. ### Where the number misleads The dangerous reading is directional. Everyone anticipates "something we fixed regressed"; almost nobody plans for "an open finding silently closed". The second is equally damaging: the report still says the finding is open, an engineering team is funding a fix for a behaviour that no longer reproduces, and when they fail to reproduce it, the discount lands on your programme's credibility rather than on the finding. And that closure is not a control anyone owns — the next repoint can reopen it with nobody watching. The other misleading reading is the pass itself. A filed "blocked, 0 of 20" describes one configuration on one date; read six months later as "this attack does not work here", it is a claim the evidence never supported, and it was already a weak claim on the day it was written. ### What it costs to stay honest Revalidation is metered. A modest suite of 60 attack prompts at 20 trials each is 1,200 requests per pass; multi-turn probes multiply that by the turn budget, and any model-scored judging adds a second call per attempt. Analyst triage, not tokens, is usually the larger line. That arithmetic is exactly why the answer is not "rerun everything whenever anything changes" — it is to know precisely which filed results a given change put in doubt. ### What you check Stamp every filed result with a configuration fingerprint: the `model` value requested, whatever model or version field the response envelope returned recorded verbatim, a SHA-256 of the system prompt and the tool definitions, the guard names, versions and thresholds, and the decoding parameters. Void a result when its fingerprint stops matching production, not when the calendar expires. Ask the platform team to call a pinned dated identifier so a swap becomes a pull request; subscribe to the repository or config pipeline that carries prompt and guard changes; run a small benign canary set continuously for the swaps nobody announces. And never ask the model which model it is: that answer is generated text, shaped by the very system prompt you are trying to verify.

  • Why is a finding that silently CLOSES also a problem, not just good news?
    Because your report says it is open and someone is funding a fix for it, and because the change that closed it was not a control anyone owns — the next swap can reopen it with nobody watching.
  • What single field would you push the platform team to expose to make this tractable?
    A resolved configuration fingerprint on the serving path — the concrete model identifier actually used, the guard versions and a hash of the system prompt — returned or logged per request, so runs can be diffed instead of guessed at.

The endpoint is a street address, not a tenant. The sign on the door never changes, so your notes about who lives there quietly stop being true, and no letter ever arrives marked “new occupant”.

saying these in an interview costs you the question

  • Assuming a stable endpoint and a passing smoke test mean the tested system is unchanged.
  • Treating a filed red-team report as valid until its calendar expiry regardless of stack changes.
  • Only worrying about regressions and ignoring that a finding can silently close, making the report wrong in the other direction.
  • Believing the model can be asked which model it is, and trusting the answer as a version check.

context