Nobody tells you when a model provider swaps the model served behind an unchanged hosted endpoint. What can a red-team programme actually do to detect that a swap has happened?
answer
- pin dated id, not floating alias
- diff response version metadata
- benign canary set, fixed decoding
- fingerprint: refusals, length, latency
- never trust self-reported identity
basics
~20 sAsk for a pinned dated model identifier instead of a floating alias, and record whatever version field the response carries, diffing it per run. Where neither exists, run a small fixed canary prompt set on a schedule and watch its behavioural fingerprint: refusal wording, output formatting, length and latency. Drift there means rerun.
solid answer
~50 sThere are three tiers, and you want all three because each fails differently. 1. **Contractual.** Push the platform team to call a pinned, dated model identifier rather than a floating alias, so a swap becomes a pull request instead of a surprise. 2. **Metadata.** Log whatever model or version field the provider returns and diff it between runs. Cheap and exact when it exists — but a gateway may strip it, and a same-name weight update will not move it. 3. **Behavioural canaries.** Keep a small fixed set of harmless prompts, run them on a schedule with decoding pinned as tightly as the API allows, and track a fingerprint: refusal phrasing, formatting, response length, latency. Alert on distribution shift, not on any single response. The trap to name: never ask the model which model it is. Self-reported identity is generated text, shaped by the system prompt, and wrong often enough to be useless.
go deeper
Says to record the model identifier the app requests and check response metadata for a version field.
Adds behavioural canaries with pinned decoding, explains that the signal is a distribution shift across a set rather than one changed answer, and rejects asking the model its own name.
Discusses baselining the fingerprint before alerting, gateway layers that strip metadata, the ongoing metered cost of canaries, and what a hit triggers downstream.
Argues the real fix is contractual — pinned identifiers and provider notice — and treats detection as the fallback for the surfaces where pinning was refused.
You are trying to detect a change that the system is designed not to announce, and there is no ground truth to check yourself against. So you build layers that fail in different ways and accept that none of them is proof. ### Layer 1 — remove the silence rather than detect it The strongest control is not a detector. If the application sends a pinned, dated, immutable identifier in the request's `model` field instead of a floating family alias, a swap stops being a provider decision and becomes a config change in your own repository, with a diff, a reviewer and a date. Getting that change made is a negotiation with the product team about who absorbs the risk of an unannounced behaviour change, and it is worth more than anything you can instrument. The trade the product team will raise is real: pinning means they stop getting quality and price improvements for free, and it means someone must own the migration when the pinned build is retired. ### Layer 2 — diff the metadata you already receive Persist the whole response envelope for at least one request per run, not just the assistant text. Where the envelope carries a model or version field, a per-run diff of that string is an exact swap detector at zero marginal cost. Know its two failure modes: an internal gateway, router or caching proxy may normalise or rewrite that field, so what you are diffing is the proxy's opinion rather than the provider's; and a same-name weight or serving-stack update will not move it at all. ### Layer 3 — behavioural canaries Pick ten to thirty benign prompts whose outputs are cheap, short and stylistically stable — formatting requests, list and arithmetic behaviour, a mild refusal-triggering but harmless request, one long-output prompt. Pin `temperature` to 0, pin `top_p` and `max_tokens`, and pass a `seed` if the endpoint accepts one, recording whether it was honoured. Run the set on a fixed schedule and store the raw outputs. The signal is not any one answer: it is a joint shift in a fingerprint computed across the set — mean and spread of output length, exact-match rate against the previous run, similarity of refusal phrasing, median latency. ### What it costs Twenty canaries once a day is roughly 7,300 requests a year: negligible money, but a permanent line item and a permanent alert surface. The real cost is triage. A threshold tuned before you know the fingerprint's day-to-day variance pages someone weekly, and a canary that pages weekly is muted within a month, after which you have the cost and none of the detection. Budget canaries to the assurance line, not to whichever engagement happens to be open. ### Where the number misleads A canary shift is **not attributable**. A system-prompt edit upstream, a guardrail threshold change, a routing change to a different region or a differently quantised serving stack, and an actual model swap all move the same fingerprint in the same direction. Reporting "the model was swapped" on canary evidence alone is a claim the data cannot support; the honest statement is "the serving path changed". The reverse error is worse because it is invisible: a quiet canary is weak evidence. Two builds in the same family can share length, formatting and refusal style closely enough to leave the fingerprint flat, and your false-negative rate is unknown and, without provider ground truth, unmeasurable. Treat six quiet weeks as "no change large enough to move this fingerprint", never as "nothing changed". And the classic: asking the model to state its own name and version. That output is generated text, conditioned on the system prompt and on training data that predates the current build. It is wrong often enough to be worthless as a version check, and confidently wrong, which is worse. ### What you check Baseline for two to four weeks of quiet before you let anything page. Measure the fingerprint's own variance in that window and set thresholds from it. Confirm whether `seed` is honoured by sending one identical request twice. Compare the envelope you receive against a direct provider call where policy allows, to find out whether a gateway is stripping metadata. When a canary does fire, the output is a trigger, not a finding: rerun the high-severity tier, ask the platform team which of the three surfaces moved, and mark filed results measured under the old fingerprint as stale.
- Why must the canary prompts be benign rather than a few of your best attack prompts?Attack prompts are the thing whose success rate you are measuring; using them as a tripwire runs them constantly against production, burns them into any logging or abuse-monitoring path, and confuses drift detection with the finding itself.
- Your canary fingerprint shifts but the response metadata version string is unchanged. What do you conclude?Not enough to conclude anything alone — it could be a same-name weight update, a routing change, or a system-prompt or guard change upstream of the model. Treat it as a rerun trigger and ask the platform team which of the three moved.
saying these in an interview costs you the question
- Prompting the model to state which model it is and treating the answer as a version check.
- Alerting on any single canary response differing, which fires constantly on a sampling endpoint.
- Assuming an unchanged model name string proves unchanged weights.
- Building an elaborate detector while never asking the platform team to pin a dated identifier.