A provider offers both a floating alias for a chat model and a pinned snapshot identifier of the same model. Which of the two should a recurring safety-benchmark suite be pointed at, and what does each choice cost you?
answer
- alias measures risk, pin measures the instrument
- control run on a subset
- control moves means harness moved
- pins get retired, re-baseline
- pin does not freeze the filter stack
basics
~20 sPoint at both if you can. The alias is what production calls, so running it is your drift detector. The pinned snapshot is a control: same weights every quarter, so a moved rate there means your harness or judge moved. Costs are double the queries, and pinned snapshots eventually get retired.
solid answer
~1 minThey answer different questions, so the choice depends on what the suite is for. - **Floating alias** — measures what your users actually reach. It is the only run that detects a silent model or filter change, and it is the number you would put in a risk report. Its weakness is that it is a moving target: every quarter-over-quarter comparison mixes model drift with your own harness drift. - **Pinned snapshot** — reproducible. Because the weights are held fixed, any movement in its rate is attributable to your harness, judge or sampling. That makes it a **control run** rather than a risk measurement, and it is also what a published result should cite if it is meant to be reproducible. The practical arrangement is a small pinned control run plus the full alias run: the control validates the instrument, the alias measures the risk. Costs: roughly double the query budget on a metered endpoint unless you cut the control to a subset, and pinned snapshots are retired by providers, so a long-lived control eventually vanishes and the baseline has to be re-established. Whichever you use, the manifest must say which — 'same endpoint' is not a claim you can make from the alias alone.
go deeper
Knows a pinned identifier is reproducible and a floating alias may change, and that production uses the alias.
Explains the two-run arrangement: alias for risk, pinned subset as a control that catches harness and judge drift.
Sizes the control against a metered query budget, plans for pin retirement and re-baselining, and knows a pin does not freeze provider-side filtering.
Sets the standard: every recurring safety suite carries a control, the manifest names the target precisely, and series discontinuities are recorded rather than smoothed.
### The idea underneath: separate the instrument from the subject A recurring safety suite has two jobs that pull against each other. One is to detect change in the thing being measured — the endpoint your users actually reach. The other is to prove the measuring apparatus still reads true, so that when the number moves you can say the *subject* moved. A single target cannot do both, and that, rather than any property of the provider's naming scheme, is why the answer is usually "both, sized differently". ### The floating alias — the risk measurement A floating alias is the product-level name production calls. Whatever the provider is currently routing that name to answers the request, including any serving-side filtering — an input or output classifier in front of the model — that the account has enabled. It is therefore the only run whose number describes deployed exposure, and the only run in which a silent model update or a newly enabled filter can show up at all. Its weakness is that it is a moving target: if the alias run moves and nothing else is held fixed, you cannot say whether the subject changed or your own instrument did. ### The pinned snapshot — the control A pinned snapshot identifier holds the weights fixed across quarters. Its rate should therefore be stable up to sampling noise. If it moves, you changed something: the harm judge's revision or threshold, the prompt template, temperature or `max_tokens`, the behaviour subset, the error-counting convention. That makes it a **control run**, an instrument check, not a statement about risk. It is also the target a published result should cite if the result is meant to be reproducible by anyone else. ### How to size the pair | | target | scope | what it answers | |---|---|---|---| | risk run | floating alias | full behaviour set | what is our current exposure | | control run | pinned snapshot | small fixed subset | is the instrument still reading true | A control does not need coverage; it needs sensitivity to harness change. A tenth of the behaviours, fixed and never edited, run in the same job with the same judge, is usually enough. The operating rule: **if the control moved, stop and fix the instrument before reading the alias result at all** — otherwise you will spend the quarter explaining a number the harness invented. ### What it costs Running both roughly doubles the query budget only if the control is full-size, which is why it should not be. Concretely, a 200-behaviour alias run at five attempts is about 1,000 target completions plus 1,000 judge calls; a 20-behaviour control adds about 100 of each, a ten percent surcharge, plus a few minutes of wall clock. Against that, the control buys you the ability to attribute movement at all, which otherwise costs an analyst days of retrospective archaeology. The larger recurring cost is human triage of alias-run hits, and the control does not add to it. ### Where each choice misleads - **Reporting the pinned run as production exposure.** The most common error. Users do not call the pin; whatever the alias resolves to today answers them, filters and all. A pinned number presented as risk is a statement about a build nobody in production is talking to. - **Reading alias movement as model change.** It bundles model change, filter change, judge drift and harness drift. Without a control you have no way to peel any of them off. - **Believing a pin freezes everything.** It freezes weights, not the serving stack. Provider-side filtering can be changed in front of a frozen model, so a moving control is not automatically your bug — check your own manifest diff first, then consider this. - **Splicing across a retired pin.** Providers retire snapshots on their own schedule, and the control will eventually break. Re-baseline deliberately and record the discontinuity in the tracker rather than drawing one continuous line through it. ### What you would check That the manifest names which target each number came from — "we tested the endpoint" is not reproducible. That the control subset has not been quietly edited (edit it and it stops being a control). That any served-revision identifier the provider returns is captured for both runs. And where no pinned option exists at all, that a locally held open-weights model, whose artefact you can hash, is standing in as the harness control: it says nothing about the hosted endpoint, but it keeps the instrument honest.
- Your control run's rate moved but the alias run's did not. What now?Suspect the harness or judge first: diff the manifests and re-score old transcripts. Until the control is explained, the alias number is not trustworthy either, since both used the same instrument.
- The provider offers no pinned option. What is the nearest substitute?A locally-run open-weights model held constant as a harness control, plus pinning your judge and dataset revision. It controls the instrument, though it says nothing about the hosted endpoint.
The pinned snapshot is the known-reference sample a testing lab runs alongside the real ones each morning. Nobody cares about its result for its own sake; it exists so that when a real sample reads oddly, you already know the assay was working.
saying these in an interview costs you the question
- Tests only a pinned snapshot and reports it as production exposure.
- Tests only the alias and treats quarter-over-quarter movement as model change.
- Assumes a pinned identifier freezes provider-side filtering too.
- Does not record which of the two the run used.
- Splices a series across a retired pin as if it were continuous.