skip to content

Anthropic retires the Claude Opus snapshot your service pins — how do you run the migration?

level: principalimportance: should knowfreq 30%

answer

  1. a notice starts a clock
  2. retired IDs fail, they do not fall back
  3. find every place the string is written
  4. measure the successor before switching
  5. one owned mapping beats twelve scattered pins

basics

~20 s

Treat it as a dependency upgrade with a hard deadline: inventory every place the retired Opus snapshot ID appears, run a golden-set eval against the successor snapshot, fix prompt and parsing regressions, canary the new pin, and finish before the shutdown date, because retired IDs return errors rather than falling back.

solid answer

~50 s

A retirement notice starts a clock, not an emergency. First, inventory: every service, prompt template, batch job, eval harness and config layer that names the retired Opus snapshot — model IDs sprawl further than people expect. Second, measure: run your golden set against the successor snapshot and diff on the things that actually break, which are rarely benchmark scores — output formatting, tool-call eagerness and argument filling, refusal edges, and response-length distribution that moves p95 latency. Third, fix what regressed, in prompts and parsers rather than by pinning harder. Fourth, roll out behind a flag or percentage canary with the old pin still reachable, and watch business metrics, not just error rates. Finish well before the date: after retirement a request naming that model fails outright — the API does not silently route you to a newer Opus.

go deeper

for a junior

Know that a pinned model version does not live forever, that Anthropic publishes retirement dates in advance, and that calls naming a retired model fail rather than being served by a newer one.

for a middle

Explain the mechanics of the switch: find every place the model ID is written, run the existing prompts against the successor, and look for changed output formatting or tool-call behaviour before shipping the new pin.

for a senior

Show the rollout discipline — golden-set eval diff, fixes made in prompts and parsers rather than in pinning, flag or canary with the old pin reachable, and product-level metrics watched because regressions rarely surface as errors.

for a principal

Own the governance: one logical model name with a single accountable owner, an evidence gate for any pin change, deprecation notices routed to a human with authority, and a rehearsal habit so the next retirement is a mapping edit rather than a project.

## What retirement actually means A dated Opus snapshot is immutable but not eternal. Anthropic publishes deprecation notices with a retirement date, and after that date a request naming the retired model does not quietly get served by something newer — it fails. That is the correct design, because a silent substitution would swap the weights under a production workload without anyone noticing, but it means the deadline is real and a missed migration is an outage, not a degradation. The first instinct of many teams — switch to a floating alias so this can never happen again — trades a scheduled, testable migration for an unscheduled, untested one. It converts a date you control into a surprise you do not. ## Step one: inventory Model identifiers sprawl. Beyond the obvious service config there are usually batch and offline jobs, evaluation harnesses, notebooks that became load-bearing, infrastructure-as-code defaults, per-tenant overrides, feature-flag payloads, and documentation and runbooks that will send someone down the wrong path later. Grep the organisation, not the repository. The output of this step is a list with an owner per entry, because the migration's real cost is coordination, not code. This is also the moment to notice how many distinct pins you carry. A fleet with one logical name resolving to one ID migrates in an afternoon; a fleet where twelve teams each chose their own is where the schedule goes. ## Step two: measure before you change anything Run the successor snapshot against a golden set that reflects your traffic — real prompts, real documents, real tool schemas — and diff against the incumbent. Aggregate quality scores are the least informative signal here. Watch instead for: - **Format drift.** Added preambles, changed heading or list style, fenced output where bare JSON was expected. Anything downstream that parses text is where breakage concentrates. - **Tool-call behaviour.** Different eagerness to call, different handling of optional arguments, different batching per turn. In an agent loop this changes the number of round trips and therefore latency and cost, even when each call is individually valid. - **Refusal boundaries.** Prompts near the line can flip, creating a whole class of empty or hedged responses in one workflow while everything else looks healthy. - **Length and latency distribution.** Longer average responses move p95 and can trip timeouts that were comfortable before. Record these as an artefact. The same harness is what you will run for the next migration, so the investment compounds. ## Step three: fix in the right layer Regressions belong in prompts, output schemas and parsers — the layers you own — not in more elaborate pinning. If a downstream parser only worked because a particular build never emitted a preamble, the parser was always fragile and the migration merely revealed it. Prefer structural fixes: explicit output contracts, tolerant parsing, tool schemas that make the required arguments unambiguous. ## Step four: roll out like a dependency upgrade Flag or percentage canary, old pin still reachable, rollback as a config revert rather than a redeploy. Watch business-level metrics — task success, human-escalation rate, downstream error rate — because model regressions rarely show up as HTTP errors. Ramp over days, not minutes, and keep the ability to hold at a partial split. Budget the calendar backwards from the retirement date with real slack. Aim to be fully migrated well before it, so that discovering a regression late leaves room to fix rather than forcing a bad choice at the wall. ## Step five: fix the governance, not just this migration The organisational answer is a single source of truth: a logical model name owned by one team, mapped to a concrete pinned ID, consumed by everyone else. Then a retirement is one mapping change plus one eval run, and the inventory step disappears. Pair it with a policy on who may change a pin and what evidence gates the change, and subscribe someone to the deprecation announcements so the notice reaches a human with authority rather than a shared mailbox. ## What good judgment looks like here The defensible position is not "always pin" or "always float" but "pin, with an owned, rehearsed upgrade path". Pinning without a migration muscle is how teams end up doing an emergency swap the week of the deadline with no evals at all — the worst of both strategies. Rehearse by moving a low-stakes workload to each new snapshot early: it keeps the harness honest and surfaces format drift long before the clock matters.

  • Why is switching everything to a floating alias a poor response to a retirement notice?
    It replaces a scheduled, testable migration with an unscheduled one. The alias will still move — just on Anthropic's calendar rather than yours, without an eval run, a canary, or anyone watching. You trade a deadline you can plan against for a behaviour change that lands silently in production, and you lose the ability to reproduce past responses.
  • What signals tell you a snapshot migration regressed, if error rates stay flat?
    Product-level ones. Task completion or acceptance rate, human-escalation or retry rate, downstream parse failures, tool-call counts per session, and the p95 of response length and latency. Model regressions usually present as successful HTTP calls with subtly worse content, so a rollout watched only on 5xx rates will look perfectly healthy while quality slides.
  • How would you make the next retirement cheap rather than repeating this work?
    Collapse the fleet onto one logical model name owned by a single team and resolved to a concrete pinned ID in one place, keep a golden-set eval harness that any candidate snapshot can be run through, and rehearse by moving a low-stakes workload onto each new snapshot early. The migration then becomes one mapping change plus one eval run.
  • Who should own the decision to move the pin, and what evidence should gate it?
    A single accountable team, with the eval diff as the gate: golden-set results on the successor, an explicit list of format or tool-behaviour changes found, and a canary plan with rollback. Distributing that decision to every consuming team guarantees drift, divergent pins, and a much larger inventory the next time a retirement notice arrives.

saying these in an interview costs you the question

  • Assumes a retired model silently falls back to a newer one
  • Moves everything to a floating alias to avoid future notices
  • Swaps the pin without re-running evals because scores improved
  • Watches only HTTP error rates during the rollout
  • Starts the migration the week of the retirement date

context