skip to content

When a claim-severity model must be rolled back to the previously promoted version, what must the registry already hold for that to take minutes?

level: seniorimportance: must knowfreq 80%

answer

  1. roll back the bundle, not the file
  2. artifact, input contract, environment lock
  3. promotion is a pointer resolved at load
  4. retention can delete your fallback
  5. done when every instance reports the digest

basics

~20 s

Rollback is fast only when the previous version is still fully resolvable: retained artifact bytes under a verified digest, the feature-definition and input-schema version it was trained against, its environment lock, and a promotion pointer that serving resolves at load - so the act is a pointer flip, not a rebuild.

solid answer

~50 s

Treat a model version as a **bundle**, not a file. To go back in minutes the registry must already hold, for the previous version: the artifact bytes retained under a digest that is verified on load; the feature-definition version and input schema it was trained against; its environment lock; and its evaluation record. Serving must resolve the promoted version through a **pointer read at load**, so reverting is a registry write plus a reload rather than a build or a retrain. The classic failure is rolling the artifact back while the candidate's new feature definitions stay live: the old model then receives columns it never saw, or loses ones it needs, and for claim severity it will happily return plausible-looking numbers rather than erroring. The second classic failure is a retention job that deleted the fallback.

code

pseudocode · 19 lines
pseudocode
function activate(model, targetVersion):
    record = registry.lookup(model, targetVersion)
    if record.stage != "servable":
        fail("version not servable: " + targetVersion)

    bytes = artifactStore.fetch(record.artifactDigest)
    if sha256(bytes) != record.artifactDigest:
        fail("artifact digest mismatch")

    features = servingRuntime.featureDefinitions()
    if features.version != record.featureDefinitionVersion:
        fail("feature definitions at " + features.version +
             ", version expects " + record.featureDefinitionVersion)
    if features.schemaHash != record.inputSchemaHash:
        fail("input schema does not match trained contract")

    servingRuntime.load(bytes)
    registry.setPointer(model, "production", targetVersion)
    emit("activated", model, targetVersion, record.artifactDigest)

go deeper

for a junior

Know that going back to the previous model means the previous model must still exist: its bytes kept, its version still listed, and serving able to pick it without anyone rebuilding or retraining anything.

for a middle

Explain why a model version is a bundle - artifact, feature-definition version, input schema, thresholds - and why reverting only the artifact can leave a model fed inputs it was never trained on.

for a senior

Show the mechanism and the proof: a promotion pointer resolved at load, digest and contract checks before serving, reference-driven retention that protects the fallback, and a rehearsed, timed rollback drill.

for a principal

Set the guarantee. The published rollback time is what makes promotion a cheap, reversible decision; the cost is retention and the discipline of versioning inputs alongside models, and it is worth paying explicitly rather than discovering during an incident.

## Rollback is a property you build, not a decision you make By the time someone decides to roll a claim-severity model back, every input to the decision is already fixed. The question is not *should we* but *can we, and how long does it take*. That duration was determined months earlier by what the registry retained and by how serving resolves a version. A useful framing: a rollback should be the **same mechanism as a promotion, pointed backwards**. If going forward is a registry pointer write and a reload, going back is too. If going forward involves a build, a retrain or a hand-edited configuration, going back involves all of them, under pressure, at whatever hour the problem surfaced. ## What the previous version must still resolve to 1. **Artifact bytes**, retained and addressed by content digest, verified when loaded. 2. **The input contract** - the feature-definition version and the input schema hash the model was trained against. This is the part teams forget, and it is the part that causes silent wrongness. 3. **The environment lock** for that version, so a serving image can be rebuilt if the running one is also implicated. 4. **The evaluation record** - which set, which metric, which value - so the decision to go back is defensible afterwards. 5. **A servable stage**, meaning retention and archival policy have not moved the version into a state whose artifact was dropped. ## The failure that does not error The candidate introduced two new features and renamed a third. It is promoted, then found to be over-reserving. The artifact is rolled back to the previous version. The feature definitions were not rolled back, because they live in a different system and nobody pinned them to the model version. Now the previous model is fed an input vector it never saw: - a renamed column arrives as missing and is imputed to a default; - a new column is ignored or, worse, shifts positional ordering; - nothing raises an exception, because the shapes still line up. The result is a model producing numbers in the right range that are systematically wrong. Error rates stay flat, latency stays flat, and the only signal is in the predictions themselves. This is why the registry entry must pin the input contract, and why a rollback validates the contract before it loads. ## Rollback scope: what actually has to revert | what changed with the candidate | reverts with the artifact? | what to do | |---|---|---| | model weights | yes | retained artifact, verified digest | | feature definitions or transforms | no | pin the version, revert together or keep both live | | input schema | no | store a schema hash on the version and check it | | decision threshold or post-processing | no | version it as part of the bundle | | serving image or runtime | no | keep the environment lock, redeploy if implicated | | the training data snapshot | not applicable | it produced the artifact; it is not needed to serve it | The first row is the only one people plan for; rows two to four are where incidents live. ## Retention is a rollback control An age-based retention policy is not safe here, because age is uncorrelated with whether a version is a live fallback. Drive retention by **references**: the currently promoted version, the previous promoted version, anything an open approval or investigation points at, and anything a compliance window requires are all exempt; everything else ages out. A store asked to delete a referenced digest should refuse rather than comply. ## Prove it before you need it - **Rehearse.** Roll a promoted version back to its predecessor on a schedule, on the real path, and record how long it took. Publish that number as a property of the registry. - **Check completion, not initiation.** The rollback is done when every serving instance reports the expected digest, not when the pointer was written. A partially reloaded fleet serves both models at once. - **Make the reverse trip cheap too.** After a rollback you will want the candidate back once fixed; if that requires a fresh build, the incident has a long tail. - **Record the rollback as an event** on the version, with who and when, for the same reason a promotion is recorded: the next person needs to know the currently promoted version is not the newest one. A registry that satisfies all of this turns the worst moment of the quarter into a two-line change and a reload - which is also what makes a promotion approvable in the first place, because an easily reversible decision is a cheap decision.

  • What if the candidate also changed the input schema in a backwards-incompatible way?
    Then the rollback is not a pointer flip and you need to know that before promotion, not during the incident. Either keep the previous feature definitions live in parallel so both versions remain servable, or promote the schema change as its own versioned step that can be reverted independently. Pinning a model version to a feature-definition version is what surfaces the incompatibility early.
  • How do you know the rollback path works before you need it?
    Exercise it. Roll a promoted version back to its predecessor on a schedule, on the real serving path, and time it end to end - from the registry write until every instance reports the expected digest. Publish that duration. A rollback that has never been executed is a claim about the system, not a capability of it.
  • What should retention do with the previously promoted version?
    Exempt it explicitly. Retention must be reference-driven: the promoted version, the one before it, anything an open approval or investigation references, and anything inside a compliance window are all pinned, while unreferenced candidates age out on a stated schedule. Deleting a referenced digest should be refused by the store rather than allowed by policy.

Restoring last year's recipe is useless if the kitchen has since relabelled every ingredient jar: the instructions run to completion and the dish is quietly wrong.

saying these in an interview costs you the question

  • Rolls the model artifact back but leaves the candidate's feature definitions live
  • Calls retraining the previous configuration a rollback
  • Assumes the previous artifact is still there because nobody deleted it deliberately
  • Keeps the previous version but not the input schema it expects
  • Treats the rollback as complete when the registry pointer is written
  • Plans rollback of the weights only, ignoring thresholds and post-processing