skip to content

Rolling the dropout early-warning model back to its previous version mid-incident: what must that rollback cover beyond the trained artifact?

level: middleimportance: must knowfreq 68%

answer

  1. a version is a set, not a file
  2. artifact, features, threshold, contract
  3. flip a pointer, do not rebuild
  4. the old version must still load
  5. code rollback is not model rollback

basics

~20 s

A model version is a pin set, not a file: the artifact plus the preprocessing and feature-definition versions it reads, its decision threshold and its output contract. Roll back all of them together, by flipping one pointer to a retained, still-loadable version.

solid answer

~50 s

The trained artifact is only one of the things the newer version changed. It arrived with preprocessing code, a feature-definition version, a decision threshold and often a different output shape, and restoring the artifact alone leaves a mismatched combination that nobody ever evaluated. Treat the version as a **pin set** and flip all of it at once. Two preconditions have to already exist for that flip to work: the older artifact must still be retained and loadable rather than aged out, and the features it reads must still be computed - if the newer version's feature definitions replaced the older ones, the restored model has no inputs. Note also that the model is usually fetched at runtime by a version pointer, so redeploying the previous service release rolls back the code and leaves the model exactly where it was.

code

json · 19 lines
json
{
  "service": "dropout-early-warning",
  "active": {
    "modelVersion": "2026-03-11.4",
    "featureDefinitionVersion": "term-features-v7",
    "preprocessingVersion": "prep-2026-03",
    "decisionThreshold": 0.62,
    "outputContractVersion": 3,
    "weeklyCaseloadCap": 400
  },
  "rollbackTarget": {
    "modelVersion": "2026-01-08.2",
    "featureDefinitionVersion": "term-features-v6",
    "preprocessingVersion": "prep-2025-11",
    "decisionThreshold": 0.55,
    "outputContractVersion": 3,
    "weeklyCaseloadCap": 400
  }
}

go deeper

for a junior

Know that the trained artifact and the service code are deployed separately, so rolling back one of them usually leaves the other exactly where it was.

for a middle

Explain the pin set: artifact, preprocessing, feature-definition version, threshold and output contract move together, because only that combination was ever evaluated.

for a senior

Demonstrate the preconditions - retained loadable artifacts, feature definitions added rather than replaced, a flip that needs no build - and name when rollback is the wrong lever entirely.

for a principal

The tradeoff is retention cost and the discipline of never mutating a feature definition, paid continuously against a rollback capability used a few times a year.

## What "the model version" actually names Ask three engineers what they would roll back and you get three answers: the code revision, the trained artifact, the feature-definition version, or the version stamped on decisions already made. In a serving tier, the thing that produces a prediction is none of these alone. It is a **pin set** - a small bundle of things that were evaluated together and are only known to be correct together: - the **trained artifact** (weights or tree structure, plus whatever the loader needs); - the **preprocessing** that turns a raw request into the model's input, which usually ships as code; - the **feature-definition version** the online path reads, which decides what `attendance_rate` means this month; - the **decision threshold** that turns a score into a flag, tuned per version against a target caseload; - the **output contract** the downstream consumer parses. A rollback that restores only the first item produces a combination that was never evaluated: last quarter's artifact scoring this quarter's feature definitions through this quarter's threshold. That is a new, untested configuration created in the middle of an incident. ## Rolling back a pointer, not rebuilding a build The fast rollback is a **pointer flip**: the serving configuration names an active pin set, you set it to the previous one, and replicas pick it up. It is not a retrain, and it should not be a build. Anything that requires compiling, re-fitting or re-deriving features puts a pipeline between you and the mitigation, and pipelines fail at exactly the wrong moment. The common trap is that code and model roll back independently: | What moved | Rolled back by | What is left behind if you forget | |---|---|---| | Service code | Redeploying the previous release | The pointer still resolves to the newer model | | Model artifact | Flipping the version pointer | New preprocessing still runs on old weights | | Feature definition | Restoring the prior definition version | Online values silently change meaning | | Decision threshold | Restoring it with the pin set | Caseload volume jumps even on a correct model | Because of the first row, "we redeployed the last good release and the predictions did not change" is a routine and confusing incident moment. It is the expected outcome whenever the model is fetched at runtime by a pointer. ## Preconditions that must already exist A rollback is a capability you build before the incident, not a command you invent during one: 1. **The prior version is retained and loadable.** Retention policies that age out artifacts, and loaders that only accept the current serialisation format, both turn a rollback into a retrain. 2. **Its feature definitions are still computed.** If the newer version replaced a definition rather than adding one beside it, the older model's inputs stopped existing. Adding definitions rather than mutating them is what keeps the previous version runnable. 3. **The pin set is stored as one unit.** If the threshold lives in a config file, the features in a platform, and the artifact in a store, nothing guarantees they move together. 4. **The flip needs no build.** One configuration change, effective on the next request or the next scheduled run, with a known propagation time you have measured. 5. **Someone reachable can do it.** A rollback only the model's owning team can perform is unavailable at the hours incidents favour. ## When a version rollback is the wrong lever Rolling back reverses a change **you** made. It does nothing when the cause is upstream and shared: a calendar feed that changed meaning, a schema that shifted, a source system that stopped emitting an event. Every retained version reads that same broken input, so every one of them is wrong in the same way, and the pointer flip buys nothing but a false sense of mitigation. The tell is simple and worth checking before you flip: *did anything about this model change, or did the world under it change?* If the last model change predates the suspected start, the lever you want is not the prior version - it is switching off the model path entirely and serving from something that does not read the broken input. The second limit is scope. A rollback restores prediction quality going forward. It does not touch decisions already delivered, and it does not re-score the weeks that went out wrong. Those are a separate, deliberate remediation step, and skipping it is how an incident gets closed while its effects are still working their way through an advising team.

  • The previous service release was redeployed but the predictions are unchanged. What happened?
    The model is almost certainly fetched at runtime by a version pointer rather than baked into the release, so the redeploy restored the code path and left the pointer resolving to the newer artifact. They are two independent rollbacks. The fix is to flip the pointer as well - and, if the older code expects the older preprocessing, to flip both in the same change rather than one at a time.
  • The rollback completed but the flags are still wrong. What does that tell you?
    That the cause is not the model version. Either a precondition was missed - the older feature definitions are no longer computed, so the restored model is reading substituted or stale values - or the failure is upstream of every version, in an input all of them share. Check whether any model change predates the suspected start; if none does, stop flipping versions and switch the model path off.

saying these in an interview costs you the question

  • Rolling back the service release rolls back the model too
  • A rollback means retraining the previous version from scratch
  • Restoring the artifact alone restores the previous behaviour
  • The old threshold can stay while the old artifact comes back
  • A retention policy that deletes old artifacts is harmless