A candidate send-time model is rolled back to the previous artifact, yet send hours stay wrong and nothing errors — what else must the rollback restore?
answer
- the artifact is not the release
- two versions move together
- features the candidate introduced
- same type, different meaning
- drain what is already queued
basics
~20 sThe feature definitions the candidate introduced, and the queued work planned under it. Reverting only the artifact leaves the previous model reading features whose meaning changed, which produces plausible wrong hours with no error, and leaves already-queued sends carrying the candidate's decisions.
solid answer
~40 sA model version is not a release on its own. The candidate almost certainly shipped with a changed **feature-definition version** — a recomputed engagement-by-hour aggregate, a new window, a different event filter — and if that pipeline was upgraded in place, the restored model now reads values it was never trained against. Nothing throws: the field is present, well-typed and in a plausible range, so scoring succeeds and the send hour is simply wrong. A rollback therefore has to restore one pinned unit: the model artifact version, the feature-definition version together with its materialisation in the online store, and the serving config including the ramp percentage. It also has to drain or re-plan the sends already queued under the candidate, or the bad behaviour continues for the queue's horizon.
code
json · 13 lines{
"release": "send-hour-2026-09-18",
"modelArtifactVersion": "send-hour-v7",
"featureDefinitionVersion": "send-features-v4",
"onlineStoreMaterialization": "send-features-v4",
"servingConfig": {
"rampPercent": 25,
"rampId": "send-hour-ramp-3",
"fallbackHour": "user-local-09"
},
"queueHorizonHours": 24,
"previousRelease": "send-hour-2026-09-11"
}go deeper
Remember that a model version is served together with the features it was trained on, so putting the old model back does not by itself put the old behaviour back.
Explain the training-serving mismatch a feature redefined in place creates, and why it produces well-formed values that the restored model interprets wrongly.
Define the rollback unit — artifact, feature-definition version and its materialisation, serving config, work in flight — and say how you would have caught the silent version quickly.
Set the platform rule that makes this impossible: additive feature versioning with a retention window, explicit version pinning at read time, and a tested drain path for queued work.
## The release is not the artifact When a candidate model version is trained, it is trained against a specific set of **feature definitions** — how "opens in the last seven days by hour of day" is computed, which events count, what window is used, how missing hours are filled. Those definitions are produced by a feature pipeline and materialised somewhere the serving or planning path can read quickly. A candidate that improves on the incumbent very often does so *because* it was given new or redefined features, so the candidate's release commonly moves two versions at once. If the feature pipeline is upgraded **in place** — the same feature name, new semantics — then reverting the model artifact alone puts the previous model version in front of data it has never seen: - the previous artifact was trained against definition **N**; - the online store now holds values computed under definition **N+1**; - the previous artifact reads them anyway, because they have the right name and the right type. This is a training-serving mismatch created by a rollback, and it is why the symptom in the question is possible at all. ## Why the failure is silent | Signal | What it shows | |---|---| | Error rate | Flat — scoring succeeded on every user | | Latency | Flat — the same fetch and the same forward pass | | Feature null rate | Flat — the field is populated | | Prediction distribution | Shifted, if anyone is watching it | | Send hour correctness | Wrong, and only a human notices | A wrong feature *meaning* does not look like a fault. A value of 0.31 computed under the new definition is as well-formed as a value of 0.31 under the old one; only the model's interpretation of it is now false. Anything that gates on exceptions and percentiles reports a healthy system. ## What one rollback unit has to pin 1. **The model artifact version** — the weights or the tree ensemble actually being served. 2. **The feature-definition version** — plus the pipeline that produces it, and the guarantee that the online store is materialised under that definition again. Until the re-materialisation catches up, the restored model reads stale or missing values and its fallback path takes over. 3. **The serving config** — the ramp percentage, the assignment's ramp identifier, the fallback hour, any threshold on the score. 4. **The queued work** — sends already planned and enqueued carry the candidate's predicted hour and the candidate's stamped version. A rollback that only changes what future planning runs do leaves those sends in flight, so the harm continues for as long as the queue horizon. The practical form of this is a release manifest that names every one of those versions, so "roll back" means "apply the previous manifest" rather than "figure out what moved". ## The design that removes the problem - **Add feature definitions, never edit them in place.** A new definition gets a new version and is materialised alongside the old one for the whole ramp plus a retention window. Then both models can read the definition each was trained against, and reverting the artifact really is sufficient. - **Pin by version at read time.** The scoring path asks for a definition version explicitly rather than for whatever the pipeline last wrote, so a model can never silently be served the wrong semantics. - **Make the queue horizon a stated number.** If sends are planned a day ahead, then a day is the floor on how long a revert takes to be fully effective — unless draining and re-planning is an operation somebody has actually built and tested. - **Watch the prediction distribution, not only errors.** A shift in the distribution of predicted send hours is the one signal that would have caught this quickly, precisely because the error-based signals cannot. ## What to say in the round The checkable answer is: *the artifact is the smallest part of the release*. A model rollback must restore everything the candidate's predictions depended on and everything the candidate already wrote down — the feature definitions and their materialisation, the serving config, and the work in flight. A team that can revert a model version but cannot revert the feature definitions it shipped with has a rollback that works only for the subset of candidates that changed nothing but weights.
- What makes this failure silent rather than an error?The feature is present, well-typed and in a plausible range; only its meaning changed. Scoring succeeds, latency is unchanged and the null rate is unchanged, so every health signal computed from exceptions and percentiles reads clean. Only the distribution of predicted send hours — or a human reading their notifications — reveals it.
- How would you design so that reverting the artifact alone is enough?Version feature definitions additively instead of editing them in place, keep the previous definition materialised for the whole ramp plus a retention window, and have the scoring path request a definition version explicitly. Then each model version always reads the semantics it was trained against, and the artifact becomes the only thing that has to move.
- What happens to sends already queued under the candidate when you revert?They still carry the candidate's predicted hour and stamped version, so they keep delivering the behaviour you just reverted for as long as the queue horizon. Either drain and re-plan them as an explicit operation, or state plainly that the revert is only fully effective after the horizon has passed.
Putting yesterday's recipe back in the kitchen does not help if the pantry was relabelled overnight. The cook follows the old steps successfully, with no complaint and no alarm, and plates the wrong dish.
saying these in an interview costs you the question
- Believes reverting the model artifact version is the whole rollback
- Assumes a wrong prediction shows up as errors or latency
- Edits a feature definition in place during a ramp
- Forgets the sends already queued under the candidate version
- Expects the restored model to read correct features immediately
- Cannot name what a release pins besides the model weights