A service runs from an immutable image identified by digest, but its environment configuration is applied separately and always from the config repository's HEAD. Redeploying the previous image digest did not restore the previous behaviour. Why, and what would you change so that one action reverts the whole release?
answer
- a release is more than an image
- only versioned things can be reverted
- config from HEAD does not move
- pin digest and config revision together
- one revert of one desired-state record
basics
~20 sOnly the code was versioned. Configuration applied from a moving HEAD stays at its new value, so half the release is still live after the rollback. Bind the artifact digest and the config revision into one versioned release record you revert as a unit.
solid answer
~50 sA release is not just an image. It is code plus the configuration, schema and assets that shipped alongside it, and only the parts you version can be rolled back. If configuration comes from a repository's HEAD, or from a console someone edited by hand, then redeploying the previous digest reverts the code and leaves the timeout, pool size or routing change exactly where it was - and if the bad change *was* the config, the rollback does nothing at all. The fix is to make the deployable unit a single versioned record that pins the artifact digest and the configuration revision together, so "roll back" means "restore the previous revision of that record" - one revert, one action, both halves. Anything genuinely runtime-mutable, such as flags, is deliberately decoupled and must have its own audited, revertible history. Then rehearse it, because the failure mode only shows up when you try.
go deeper
Be able to say that a release is code plus configuration, and that redeploying an old image leaves a separately applied configuration change untouched.
Explain concretely how to bind the artifact digest and the configuration revision into one versioned record so a single revert restores both, and why a moving HEAD makes the running pair time-dependent.
Show that you rehearse rollbacks outside incidents and that you can name what still does not come back - cached client assets, published messages, written data - and how you keep the previous version compatible with them.
Own the standard that any change reaching production goes through a versioned, revertible path, and be ready to defend which levers stay deliberately decoupled and what audit and drill requirements you attach to them.
## The release is bigger than the artifact When a team says "we can roll back", they usually mean "we can redeploy the previous image". But a release is everything that changed in production as part of shipping that version: - the **code artifact** (image, package, bundle), - the **configuration** it reads - timeouts, pool sizes, endpoints, routing weights, resource limits, - **schema and data** changes, - **client-side assets** already cached in browsers or CDNs, - **messages already published** in a new format, - and any **runtime state** the new version created. A rollback reverses only the parts that are versioned and coupled to the thing you are reverting. Everything else stays where it is. The specific failure in the question is the most common one: the image is content-addressed and precise, and the configuration is not versioned with it at all. ## Why moving-HEAD configuration breaks rollback If configuration is rendered from whatever the config repository currently holds, then the pair (code, config) that actually runs is a function of *when* you deploy, not of *what* you deployed. Two consequences follow: 1. **Rollback is incomplete.** Redeploying the old digest gives you old code against new configuration - a combination that may never have been tested and can be worse than either release. 2. **Rollback can be a no-op.** If the regression came from a configuration change (a timeout cut from 5s to 500ms, a connection pool halved, a traffic weight moved), the image is innocent and rolling it back changes nothing. Teams then conclude "the rollback didn't work" and start debugging the wrong layer, which is expensive precisely when it is most expensive. The same reasoning applies to configuration edited directly in a cloud console or an admin UI: there is no revision to return to, only somebody's memory of the previous value. ## Making the release one revertible unit The goal is that a release has a single identity, and that identity resolves to everything the release changed. Practically: - **Pin both halves in one record.** The deployable unit is a manifest that names the artifact digest *and* the exact configuration revision - not "latest config", a specific commit or version. Deploying release N means applying that record; rolling back means applying record N-1. - **Store it as versioned desired state.** If your production state is expressed as a repository or an equivalent versioned store, the rollback is a revert of one commit that moves both digest and config together, and the whole change is visible in one diff. - **Make config changes go through the same path as code.** A configuration change that skips the release process is a deploy with no rollback story. Reviewed, versioned, and rolled out the same way is what makes it reversible. - **Treat runtime-mutable levers deliberately.** Some things are *supposed* to change without a deploy. That decoupling is their purpose, but it also means the artifact rollback will not move them - they need their own change log and their own way back. ## The parts that still will not come back Even a perfectly coupled release record does not undo everything: - **Client-side assets** already downloaded keep running in users' browsers until their cache expires, so a rolled-back server must still tolerate the previous release's clients. - **Messages already published** in a new format sit in queues and topics; the rolled-back consumer has to be able to read or safely park them. - **Data written by the new version** stays written. The practical rule is that a rollback restores your deployment, not the world's memory of it. Design each release so the previous version can survive contact with whatever the new one left behind - that compatibility, not the deploy tooling, is what makes rollback safe. ## How to know it works This class of defect is invisible until you exercise it. Roll back a release deliberately in a non-emergency, and check that both the code version and the effective configuration values match the previous release. Teams that only try during an incident discover the coupling gap at the worst possible moment.
- If the regression came from a configuration change rather than the code, how should the rollback differ?It should not differ - that is the point of coupling them. If the release record pins both, restoring the previous record reverts whichever half was at fault without you needing to know which. When config is decoupled, you first have to diagnose the layer before you can act, which adds minutes to an incident where the whole value of a rollback is that it needs no diagnosis.
- Some configuration is deliberately runtime-mutable and does not roll back with a deploy. How do you keep that from being a hole?Give it the properties the deploy path already has: every change recorded with who, what and when, a previous value you can restore in one action, and change windows visible on the same timeline as deploys. The decoupling is intentional and useful, but it means the artifact rollback does not touch it, so those levers need their own audited way back and need to be drilled.
- Why does redeploying an old release not help with assets already sitting in users' browsers?Those bytes were delivered and cached before the rollback; a server-side change cannot recall them. Until the cache expires or the client reloads, the rolled-back backend is serving the previous release's front end. That is why servers must stay compatible with recently shipped clients, and why asset caching lifetimes are effectively part of your rollback window.
saying these in an interview costs you the question
- Rolling back the image rolls back the entire release
- Config changes are low risk, so they need no version history
- We can just remember the old value and set it back
- A config edit in the console is not a deploy
- Rollback undoes everything the release did, including data