skip to content

In an immutable infrastructure model, rollback is often described as simply redeploying the previous image. What property of the model makes that possible, and which parts of a real production system does it fail to restore?

level: seniorimportance: should knowfreq 50%

answer

  1. nothing to invert, because nothing was mutated
  2. the artifact still exists, untouched
  3. compute rolls back, state rolls forward
  4. additive first, remove a release later
  5. irreversible effects have already left the building

basics

~20 s

Rollback is cheap because the previous version still exists untouched as an addressable artifact, so you launch it rather than reverse a change. It restores compute and configuration only — not migrated data, consumed messages, external side effects or deleted infrastructure.

solid answer

~50 s

It works because nothing was mutated: the previous image is still on disk exactly as it was tested, so rolling back is launching a known artifact rather than computing and applying the inverse of a change. That matters because many mutations have no inverse — uninstalling a package does not restore the config file it overwrote, and the old version may no longer exist in the repository. What the image does not carry back is everything stateful: schema migrations already applied, data written in the new format, messages other services have already consumed, emails sent and payments taken, secrets rotated, and any infrastructure the change destroyed. So the artifact must stay compatible with the state the newer version left behind — usually via expand-and-contract migrations and one-version backward compatibility — and the previous image has to be retained and still launchable.

go deeper

for a junior

Know that rolling back means launching the previous image again, because that image still exists unchanged rather than having to be rebuilt or un-patched.

for a middle

Explain why there is nothing to invert in an immutable model, and name the preconditions: the artifact is retained, referenced by an immutable identifier, and still able to start against current dependencies.

for a senior

Demonstrate the state boundary in production terms — migrations, consumed messages, irreversible external effects, deleted resources — and describe expand-and-contract plus one-version compatibility as the discipline that keeps rollback viable.

for a principal

Own the policy: how long a version stays a rollback target, what retention and referencing rules enforce it, when fixing forward is safer than rolling back, and how the organisation verifies the path works before it is needed.

## Why the immutable model makes rollback cheap Rolling back a mutable server means computing and applying the *inverse* of a change, and that is harder than it sounds. Uninstalling a package does not restore the configuration file its install script overwrote. The previous version may have been removed from the package repository. A migration script that ran halfway leaves a state that neither the old nor the new code understands. Every rollback is a bespoke piece of work performed under pressure, and it is exercised for the first time during an incident. In the immutable model there is nothing to invert. The previous image was never modified — it is still a distinct, addressable artifact, byte-identical to the one that was tested and running yesterday. "Rollback" is therefore the same operation as "deploy", with a different version identifier: launch instances from the earlier image. This is why immutability is often argued for on reliability grounds rather than tidiness grounds — it converts a rare, improvised operation into the most-exercised path you have. (How traffic is actually moved between the old and new instances is a release-strategy question with its own answers; here the point is only that the artifact you need still exists and still runs.) ## Three preconditions people forget The property holds only if you maintain it: - **Retention.** The previous image must still be in the registry or image store. Aggressive garbage collection, a lifecycle rule, or a registry cleanup that keeps only the last tag will delete your rollback target. - **Immutable references.** If the version identifier is a moving tag rather than a fixed digest or version, "the previous image" may not be the bits you tested. Roll back to an identifier that can only mean one artifact. - **Still launchable.** An image that depends on a resource which has since been deleted or renamed — a security group, a subnet, an IAM role, a config key — will start and immediately fail. Rollback of compute alone is not enough if the surrounding declaration moved on. ## What redeploying the previous image does not restore This is the part interviewers are listening for. The image is compute and configuration. It carries nothing about the state the newer version produced. **Database schema and data.** If the new version applied a migration, the old code now faces a schema it was not written against — a dropped or renamed column is fatal. Worse, rows written by the new version may be in a format the old version cannot parse. The standard discipline is expand-and-contract: deploy the additive change first, make the code tolerate both shapes, and only remove the old shape a release later, once you no longer intend to roll back past it. **Messages and events already consumed.** Anything the new version read from a queue is gone; anything it published in a new format is already in other services' hands. Rolling back your service does not roll back its consumers. **External side effects.** Emails sent, payments captured, webhooks delivered, third-party records created. These are irreversible by construction, and a rollback that pretends otherwise produces duplicate effects when the old code retries. **Rotated credentials and secrets.** If the deploy rotated a key, the old image's expectations may no longer match what the secret store now holds. **Destroyed infrastructure.** If the change deleted a resource — a queue, a bucket, a table, a load balancer — redeploying the earlier compute image does not bring it back, and if the resource held data, no rollback will. **Caches and clients you do not control.** DNS records with a long TTL, CDN caches, browser bundles already downloaded, mobile app versions in users' hands that now expect the new API. Server-side rollback does not reach them. ## The rule that ties it together Compute rolls back; state rolls forward. Design changes so that the previous artifact can still operate correctly against the state the newer one left behind — which in practice means one-version backward compatibility as a standing rule, additive-first schema changes, and tolerating both data shapes during the overlap. And give the rollback a deadline. Backward compatibility is a cost you carry, so decide explicitly when a version stops being a rollback target — usually once the contract phase of a migration runs — and make that visible, so nobody discovers mid-incident that the artifact they were counting on can no longer run. ## What to say about testing it A rollback path that has never been executed is a hypothesis. Exercise it deliberately: deploy a version, roll back, and confirm the service is healthy against the current data — ideally in a lower environment as part of the release process, so the discovery that a migration is not reversible happens on a Tuesday afternoon rather than during an outage.

  • What migration discipline keeps the previous image deployable after a schema change?
    Expand and contract. First deploy the additive change — add the new column or table while leaving the old shape intact — then deploy code that writes both and reads either, and only in a later release remove the old shape. During the overlap either version can run against the database, so rollback stays a deploy rather than a data-recovery exercise. The contract step is the point at which you consciously give up the ability to roll back past it.
  • A team keeps only the most recent image in their registry to control storage costs. What have they given up?
    Their rollback target. The model's cheap rollback depends entirely on the previous artifact still existing and still being launchable, so a retention policy is a reliability control, not just a cost control. Keep enough versions to cover your realistic rollback window, pin deployments to immutable digests rather than moving tags, and make sure the retention rule cannot delete an image that is currently running or was running recently.
  • Which failures are made worse rather than better by rolling back?
    Anything where the new version already changed shared state irreversibly: data written in a new format the old code cannot read, messages consumed from a queue, payments or emails already dispatched, or a deleted resource. In those cases rolling back reintroduces old behaviour on top of new state, which can double-apply side effects. Fixing forward with a small targeted change is often the safer call, and knowing which situation you are in is the judgment being tested.

saying these in an interview costs you the question

  • Believes redeploying the old image also reverts database changes
  • Assumes the previous image is always still in the registry
  • Deploys from a moving tag and calls it a pinned rollback target
  • Thinks rollback undoes emails, payments or consumed messages
  • Never tests the rollback path before needing it in an incident

context