skip to content

A production release is bad and you decide to roll back. Your delivery system can either redeploy the previously built artifact or revert the commit and let CI build and deploy a fresh one. Which one is the rollback, and why does the other still have to happen?

level: juniorimportance: should knowfreq 55%

answer

  1. deployment action, not a source action
  2. the bytes that already ran
  3. rebuild produces a third, untested version
  4. source is the desired state
  5. reconciler puts the bad version back

basics

~20 s

Redeploy the previously built artifact - it already exists, already ran in production, and returns in minutes instead of a full build. Revert the source commit as well, or the next deploy reships the bad code.

solid answer

~50 s

The rollback is redeploying the artifact that was already running before this release: those exact bytes were built once, passed the pipeline, and served real traffic, so redeploying them is the fastest change with the least new risk. Reverting the commit and rebuilding is not a rollback - it produces a *third* version nobody has run. Between then and now the branch may have taken other merges, and unpinned dependencies or base images can resolve differently, so you would be shipping an untested build in the middle of an incident. But the revert still has to land in source control, because source is the desired state: without it the next merge redeploys the bad code, and if a controller continuously reconciles the cluster against the repository it will simply put the bad version back. So: artifact first to stop the bleeding, revert second to make the fix stick.

go deeper

for a junior

Know that a rollback means redeploying the artifact that already ran, not rebuilding old source, and say that the bad commit still has to be reverted afterwards.

for a middle

Explain why a rebuild is a different version - other merges on the branch, unpinned dependencies, a moving base image - and why a reconciling deployment platform will undo a manual redeploy that source control contradicts.

for a senior

Show that you treat time-to-rollback as a measured number and that you have made the path immutable, one-action and exercised, rather than a runbook nobody has run under pressure.

for a principal

Be ready to argue where the organisation should spend to keep this path real - reference immutability, retention of rollback-eligible releases, and the honest cases where no previous artifact can run against today's state.

## What "roll back" actually means A rollback is a *deployment* action, not a source-control action. The thing you want back is the exact software that was serving traffic before the bad release - a specific container image digest, a specific package version, a specific bundle. That thing already exists in a registry or artifact store. Pointing production at it again is a single operation whose duration is measured in the time to pull an image and cycle the instances. Reverting the commit and letting CI build produces something different: a new artifact from a new source tree. It may or may not be equivalent to the old one, and you will not know until it is running in production. ## Why the rebuild is not the old version Three things drift between the moment a release was built and the moment you would rebuild its predecessor: - **The branch moved.** Reverting your bad commit on top of `main` also carries every other commit merged since. That is not the previous release; it is the previous release plus whatever else landed today. - **Inputs are rarely pinned end to end.** Version ranges in a lockfile-free dependency, a floating base image tag, a toolchain installed at build time - any of these can resolve differently now than they did last week. Reproducible builds are an achievement, not a default. - **Time.** A pipeline that runs full verification takes tens of minutes. If your service is down, that entire window is added to the incident, and the pipeline itself may be queued behind other work. During an incident you want the change with the *smallest* unknown. The known-good artifact has an unknown of nearly zero; a fresh build has an unknown you cannot bound. ## Why the revert still matters Stopping at "we redeployed the old image" leaves production and source disagreeing, and something will eventually reconcile that disagreement against you: - The next person to merge anything triggers a deploy from `main`, which still contains the bad commit - the outage returns, now with a confusing cause. - If your platform continuously reconciles running state against a repository, it treats your manual redeploy as *drift* and reverts it back to the bad version, sometimes within seconds. The rollback appears to "not stick". - The audit trail matters later: the postmortem needs the record of what shipped and what was withdrawn. So the sequence is: redeploy the known-good artifact to restore service, then revert (or fix and re-land) in source so the desired state matches reality. ## Making the artifact path real For "redeploy the previous artifact" to be a real option and not a hope, a few properties have to be engineered in advance: - **Immutable, content-addressed references.** "Redeploy the tag `v2.3`" is unsafe if tags can be repointed - deploy by digest or by an identifier that can never mean different bytes. - **The previous artifact must still exist.** Whatever release you might roll back to has to still be in the store; a release that has been garbage-collected is not a rollback target. - **One action, not a runbook of twelve steps.** Rollback time is part of your recovery time. If it takes fifteen minutes of manual work, it is a fifteen-minute outage floor. - **Exercise it.** A rollback path that has never been run is a claim, not a capability. ## When rebuilding is the only option Sometimes the previous artifact genuinely cannot be run again - a completed data migration means the old code no longer understands the database, or the old build carried a credential or certificate that has since been rotated. That is the moment you discover that rollback readiness is an engineering property of the release, not a button in a tool: if you never kept the previous version able to run against today's state, your only path is forward, and the length of your pipeline becomes the floor on your recovery time.

  • Your rollback instruction says "redeploy the image tagged v2.3". What is wrong with that instruction?
    Tags are mutable in most registries: `v2.3` can have been repointed to different bytes since it first shipped, so the instruction does not identify a specific release. Roll back by immutable digest, or by an identifier your build system guarantees is never reused. The point of a rollback is to run something you have already run; a reference that can silently mean something else defeats it.
  • Why is time-to-rollback worth measuring as its own number?
    Because it is a floor on recovery time for any incident a rollback would fix. If deploying takes four minutes and rolling back takes twenty because it involves a manual approval and a runbook, then every bad release costs at least twenty minutes regardless of how fast you detect it. Measuring it turns "we can roll back" into a number you can improve and drill.

saying these in an interview costs you the question

  • Reverting the commit and rebuilding is the same as rolling back
  • A rebuild of the old commit produces the identical artifact
  • Once the old image is redeployed, source control needs no change
  • We can always rebuild any old version if we need it
  • Rollback speed does not matter because deploys are fast

context