skip to content

Rolling back by redeploying the image tag `indexer:stable` brings the same broken build back. Why?

level: seniorimportance: should knowfreq 58%

answer

  1. The release moved the thing you rolled back to
  2. A mutable name in the rollback path
  3. Local images are reused without asking the registry
  4. Compare image IDs, not tag strings
  5. The registry keeps no history of a pointer

basics

~20 s

Because the release already re-pointed that alias at the broken build, so redeploying it fetches the same manifest. A rollback has to name a tag nobody moved — the previous build's own immutable tag — not the alias the release just updated.

solid answer

~50 s

`stable` is a mutable name, and the deploy that broke production re-pointed it. Redeploying the same name therefore resolves to the same manifest, and if the host already has an image under that tag, `docker run` will not even re-check the registry unless you `docker pull` first or pass `--pull=always`. Confirm it in seconds: `docker image inspect --format '{{.Id}}' registry.internal:5000/indexer:stable` on the host, then pull the alias again and compare — an unchanged image ID tells you the alias is the problem, not the deploy tooling. The fix during the incident is to deploy the previous build by its own immutable tag, for example `registry.internal:5000/indexer:7e1c94d`, and only afterwards re-point `stable` back. The fix afterwards is structural: deployments and rollbacks reference build tags, and the alias moves as a consequence of a successful deploy, never as the deploy itself.

code

bash · 7 lines
bash
# is the running container really the build we think it is?
docker inspect --format '{{.Image}}' indexer

# does the alias still resolve to the same bytes after a fresh pull?
docker image inspect --format '{{.Id}}' registry.internal:5000/indexer:stable
docker pull registry.internal:5000/indexer:stable
docker image inspect --format '{{.Id}}' registry.internal:5000/indexer:stable

go deeper

for a junior

Understand the core fact: an alias points at whatever was pushed to it last, so redeploying it after a bad release gives you the bad release again. Rollback needs the previous build's own tag.

for a middle

Explain the mechanics you would check — the local image cache being used without contacting the registry, and comparing image IDs before and after a fresh pull to prove where the name resolves.

for a senior

Show the incident method: confirm what is running, prove the alias moved, roll back by immutable tag, re-point afterwards, and identify that the missing deploy record is what made the incident slow.

for a principal

Own the structural rule that keeps mutable names out of the rollback path, and make sure some system other than the registry records which build each alias move pointed at.

### What actually happened A release of a document-indexing service shipped build `7e1c94d`, and the pipeline did what it always does: pushed the build tag, then re-pointed the alias with `docker tag ... indexer:stable` and a second push. The build turned out to hang on shutdown — the Elixir release ignored SIGTERM, so every stop hit the 4-second grace period and was killed, losing partial index state. Twenty-three minutes into the incident somebody runs the rollback: redeploy `indexer:stable`. The same broken container comes back. This is not a tooling bug. It is the defining property of an alias. `stable` is not "the last good build" — it is "whatever was pushed to that name most recently", and the release you are trying to undo is what pushed it. **You cannot roll back to a name that the thing you are rolling back moved.** ### Why it can look even stranger than that Two effects often overlap and make the diagnosis muddier. First, **local caching**. `docker run registry.internal:5000/indexer:stable` uses a local image with that tag if one exists; it does not contact the registry to check whether the tag has moved. So on a host that already pulled the image, a rollback can appear to "work" while doing nothing at all. Force the question with an explicit `docker pull`, or `docker run --pull=always`. Second, **skew across hosts**. Hosts that pulled the alias before the release still hold the old bytes; hosts that pulled after hold the new ones. The fleet is then running two different builds under one name, and "is the rollback done?" has no single answer. This is the same root cause wearing a different costume. ### Diagnosing it in under a minute Work from the identity of the bytes, not from the name: ``` # what is this host actually running? docker inspect --format '{{.Image}}' indexer # what does the local tag point at, and does the registry agree? docker image inspect --format '{{.Id}}' registry.internal:5000/indexer:stable docker pull registry.internal:5000/indexer:stable docker image inspect --format '{{.Id}}' registry.internal:5000/indexer:stable ``` If the pull changes nothing and the container is still the bad build, the alias is pointing at the bad build. If the pull *does* change it, you had a stale cache as well. Either way, the fix path is the same. ### The fix during the incident Deploy the previous build by a name that nobody moved: ``` docker pull registry.internal:5000/indexer:2b83f0e docker run -d --name indexer registry.internal:5000/indexer:2b83f0e ``` That requires knowing which build was previous, which is the part teams most often lack. The registry does not keep the history of an alias — one pointer, no log. The answer has to come from somewhere that does keep history: the deployment record that says which immutable tag was rolled out and when. If your deploy tool records only "deployed indexer:stable", it has recorded nothing useful, and the incident now includes an archaeology step. Move the alias back **after** the rollback is verified healthy, not before. Otherwise anybody pulling `stable` mid-incident — a developer, a smoke test, a host that happens to restart — gets the build you are trying to remove. ### The fix afterwards Three changes turn this class of incident off: 1. **Deployments name immutable build tags.** The alias may exist, but nothing in the deploy path resolves it. This is the single change that does most of the work. 2. **Deploy records store the resolved build tag**, so "the previous release" is a lookup rather than a reconstruction. Ideally the rollback command reads it directly. 3. **Alias moves happen after a healthy deploy**, as a separate step, and are equally logged. A useful sanity question for any release process: *if the alias were deleted right now, could we still redeploy every release from the last month?* If the answer is no, the aliases are load-bearing and rollback is a rebuild in disguise. ### The distinction worth stating out loud Mutable names are fine — they are convenient and people want them. What is not fine is a mutable name in the rollback path. Rollback is the operation you perform when you are already wrong about something, under time pressure, so it must depend on the fewest possible things that can have changed since. A name written once by a build that has already run is about as few as it gets.

  • How do you find out which build was running before, when the registry keeps no history for the alias?
    The registry cannot help — a tag is a single pointer with no log. The answer has to come from a system that records history: the deploy log, the pipeline run that produced the previous release, or an internal release ledger mapping each alias move to a build tag and a timestamp. If none of those record the resolved build tag, that is the first thing to fix after the incident, because it converts a lookup into archaeology.
  • Why can a rollback appear to succeed on some hosts and not on others?
    Because each host resolves the alias independently and caches the result. Hosts that pulled before the bad release still hold the old bytes and look fine; hosts that pulled after hold the new ones. Under one tag the fleet is split. Comparing the running container's image ID per host exposes it, and deploying by an immutable build tag with an explicit pull removes the ambiguity entirely.
  • Is it ever right to roll back by re-pointing the alias backwards?
    Only as a follow-up, never as the rollback itself. Re-pointing tells future pulls what is good, but it does not change anything already running, it leaves a window where the fleet is split, and it depends on the same mutable name that just failed you. Deploy the previous build tag first, verify health, then move the alias so it once again describes reality.

saying these in an interview costs you the question

  • Assumes redeploying an alias fetches an older image
  • Trusts the tag name over the running image ID
  • Forgets that docker run reuses a cached local image
  • Expects the registry to hold a history of tag moves
  • Re-points the alias before verifying the rollback
  • Proposes rebuilding the old commit as the rollback

context