skip to content

On a single Docker host, how do you replace a running container with a new image version?

level: juniorimportance: must knowfreq 58%

answer

  1. The old one serves during the slow part
  2. Downtime is the swap, not the download
  3. Two image versions on disk at once
  4. A tag that never moves under you
  5. docker pull before docker stop

basics

~20 s

Pull the new image while the old container still serves, then stop it, remove it and run a new container on the new tag. Pin an immutable tag, and keep the previous image for rollback.

solid answer

~50 s

A container is one execution of one image; nothing mutates it in place. `docker restart` re-runs the **same** container from the same image ID, and `docker pull` only updates the local image store, so a deploy means creating a new container. Order the steps as pull-then-replace: `docker pull invoice-render:2026.08.19-7c41ab2` runs while the old container keeps answering, so the outage is only the stop plus start, not the download, and a failed or slow pull leaves production untouched. Then `docker stop`, `docker rm` (or run with `--rm`), and `docker run` with exactly the same flags. Keep those flags in one place — typically a systemd unit whose `ExecStartPre` does the pull — so the deploy is one `systemctl restart`. Pin a build-specific tag or a digest rather than a floating one, and do not delete the previous image: re-running it is your rollback.

code

bash · 13 lines
bash
set -euo pipefail
TAG=invoice-render:2026.08.19-7c41ab2

# 1. slow, fallible step happens while the old container still serves
docker pull "$TAG"

# 2. short swap window
docker stop -t 30 invoice-worker || true
docker rm invoice-worker || true
docker run -d --name invoice-worker \
  -p 127.0.0.1:8127:8000 \
  -v invoice-render-cache:/var/cache/render \
  "$TAG"

go deeper

for a junior

Be ready to state that images are immutable and that a new version means a new container: pull, stop, remove, run. Know that docker restart reuses the old image and that docker pull alone changes nothing about a running container.

for a middle

Explain why the pull goes first — it is the slow, failing step and the old container can serve through it — and why the run flags belong in a unit file rather than in someone's shell history.

for a senior

Show the operational judgement: pinned build tags so the running version is knowable, the previous image retained so rollback is a restart, a tolerated pull failure at boot, and awareness that the container's writable layer is discarded on every deploy.

for a principal

Own the release contract for the fleet: who decides the tag, how the tag maps to a commit, how far back rollback is guaranteed to work, and when the one-host replace stops being enough and the cost of a scheduler is worth paying.

### A container is not a server you upgrade in place An image is immutable, and a container is **one execution of one image with one set of flags**, fixed at the moment it was created. Nothing you can run against an existing container changes which image it runs. `docker restart invoice-worker` stops and starts the *same* container object — same image ID, same ports, same environment. `docker pull` fetches layers into the host's image store and never touches a running process. `docker update` changes a few resource limits and nothing about the image. So "deploy a new version" on a single host has exactly one shape: **create a new container from the new image and get rid of the old one.** Everything else is about ordering that sequence so the outage window is small and so you can go back. ### Pull, then replace The naive script is the wrong order: ```bash docker stop invoice-worker && docker rm invoice-worker docker pull invoice-render:2026.08.19-7c41ab2 # <- service is down for this docker run -d --name invoice-worker invoice-render:2026.08.19-7c41ab2 ``` The pull is the slow, network-dependent, failure-prone step: it can take tens of seconds for changed layers, and it can fail outright on a registry hiccup or an expired credential. Doing it *after* the stop means the service is down for all of it, and a failed pull leaves you with no container at all. Pull first and the old container is still answering requests the whole time; if the pull fails you abort and nothing has changed. The remaining downtime is the stop plus the start — for a small Python service on `python:slim` that is typically a second or two, dominated by the application's own startup. ### Pin the tag Pull-then-replace is only reproducible if the tag is. A floating tag such as `latest` makes two questions unanswerable: *which build is on this box?* and *what does rollback mean?* Two hosts pulling `latest` a minute apart can land on different builds, and re-pulling `latest` after a bad deploy gets you the bad build again. Use a build-specific tag — `invoice-render:2026.08.19-7c41ab2`, a date plus the commit — or pin the digest with `invoice-render@sha256:…`, which is the only form that is cryptographically exact. Note that a running container is never affected by someone re-pushing a tag; the container holds a resolved image ID. The change only arrives when you pull *and* recreate. ### Put the flags where the deploy can find them The run flags **are** the deployment: published ports, mounts, environment, `--user`, memory limits. If they live in a shell history, the next deploy loses one of them silently. On a single host the usual home for them is a systemd unit that owns the container, with the pull as a pre-start step: ```ini [Service] ExecStartPre=-/usr/bin/docker pull invoice-render:2026.08.19-7c41ab2 ExecStartPre=-/usr/bin/docker rm -f invoice-worker ExecStart=/usr/bin/docker run --rm --name invoice-worker \ -p 127.0.0.1:8127:8000 invoice-render:2026.08.19-7c41ab2 ExecStop=/usr/bin/docker stop invoice-worker ``` Deploying is then editing the tag and running `systemctl restart invoice-worker`. The leading `-` on the pre-start lines matters: without it, a boot while the registry is unreachable fails the unit and you get no service at all, when the cached image on disk would have been perfectly good. With it, a failed pull is tolerated and the host boots on what it already has. The `docker rm -f` line clears a stale container left by an unclean stop — otherwise `docker run --name invoice-worker` fails with `Conflict. The container name "/invoice-worker" is already in use`. ### Keep the old image Rollback on one host is `docker run` on the previous tag — seconds, and no registry round trip — **provided the previous image is still on disk**. A deploy script that prunes images immediately after a successful start has thrown that away and turned a ten-second rollback into a re-pull that may not even succeed. Keep the last few builds and clean up on a schedule instead. ### What replacement destroys Because the deploy deletes a container, everything in that container's writable layer goes with it: temporary files, anything the app wrote to a path that is not a mount. Any state that must survive a deploy has to live in a named volume or a bind mount, or outside the host entirely. This is worth saying out loud in an interview, because it is the difference between "we redeploy weekly" and "we redeploy weekly and lose the render cache each time". ### What this pattern does not give you Pull-then-replace shrinks the outage; it does not remove it. The stop and the start do not overlap, so there is a window — short, but real — in which nothing is listening on the port and callers get connection refused. Removing that window means running both containers at once behind something that routes, which is a different technique with its own constraints.

  • Why does `docker restart` not pick up an image you just pulled?
    `docker restart` operates on the existing container object, which was created from a specific image ID and holds its own flags. Restarting stops and starts that same execution; it never re-resolves the tag. The new image only takes effect when a new container is created from it, which is why deploys are remove-and-run rather than restart.
  • Your `ExecStartPre` pull fails at boot because the registry is unreachable. What should happen?
    The host should still start the service from the image already in its local store, because a cached build serving traffic beats no service at all. Making the pull non-fatal — in a systemd unit, prefixing the line with `-` — gets that. The tradeoff is that a genuinely failed deploy can then start silently on the old image, so the deploy path itself should check the pull's exit status rather than ignoring it.
  • How would you make a rollback on this host take seconds rather than minutes?
    Keep the previous image tagged and on disk, and keep the run flags in one file with the tag as the only thing that changes. Rollback is then editing the tag back and restarting the unit: no registry, no download, no reconstruction of forgotten flags. Pruning images as the last step of a deploy is what usually destroys this property.

saying these in an interview costs you the question

  • Says `docker restart` makes the container use the newly pulled image
  • Uses `:latest` and cannot say which build is running
  • Runs `docker pull` after stopping the container, extending downtime
  • Thinks `docker pull` updates a running container
  • Prunes the old image immediately, leaving no rollback
  • Recreates the container without the original run flags

context