skip to content

Walk through what `docker service update --image myapp:2.0` does to a 10-replica swarm service, and which settings control the blast radius if the new image is broken.

level: middleimportance: should knowfreq 30%

answer

  1. tasks are replaced, never mutated
  2. parallelism + delay = batch size and soak
  3. stop-first dips capacity; start-first runs both versions
  4. monitor window + healthcheck decide failure
  5. default failure_action=pause leaves it half updated

basics

~20 s

The manager updates the service spec and replaces tasks in batches. update_config sets parallelism, delay, order (stop-first or start-first), monitor window, max failure ratio and failure_action (pause, continue, rollback). docker service rollback restores the previous spec.

solid answer

~40 s

The manager records a new service spec, then reconciles: tasks are replaced in batches of `--update-parallelism` (default 1), waiting `--update-delay` between batches. `--update-order` decides whether each replacement stops the old task first (`stop-first`, the default) or starts the new one first (`start-first`, needed for zero downtime but temporarily running both versions). After a task starts, swarm watches it for `--update-monitor` (default 5s). If it exits or its `HEALTHCHECK` fails within that window the step counts as failed; once failures exceed `--update-max-failure-ratio`, `--update-failure-action` decides what happens — `pause` (the default, leaving the service half updated), `continue`, or `rollback`. The safe production shape is parallelism 1, a real healthcheck, a monitor window longer than start-up, and `failure_action=rollback` with a matching `rollback_config`. `docker service ps` shows per-task history; `docker service rollback <svc>` reverts manually.

code

bash · 8 lines
bash
docker service create --name api --replicas 10 \
  --update-parallelism 1 --update-delay 20s \
  --update-order start-first --update-monitor 30s \
  --update-failure-action rollback \
  --rollback-parallelism 2 --rollback-order stop-first \
  --health-cmd 'curl -fsS localhost:8080/healthz || exit 1' \
  --health-interval 5s --health-retries 3 \
  registry.example.com/api:1.0

go deeper

for a junior

Know that the update replaces tasks gradually, that parallelism and delay control how quickly, and that rollback exists.

for a middle

Explain every knob — parallelism, delay, order, monitor, max_failure_ratio, failure_action — and why the pause default leaves a mixed-version service.

for a senior

Tie it to health signals and compatibility: a healthcheck is what makes the monitor window meaningful, and start-first demands version compatibility and spare capacity.

for a principal

Set the organisation's default policy — automatic rollback, digest pinning, mandatory healthchecks — and be explicit about which classes of change, such as schema migrations, a rolling update cannot handle at all.

## What actually happens `docker service update` mutates the service specification held in the Raft store and bumps its version. The orchestrator then reconciles the running tasks with the new spec. Because tasks are immutable, updating means shutting down old tasks and creating new ones — replicas are replaced, never mutated. The replacement is deliberately incremental: - **parallelism** — how many tasks are converted at a time. Default 1. With 10 replicas and parallelism 1 you get ten sequential steps; parallelism 0 means all at once. - **delay** — how long to wait after a batch converges before starting the next one. This is your soak time per batch. - **order** — `stop-first` (default) removes the old task then starts the new one, so capacity dips by one task per batch. `start-first` starts the replacement first and removes the old task once the new one is up, so capacity never dips but both versions run simultaneously — which requires the versions to be compatible, especially against a shared database. - **monitor** — the observation window after a task starts (default 5 seconds). A task that exits or reports unhealthy inside the window marks the step failed. This is why a healthcheck matters: without one, started means the process launched, so a container that boots and then fails to serve traffic sails through. - **max-failure-ratio** — the fraction of failed tasks tolerated before the update is considered failed (default 0). - **failure-action** — `pause` (default), `continue`, or `rollback`. ## The default is a trap With `failure_action: pause`, a broken image leaves the service *half updated*: some tasks on 2.0 failing, the rest on 1.0 serving. The update stops but nothing is undone, and `docker service ps` shows the mixture. That is safe in the sense that it does not destroy the remaining good replicas, but it is not self-healing and it needs a human. Setting `failure_action: rollback` with a `rollback_config` (its own parallelism, delay and order) makes the service return to the previous spec automatically. ## Rolling back `docker service rollback <service>` reverts to the previously stored spec — swarm keeps the prior version specifically for this. It respects `rollback_config` if present, otherwise the update settings. Note that it toggles between the last two specs, so calling it twice returns you to where you started; it is not a version history. ## Related update mechanics - **Tag resolution.** By default swarm resolves the image tag to a digest at update time and pins tasks to that digest, so every replica runs the identical image even if the tag is later moved. `--no-resolve-image` disables it; deploying by digest explicitly is better still. - **Registry credentials.** `--with-registry-auth` forwards your registry token so nodes can pull private images; without it an update of a private image fails at pull time on the nodes. - **Forcing.** `docker service update --force` recreates tasks with an unchanged spec — useful to pick up a moved tag, redistribute after adding nodes, or restart a stuck service. - **Observing.** `docker service ps <svc>` shows desired versus current state per task and the error text of failed attempts; `docker service inspect --format '{{.UpdateStatus.State}} {{.UpdateStatus.Message}}' <svc>` gives the update's own state (updating, paused, rollback_started, completed). ## A production-shaped setting Parallelism 1, delay comfortably above start-up plus warm-up, order `start-first` when the versions are compatible, a monitor window longer than the healthcheck needs to become meaningful, `failure_action: rollback`, and an actual `HEALTHCHECK` in the image. Without the healthcheck, most of these knobs are measuring the wrong thing.

  • Why does a rolling update often appear to succeed with a broken image, and what fixes that?
    Without a HEALTHCHECK, swarm only knows whether the container process started, so an image that boots and then fails every request passes the monitor window and the update rolls forward across all replicas. Adding a healthcheck that exercises a real readiness endpoint, and setting the monitor window longer than start-up, makes the failure visible in time for failure_action to trigger.
  • When is update order start-first the wrong choice?
    When the two versions cannot safely run at the same time — an incompatible database migration, an exclusive lock or licence, or a singleton that must not be duplicated. start-first also needs spare capacity, since the replica count temporarily exceeds the declared number. In those cases stop-first with parallelism 1 correctly trades brief capacity loss for version exclusivity.

saying these in an interview costs you the question

  • Believing a failed update rolls back automatically by default
  • Updating with parallelism 0 or no delay on a large service
  • Relying on the monitor window with no HEALTHCHECK in the image
  • Assuming docker service rollback is a full version history
  • Forgetting --with-registry-auth and blaming the image when nodes cannot pull

context