What does Docker Swarm mode add on top of a single Docker Engine, and how is a swarm service different from a container you start with `docker run`?
answer
- swarm init / join, managers + workers
- service = desired state, task = one container
- replicated vs global mode
- reconciliation replaces, never restarts in place
- docker service ps for task history
basics
~20 sSwarm mode turns several engines into one cluster of manager and worker nodes. A service is declared desired state — image, replica count, update policy — and managers keep that many tasks (containers) running, rescheduling them when containers or nodes fail.
solid answer
~50 s`docker run` creates one container on one engine; if it dies or the host reboots, all you get is whatever local restart policy you set. Swarm mode (`docker swarm init`, then `docker swarm join`) makes a set of engines a cluster of **managers**, which hold cluster state in a Raft log and schedule work, and **workers**, which run it. You then declare a **service**: an image, a replica count, ports, networks, resource limits, placement constraints and an update policy. The managers create that many **tasks**, each task being one container on some node, and continuously reconcile reality with the declaration — a dead container is replaced, a lost node's tasks are rescheduled elsewhere. On top of that you get service discovery by name, overlay networks spanning hosts, a routing mesh for published ports, mutual TLS between nodes, secrets and configs, and rolling updates with rollback — all in the engine, with no extra components.
code
bash · 9 linesdocker swarm init --advertise-addr 10.0.0.11
docker swarm join-token -q worker
# on each worker:
docker swarm join --token SWMTKN-1-xxx 10.0.0.11:2377
docker network create -d overlay appnet
docker service create --name api --replicas 3 \
--network appnet --publish 8080:8080 \
--constraint 'node.role==worker' registry.example.com/api:1.2go deeper
Know the vocabulary — manager, worker, service, task — and that a service is a declared replica count the cluster maintains.
Explain the reconciliation loop, replicated versus global mode, constraints, and what changes versus single-host habits (registry required, per-node bind mounts).
Bring in operational detail: task state history for diagnosis, healthchecks driving replacement and rolling updates, and how discovery and published ports actually work.
Frame it as the cheapest orchestration step up from single hosts, and be explicit about what the model does not give you before committing a platform to it.
## The gap swarm mode fills A single engine is imperative: you tell it to start a container and it does. Nothing watches the outcome across machines, nothing reschedules, nothing gives you one name that reaches several replicas. Swarm mode adds a control loop and a cluster. `docker swarm init` on the first node makes it a manager and prints join tokens — one for workers, one for managers. Other engines run `docker swarm join --token ... host:2377`. From then on the manager's CLI is a cluster CLI: `docker node ls`, `docker service ls`. ## Service, task, container The object model has three layers: - **Service** — the declaration: run `myapp:1.2`, three replicas, published on 8080, on network `appnet`, restart on failure, update two at a time. - **Task** — the scheduling unit. The manager creates one task per desired replica, assigns each to a node, and a task moves through a lifecycle (new, pending, assigned, preparing, running, complete or failed). Tasks are immutable: a task is never moved or restarted in place; a failed one is replaced by a new task. - **Container** — what the node's engine actually runs for a task. `docker service ps <svc>` shows the tasks and their history, which is where you look when replicas keep dying — you see the failure and the successive replacements. ## Replicated and global services `--replicas N` gives a fixed count spread by the scheduler. `--mode global` runs exactly one task on every node that matches the constraints, which is how agents such as log shippers or metrics collectors are deployed. Placement is influenced with `--constraint 'node.labels.zone==a'` and spread with `--placement-pref 'spread=node.labels.zone'`. ## What comes for free - **Discovery.** Every service on an overlay network is resolvable by its name from other containers on that network; the built-in DNS returns a virtual IP that load balances across healthy tasks. - **Published ports.** `--publish 8080:80` makes the port reachable on *every* node through the routing mesh, not only on nodes running a task. - **Rolling updates.** `docker service update --image myapp:1.3` replaces tasks in batches with configurable parallelism, delay and failure action, and `docker service rollback` returns to the previous spec. - **Security plumbing.** Nodes authenticate to each other with mutual TLS using automatically rotated certificates; `docker secret` and `docker config` deliver files to tasks in memory rather than baking them into images. - **Health-aware scheduling.** If the image declares a `HEALTHCHECK`, a task that goes unhealthy is replaced, and rolling updates wait for health before proceeding. ## Where the analogy to `docker run` breaks Several single-engine habits do not carry over. You cannot name the container (`--name` on a service names the *service*; task containers get generated names). You do not `docker start` a task. Logs come from `docker service logs`, which aggregates across tasks. Bind mounts still refer to *each node's* filesystem, so a service with three replicas on three nodes sees three different directories — shared state needs a volume plugin or an external store. And images must be pullable by every node, which in practice means a registry, not a locally built image.
- What is the difference between a replicated service and a global service?A replicated service runs a fixed number of tasks that the scheduler places wherever constraints allow, and scaling changes that number. A global service runs exactly one task on every node satisfying its constraints, and it grows automatically when a node joins the swarm. Global mode is the natural fit for per-node agents such as log or metrics collectors.
- A task keeps failing and being recreated. Where do you look first?`docker service ps --no-trunc <service>` lists the task history with desired and current state plus the error for each failed attempt, which usually names an image pull failure, a constraint no node satisfies, or a non-zero exit. From there `docker service logs <service>` aggregates the containers' output, and `docker node ls` confirms the target nodes are actually available.
docker run is hiring one person for one shift; a service is a staffing contract — you specify three people on duty at all times and the agency keeps refilling the slots.
saying these in an interview costs you the question
- Describing a service as just a container with a restart policy
- Thinking a failed task is restarted in place rather than replaced by a new task
- Expecting a locally built image to be available on other nodes without a registry
- Believing bind mounts give replicas on different hosts the same data
- Confusing swarm mode with the obsolete standalone Docker Swarm product