How many manager nodes should a Docker Swarm cluster run, what happens when a majority of managers becomes unreachable, and how do you recover from that?
answer
- Raft majority: 3 tolerates 1, 5 tolerates 2
- even manager counts add nothing
- no quorum = no scheduling, tasks keep running
- --force-new-cluster on a survivor
- drain managers, back up /var/lib/docker/swarm
basics
~20 sManagers replicate cluster state with Raft, so use an odd number — 3 or 5 — tolerating (N-1)/2 failures. Losing quorum freezes all cluster changes while existing tasks keep running. Recover by restoring managers, or run docker swarm init --force-new-cluster on a survivor.
solid answer
~50 sManagers hold cluster state in a Raft log, so writes need a majority. Three managers tolerate one failure, five tolerate two; even counts add no fault tolerance, and beyond seven you mostly add consensus latency. Workers scale independently — managers are about the control plane, not capacity. **Losing quorum** does not stop your applications: workers keep running the containers they already have. What stops is *change* — no scheduling, no service create/update/scale, no node joins, no rescheduling of failed tasks. The CLI reports that the swarm has no leader. **Recovery**, in order: bring the missing managers back, since state is intact and quorum simply returns. If they are permanently gone, run `docker swarm init --force-new-cluster` on a surviving manager — it keeps existing state and shrinks the Raft set to that one node — then promote replacements. Long term: spread managers across failure domains, run them with `--availability drain`, and back up `/var/lib/docker/swarm` from a stopped manager.
code
bash · 9 linesdocker node ls
# ID HOSTNAME STATUS AVAILABILITY MANAGER STATUS
# ... mgr1 Ready Drain Leader
# ... mgr2 Ready Drain Reachable
# ... mgr3 Down Drain Unreachable
docker node update --availability drain mgr1
docker node promote worker4
docker node demote mgr3 && docker node rm mgr3go deeper
Know that managers replicate state, that you want an odd number (usually 3), and that workers do the actual work.
Explain the (N-1)/2 arithmetic and the practical split between what keeps running and what freezes when quorum is lost.
Walk the recovery path end to end — restore managers, force-new-cluster, backup restore — plus the hygiene around draining, backups and autolock.
Reason about failure domains and blast radius: manager placement across zones, partition behaviour, whether the control plane's availability target justifies five managers, and how quorum loss is detected and paged.
## Why the numbers are what they are Swarm managers keep cluster state — services, networks, secrets, node membership — in a replicated Raft log. Raft commits a write only when a majority of members acknowledge it, which is what makes the state consistent across managers. With N managers the tolerated failures are floor((N-1)/2): - 1 manager: 0 failures (fine for development, no high availability) - 3 managers: 1 failure - 5 managers: 2 failures - 7 managers: 3 failures Even counts are wasteful: 4 managers tolerate one failure just like 3, while adding another machine that can fail. Beyond 7, every write waits on more acknowledgements and the control plane slows for no practical gain. Worker count is unrelated — a swarm can be 3 managers and 200 workers. Placement matters as much as count: three managers in one rack or one availability zone tolerate one *node* failure but zero *domain* failures. Spread them across domains, and remember that a network partition leaving no side with a majority freezes the whole control plane. ## What losing quorum actually does This is the point candidates most often get wrong in both directions. Losing quorum is not an outage of your applications, and it is not harmless. Still working: worker nodes keep running their assigned tasks, published ports keep serving, overlay networking on established endpoints keeps forwarding. Stopped: anything that writes cluster state. No `docker service create/update/scale`, no new nodes joining, no promotions, and — the dangerous one — **no rescheduling**. If a worker dies while quorum is lost, its tasks are simply gone until the control plane returns. So the cluster degrades quietly with every subsequent failure. ## Recovering 1. **Restore managers.** Usually the right move: the Raft state on the surviving disks is intact, so bringing back enough managers restores quorum with no data loss. 2. **Force a new cluster.** If the lost managers are unrecoverable, `docker swarm init --force-new-cluster` on a surviving manager reinitialises the Raft set as a single node while preserving services, networks and secrets. Nodes still reachable stay members. Then `docker node promote` new managers back to 3 or 5, and `docker node rm` the dead entries. 3. **Restore from backup.** With every manager lost, stop the daemon on a fresh node, restore `/var/lib/docker/swarm` from a backup taken while a manager was stopped, start the daemon and force a new cluster. Node membership will need repair, and services reconcile as nodes rejoin. ## Operating managers well - **Drain them.** `docker node update --availability drain <mgr>` keeps workload tasks off managers so a busy service cannot starve the control plane. Standard for anything beyond a toy cluster. - **Back up the Raft directory.** `/var/lib/docker/swarm` contains the whole cluster state including secrets; back it up with the daemon stopped and treat the backup as sensitive, because it is. - **Consider autolock.** `docker swarm update --autolock=true` encrypts the Raft log at rest and requires an unlock key on daemon restart. It protects the secrets in that directory at the cost of an operator step, or a stored key, on every restart. - **Watch the leader.** `docker node ls` shows manager status — Leader, Reachable, Unreachable. Alerting on anything not Reachable is the cheapest early warning available. - **Prefer replacement to repair.** A manager whose Raft state is corrupt is demoted, removed and re-added rather than debugged.
- Quorum is lost and a worker node crashes. What happens to the tasks that were running on it?They stay down. Rescheduling is a cluster-state write, so with no leader the managers cannot create replacement tasks anywhere. Surviving replicas continue serving, but the service runs below its declared count until quorum returns, which is why a quorum outage should be treated as urgent even though nothing appears broken at first.
- Why not simply run every node as a manager?Every manager participates in Raft, so each additional one increases the acknowledgements each write needs and slows the control plane while contributing nothing to fault tolerance beyond the majority arithmetic. Managers also hold the full cluster state including secrets, so widening that set widens the blast radius of a compromised node. Three or five managers plus plain workers is the standard shape.
saying these in an interview costs you the question
- Recommending an even number of managers such as 2 or 4
- Saying applications stop when quorum is lost
- Believing tasks are still rescheduled without a leader
- Running heavy workloads on managers without draining them
- Forgetting that /var/lib/docker/swarm holds secrets and must be backed up securely