skip to content

How does Docker's `overlay` network driver let containers on different hosts communicate as if on one network, and what does it require from the underlying infrastructure?

level: seniorimportance: should knowfreq 33%

answer

  1. VXLAN VNI per network, UDP 4789
  2. swarm control plane: 2377 mgmt, 7946 gossip
  3. ~50 byte header -> MTU ~1450
  4. small works / large hangs = MTU
  5. --opt encrypted = IPsec; --attachable for plain containers

basics

~20 s

Overlay networks tunnel container traffic between hosts using VXLAN: each network gets a VXLAN ID and packets are encapsulated in UDP (default port 4789) between host IPs. It needs a control plane (Docker swarm mode), the swarm management and gossip ports plus UDP 4789 open, and MTU headroom for the encapsulation header.

solid answer

~60 s

An overlay network gives containers on many hosts one flat address space. Each network is assigned a **VXLAN network identifier**; when a container sends to a peer on another host, the host wraps the layer-2 frame in a **UDP packet (default port 4789)** addressed to the peer's host, which decapsulates and delivers it. The physical network only ever sees host-to-host UDP, so no routes for container subnets are needed. The data plane needs a **control plane** to know which container IP lives on which host: in modern Docker that is **swarm mode**, whose managers distribute network state and whose nodes gossip endpoint updates. Practical requirements: - **Ports between hosts**: TCP 2377 (swarm management), TCP and UDP 7946 (node gossip), UDP 4789 (VXLAN data). - **MTU headroom** - encapsulation adds roughly 50 bytes, so overlay interfaces run at a reduced MTU; a mismatch shows up as small requests working and large ones hanging. - **Optional encryption** (`--opt encrypted`) adds IPsec between hosts at a CPU cost; the control plane is encrypted regardless. Standalone (non-service) containers can join only if the network is created `--attachable`.

code

bash · 5 lines
bash
docker swarm init --advertise-addr 10.10.0.11
docker network create -d overlay --attachable --opt encrypted app-net

docker network inspect app-net --format '{{.Driver}} {{.Options}}'
docker run --rm --network app-net alpine getent hosts api

go deeper

for a junior

Know that overlay is the multi-host driver and that it tunnels container traffic between hosts so containers appear on one network.

for a middle

Name VXLAN encapsulation over UDP 4789, the per-network VNI, and the need for a control plane plus the swarm ports.

for a senior

Diagnose from symptoms - blocked gossip versus blocked data port versus MTU - and decide on encryption and attachability deliberately.

for a principal

Weigh running Docker's own overlay against moving multi-host networking to an orchestrator or to routed underlay networking, on operability and troubleshooting cost.

## The problem Bridge networks are host-local: container IPs are meaningful only on the host that allocated them, and two hosts will happily hand out the same `172.18.0.2`. To let containers on different machines address each other directly you either teach the physical network about container subnets (routing, which needs infrastructure cooperation) or you **tunnel** - which is what the overlay driver does. ## VXLAN in one paragraph VXLAN (Virtual eXtensible LAN) encapsulates a complete layer-2 Ethernet frame inside a UDP datagram. Each virtual network gets a 24-bit **VNI**, carried in the VXLAN header, so many isolated virtual networks share one physical fabric. The endpoints doing the wrapping and unwrapping are **VTEPs** - here, the Docker hosts. On the wire your infrastructure sees only UDP between host addresses on port 4789; the container-to-container frame is the payload. Inside each participating host, Docker creates a dedicated network namespace per overlay network holding a bridge and a `vxlan` interface; each container's veth lands in that namespace's bridge. Traffic to a local peer stays on the bridge; traffic to a remote peer goes out through the VXLAN interface. ## The control plane Encapsulation alone is not enough - a host must know that container IP `10.0.1.7` currently lives on host B, and the corresponding MAC. Docker fills the forwarding and ARP tables from cluster state rather than by flooding. In current Docker that state comes from **swarm mode**: managers hold the network and endpoint definitions in the Raft-backed store, and nodes exchange endpoint updates over a gossip protocol. (Historically, standalone overlay networks used an external key-value store such as Consul or etcd; that path is legacy.) This is why the ports matter and why a partially blocked firewall produces confusing symptoms: block **2377** and nodes cannot join or receive network definitions; block **7946** and membership and endpoint propagation degrade, so some pairs of containers resolve and connect while others do not; block **4789** and the control plane looks perfectly healthy while no data flows at all. ## MTU Encapsulation costs roughly 50 bytes of header. If the underlay MTU is 1500, the overlay must use about 1450. Docker sets a reduced MTU on overlay interfaces, but the failure mode when something is inconsistent - a tunnel or VPN in the underlay with its own reduced MTU, or mixed manual configuration - is distinctive: the TCP handshake and small requests succeed, and anything that fills a full-size packet (a large POST, a TLS certificate chain, a big query result) hangs or times out. If ICMP "fragmentation needed" is filtered, path MTU discovery cannot repair it. Suspect MTU whenever "small works, large hangs". ## Encryption and isolation Overlay control traffic is encrypted by swarm. **Data** traffic is not, unless the network is created with `--opt encrypted`, which sets up IPsec between the hosts. That is a real CPU cost at high throughput, so it is a deliberate choice - typically yes when hosts communicate over untrusted or shared networks, often no inside a trusted private network where the underlay is already protected. Each overlay network is a separate VNI, so networks are isolated from one another in the same way separate bridges are. Overlay networks also get the usual embedded DNS behaviour, so containers address each other by name across hosts. ## Attachability and scope By default overlay networks created in swarm mode are usable only by swarm **services**. Creating one with `--attachable` lets plain `docker run` containers join, which is handy for debugging or mixed deployments - at the cost of a wider surface, since any container on any node can join. ## When to use it Overlay is the answer when you must keep multi-host container networking inside Docker's own tooling. If you are already moving to a full orchestrator, its cluster networking layer solves the same problem with its own plugins, and the concepts transfer directly: encapsulation, a control plane distributing endpoints, MTU headroom, and optional encryption.

  • Containers on an overlay network complete TCP handshakes and small requests, but large responses hang. What do you suspect?
    An MTU problem. VXLAN adds about 50 bytes, so overlay interfaces run near 1450; if some path in the underlay has a smaller MTU, or a manual MTU setting is inconsistent, full-size packets are dropped while small ones pass. If ICMP 'fragmentation needed' is filtered, path MTU discovery cannot recover, so you fix it by lowering the overlay MTU or repairing the underlay path.
  • Is overlay traffic encrypted by default?
    No. Swarm encrypts the control plane, but container data traffic is plain VXLAN unless the network is created with `--opt encrypted`, which establishes IPsec between the hosts. That adds CPU cost per packet, so it is a deliberate tradeoff - usually enabled when hosts talk over untrusted or shared infrastructure.

VXLAN is putting a sealed local envelope inside a national postal envelope: the postal network only routes between the two buildings, and the inner envelope is delivered to the right desk on arrival.

saying these in an interview costs you the question

  • Believing overlay works between plain Docker hosts with no cluster or control plane
  • Assuming container data traffic is encrypted by default
  • Ignoring MTU and treating 'small requests work, large ones hang' as an application bug
  • Opening only the VXLAN data port and expecting membership and discovery to work
  • Thinking any standalone container can join a swarm overlay without --attachable

context