In a Docker Swarm cluster, how does a client request reach a service replica when it hits a published port on a node that is running no replica of that service, and how do containers on different hosts reach each other?
answer
- 2377 mgmt, 7946 gossip, 4789 VXLAN data
- VIP + IPVS vs dnsrr endpoint mode
- tasks.<service> gives raw task IPs
- ingress mesh: every node accepts, SNAT hides client IP
- mode=host preserves source IP, no mesh
basics
~20 sPublished ports use the ingress overlay network and routing mesh: every node listens on the port and load balances to a task anywhere via IPVS. Container-to-container traffic runs over VXLAN-encapsulated overlay networks, with service names resolving to a virtual IP.
solid answer
~50 sTwo mechanisms, both built on overlay networking. **East-west.** An overlay network is a VXLAN layer-2 segment spanning nodes; packets are encapsulated and sent over UDP 4789 between hosts, while membership propagates over the gossip control plane on 7946 TCP/UDP. A service on that network gets a **virtual IP**; the embedded DNS resolves the service name to the VIP, and the kernel's IPVS load balances the VIP across healthy task IPs. `--endpoint-mode dnsrr` instead returns task addresses directly, for clients that want their own load balancing. **North-south.** `--publish 8080:80` attaches the service to the `ingress` overlay network and opens 8080 on *every* node. Hit any node and IPVS forwards to a task on whichever node has one — that is the routing mesh. The cost is source NAT, so the app sees a cluster-internal source address. `--publish mode=host` binds only nodes running a task and preserves the client IP, with an external load balancer in front.
code
bash · 8 linesdocker network create -d overlay --opt encrypted --attachable appnet
docker service create --name api --network appnet -p 8080:8080 --replicas 3 api:1.0
docker service inspect --format '{{json .Endpoint.VirtualIPs}}' api
# bypass the routing mesh, keep the client source IP
docker service create --name edge --network appnet \
--publish mode=host,target=443,published=443 --mode global edge:1.0go deeper
Know that overlay networks let containers on different hosts talk by service name, and that a published port answers on every node.
Explain VXLAN encapsulation, the three port numbers, and the difference between VIP and dnsrr endpoint modes.
Diagnose it: VIP versus task DNS, IPVS behind the VIP, source-IP loss through the mesh, MTU and blocked-4789 symptoms, host-mode publishing as the escape hatch.
Decide the north-south design — mesh plus a simple external load balancer versus host mode with a placement-aware one — and whether encrypted overlay's throughput cost is warranted on the network you actually run.
## Overlay networks An overlay network gives containers on different hosts one flat layer-2 segment. Each host has a hidden network namespace holding a bridge and a VXLAN interface; a packet from container A to container B on another node is bridged, encapsulated in a VXLAN header, sent host-to-host over **UDP 4789**, decapsulated on the far side and delivered. Membership and endpoint information propagate over the **gossip control plane on 7946 (TCP and UDP)**, while cluster management uses **2377/TCP** to the managers. Any firewall exercise on a swarm comes down to those three. By default the data path is unencrypted; `docker network create -d overlay --opt encrypted appnet` adds IPsec between nodes at a throughput cost. The control plane is already mutually authenticated with TLS. ## Discovery and internal load balancing Every service attached to an overlay network gets an entry in the swarm's embedded DNS. In the default **VIP endpoint mode**, the name resolves to a single virtual IP that belongs to no container. Traffic to the VIP is intercepted by **IPVS** in the kernel on the sending node and distributed across the task IPs, so failover and scaling need no client awareness and no DNS cache expiry. In **dnsrr endpoint mode** (`--endpoint-mode dnsrr`) the DNS answer contains the task addresses instead and the client picks. It suits clients that do their own pooling or need per-connection control, and it is required with host-mode publishing. The tradeoff is that clients cache DNS answers and can keep hitting departed tasks. `tasks.<service>` always resolves to the individual task IPs regardless of endpoint mode, which is handy for peer discovery in clustered applications. ## The routing mesh When a service publishes a port in the default `ingress` mode, swarm attaches it to a special overlay network called `ingress` and installs rules on *every* node so the published port is open there. A request to any node, replica or not, is destination-NAT'd into the ingress network and load balanced by IPVS to a task somewhere in the cluster. Operationally this is excellent: an external load balancer or DNS record can point at every node without tracking placement. The well-known drawback is that the client's source IP is lost — traffic is source-NAT'd on the way in, so access logs show an internal address and any IP-based logic breaks. Remedies: terminate at a proxy that adds `X-Forwarded-For` or the PROXY protocol, or publish with `mode=host`, which binds the port directly on nodes running a task, skips the mesh and preserves the client address. Host mode means only those nodes answer, so you need an external load balancer that knows where tasks are, and it caps you at one task per node per port. ## Diagnosing - `docker service inspect --format '{{json .Endpoint}}' <svc>` shows the VIP and publish mode. - `docker network inspect <overlay>` on a node running a task lists peers and container endpoints. - From a container on the network, `nslookup <service>` (VIP) and `nslookup tasks.<service>` (task IPs) separate DNS problems from load-balancing problems. - If cross-host traffic fails but same-host works, suspect UDP 4789 blocked or an MTU problem: VXLAN adds 50 bytes of header, so a path that cannot carry the resulting frames produces the classic symptom of small requests succeeding and large ones hanging. ## Summary sentence for an interview Overlay is VXLAN over UDP 4789 with gossip on 7946; discovery is DNS to a virtual IP with IPVS behind it; published ports use the ingress network so every node accepts and forwards — at the price of the client source IP unless you publish in host mode.
- Application logs show every request coming from an internal address instead of the real client IP. Why, and what are the options?Ingress-mode publishing source-NATs traffic as it enters the routing mesh, so the backend sees a cluster-internal address. Either terminate the connection at a reverse proxy that records the real address in X-Forwarded-For or the PROXY protocol before forwarding, or publish the port with mode=host so tasks bind the node's port directly and the original source IP survives — accepting that only nodes running a task will answer.
- When would you choose dnsrr endpoint mode over the default VIP mode?When the client needs to see individual task addresses: a client-side load balancer, a protocol that wants to pin connections, or a clustered application whose members must find each other. It is also required with host-mode publishing. The cost is that clients cache DNS records, so removed tasks can keep receiving traffic until the cache expires.
The routing mesh is a switchboard in every branch office: dial the number at any branch and the call is routed to whichever office actually has the person, at the cost of the recipient seeing the switchboard's number.
saying these in an interview costs you the question
- Thinking a published port only works on nodes running a replica
- Claiming overlay traffic is encrypted by default
- Confusing the VIP with a container IP, or expecting to reach tasks by pinging the VIP
- Ignoring MTU when overlay traffic hangs on large payloads
- Assuming X-Forwarded-For appears automatically through the routing mesh