Containers on a user-defined Docker bridge stall on large responses while small requests succeed — how do you diagnose it?
answer
- Small requests fine, large ones hang
- The handshake is tiny; the payload is not
- Compare two interfaces' frame sizes
- The ICMP that would fix it is filtered
- Set it at network-create time; it is immutable after
basics
~10 sThat signature points at an MTU mismatch: the bridge is 1500 while the host's real egress path is smaller, so full-size frames are dropped. Compare the container's eth0 MTU with the host's egress interface.
solid answer
~50 sConnections that establish and then hang on the first large payload are the classic MTU signature, not a DNS or routing failure — the small SYN and the small request get through, the first full-size segment does not. It happens when the Docker bridge uses a 1500-byte MTU but the path off the host is smaller (a VPN or tunnel interface at, say, 1412), and the ICMP "fragmentation needed" messages that would trigger path-MTU discovery are filtered, so the sender retransmits into a black hole. Diagnose by comparing `ip link show eth0` inside the container with the MTU of the host interface the traffic leaves by, and check `docker network inspect` for an MTU option. Fix it with `docker network create --opt com.docker.network.driver.mtu=1412 …`; options are immutable, so that means recreating the network and re-attaching.
code
bash · 12 lines# MTU seen inside a container on the network
docker run --rm --network digest-net alpine ip -o link show eth0
# MTU of the interface the host actually uses to reach the peer
ip route get 10.42.7.19
ip -o link show
# What the network was created with (empty means unset)
docker network inspect -f '{{index .Options "com.docker.network.driver.mtu"}}' digest-net
# Fix: recreate the network with an explicit MTU and re-attach
docker network create --opt com.docker.network.driver.mtu=1412 digest-net-fixedgo deeper
Know that every network interface has a maximum frame size, and that a container's eth0 gets its value from the Docker network it is attached to. Being able to run ip link inside a container is enough at this level.
Explain why a mismatch produces a stall rather than an error: the handshake fits, the first full-size segment does not, and the ICMP message that would trigger path-MTU discovery is usually filtered.
Demonstrate the diagnosis order — connection opened, so not resolution; compare container and egress MTUs; inspect the network's options — and know that the fix means recreating the network because driver options are immutable.
Own the fleet-level answer: standardise the MTU wherever networks are created, add a start-up check that fails loudly on a mismatch, and re-validate whenever host networking changes rather than waiting for intermittent build timeouts.
## Recognising the signature The symptom set is very specific and worth memorising, because it is easy to misdiagnose as an application bug: * TCP connections **establish** normally — the handshake packets are tiny. * Small requests and small responses work — `curl` against a health endpoint is fine. * Anything that pushes a full-size segment hangs: a TLS handshake carrying a large certificate chain, a multi-megabyte download, a big POST body. * The hang ends in a timeout, not a connection refused, and the application logs show nothing but a stalled read. That is an MTU problem until proven otherwise. ## Why the MTU matters here MTU is the largest payload a link will carry in one frame. A Docker bridge is an ordinary Linux bridge, and containers attached to it get an `eth0` whose MTU comes from the network's configuration — 1500 unless told otherwise. If the host then sends that packet out over an interface with a smaller MTU (a WireGuard or IPsec tunnel, a cloud overlay, a PPPoE link), the packet is too big for the next hop. In the normal case the router replies with an ICMP "fragmentation needed" message, the sender lowers its path MTU and everything recovers. In practice that ICMP is very often dropped by a firewall or a cloud security group, and you get a **path-MTU black hole**: the sender keeps retransmitting a segment nobody will forward. Because the container's own stack knows only about its 1500-byte `eth0`, the whole thing looks, from inside the container, like the peer went silent. ## Worked example A 47-node CI fleet runs the integration suite of a Ruby email-digest builder. Each job creates a per-build network, starts the worker and a Postgres container on it, and the worker pulls fixtures from an internal artefact host reached over a WireGuard tunnel whose interface MTU is 1412. Small API calls in the suite pass; the step that downloads a 38 MB fixture archive hangs and the job is killed by its 15-minute timeout. It reproduces on 11 of the 47 nodes — exactly the ones whose egress goes through the tunnel — which is what makes people blame "flaky nodes" instead of the network settings. ## Diagnosis, in order 1. **Establish that it is not name resolution.** The connection opens, so the peer's address was resolved and reachable. That already rules out a whole class of causes. 2. **Read the MTUs on both sides of the hop.** Inside a container on the network: `docker run --rm --network digest-net alpine ip -o link show eth0`. On the host: `ip -o link show` and find the interface the route to the peer uses (`ip route get <peer-ip>`). A container at 1500 leaving through a 1412 interface is the finding. 3. **Check what the network was created with.** `docker network inspect -f '{{index .Options "com.docker.network.driver.mtu"}}' digest-net` — an empty result means the option was never set. 4. **Confirm with a sized probe.** Send progressively larger payloads and find the cliff; the boundary will sit at the smaller MTU minus headers. 5. **Watch the retransmits.** A packet capture on the host shows the same large segment being retransmitted with no acknowledgement and no ICMP reply coming back. ## The fix Set the MTU on the network, at creation: ``` docker network create --opt com.docker.network.driver.mtu=1412 digest-net ``` A Docker network's options cannot be edited after creation, so applying this to an existing network means creating a replacement and re-attaching (or, for a stack, recreating it). For containers on the **default** bridge the equivalent is the daemon-wide `"mtu"` key in `/etc/docker/daemon.json` followed by a dockerd restart — which is exactly the granularity argument for not putting long-lived workloads on docker0 in the first place. Whatever the engine's default happens to be on a given version, do not assume a bridge you create has inherited the host's lower MTU. Verify it with `ip link` inside a container on that network; on a fleet, make that verification part of the node's build so a new image or kernel cannot silently change it. ## Prevention on a fleet * Bake the MTU option into whatever creates networks, so every node agrees. * Add a start-up check that compares the container MTU against the egress interface and fails loudly rather than producing intermittent hangs. * When a host's networking changes — a new VPN, a move to a tunnelled network — re-check the value; nothing in Docker will tell you it has gone stale. ## What not to say Do not reach for "the application is slow" or "the registry is flaky", do not raise the MTU hoping to push more through, and do not try to change the value on a live network — there is no command that does it.
- Why does the TCP handshake succeed if the path cannot carry the packets?Because the handshake packets are tiny. SYN, SYN-ACK and ACK carry no payload, and the first request may also fit under the smaller MTU. The path only breaks when a full-size segment appears, which is why the failure looks like a mid-transfer hang rather than a connection error, and why health checks against a small endpoint keep reporting the service as fine.
- Can you change the MTU of an existing Docker network?No. Driver options are fixed when the network is created — there is no command to update them. You create a replacement network with the right `--opt com.docker.network.driver.mtu` value, attach the containers to it, and remove the old one. For the default bridge the value comes from the daemon's `mtu` setting and changing it requires restarting dockerd.
- Only some hosts in the fleet show the problem. What does that tell you?That the bridge configuration is uniform but the egress path is not. The affected hosts route the traffic over a lower-MTU link — a tunnel, a different subnet, a different cloud network — while the rest go out over a plain 1500-byte interface. Diagnose per host by checking which interface `ip route get` picks for the peer, and fix by standardising the network's MTU to the lowest path in the fleet.
It is a low doorway on the way out of the building: people carrying nothing walk through fine, and only the ones carrying a wardrobe get stuck — and nobody sends word back that the doorway was the problem.
saying these in an interview costs you the question
- Blames the application before comparing interface MTUs
- Assumes a created bridge inherits the host's lower MTU
- Tries to change MTU on an existing Docker network in place
- Calls it a DNS failure although the connection opened
- Raises the MTU instead of lowering it to the path
- Treats intermittent per-host failures as flaky hardware