skip to content

After a Linux host is moved behind an encapsulating tunnel, small requests and ICMP echoes work fine, but SSH sessions freeze right after login and large transfers stall completely. Explain the mechanism that produces this size-dependent failure and how you would confirm it.

level: seniorimportance: should knowfreq 48%

answer

  1. small works, large does not
  2. the headers you forgot to subtract
  3. the message that never arrived
  4. don't-fragment plus a silent drop
  5. a threshold, not a slowdown

basics

~20 s

Encapsulation shrinks the usable payload below the interface MTU, so full-size TCP segments are too big for the path. When the ICMP messages that would report this are filtered, the sender never learns and simply retransmits forever — a path MTU black hole.

solid answer

~50 s

This is the classic MTU black hole. The interface still advertises an MTU of 1500, so TCP picks a maximum segment size to match and the sender emits full-size segments with the don't-fragment bit set. Once the tunnel adds its encapsulation header, those packets no longer fit the real path, so a device on the way drops them. It is supposed to signal that with an ICMP "fragmentation needed" message carrying the correct MTU, which the sender would use to shrink its segments — path MTU discovery. When a firewall filters that ICMP, the sender learns nothing and just retransmits the same oversized segment until the connection stalls. Small packets are under the limit, which is why ping and the SSH handshake succeed while the first big burst of data dies. Confirm it by sending progressively larger probes with fragmentation prohibited and finding the size at which they stop getting through.

code

bash · 3 lines
bash
ping -M do -s 1472 -c 1 10.20.0.9
ping -M do -s 1372 -c 1 10.20.0.9
ip link set dev wg0 mtu 1400

go deeper

for a junior

Recognise the signature: tiny requests succeed while big transfers hang. Know that MTU is the largest packet an interface will send and that tunnels leave less room than you expect.

for a middle

Explain how TCP derives its segment size from the interface MTU, why the don't-fragment bit turns an oversized packet into a drop, and what the ICMP fragmentation-needed message is meant to do about it.

for a senior

Diagnose it under pressure with a size sweep, distinguish a black hole from ordinary loss, and pick the right remedy — corrected interface MTU, MSS clamping at the gateway, or probing — then persist it in configuration.

for a principal

Own the policy that prevents it fleet-wide: an ICMP filtering standard that permits the messages path MTU discovery needs, and an MTU convention for every overlay so that encapsulation overhead is accounted for by design rather than discovered in an incident.

## Why the failure is size-dependent Every interface has an MTU — the largest payload a single frame may carry, conventionally 1500 bytes on Ethernet. TCP does not guess; when a connection is established each side advertises a maximum segment size derived from the MTU of the interface it will send on. For IPv4 that is the MTU minus 20 bytes of IP header and 20 bytes of TCP header, so a 1500-byte MTU yields an MSS around 1460. Now put a tunnel in the path. Encapsulation — VXLAN, GRE, IPsec, WireGuard, PPPoE — wraps the original packet in additional headers, so the bytes actually available for the inner payload are fewer than the outer link's MTU. If the host's interface MTU was never reduced to account for that, TCP will happily build segments that cannot survive the journey. The result is a sharp threshold, not a slowdown. Packets below the real limit go through untouched; packets above it are destroyed. That is exactly the observed signature: ICMP echoes are small, the SSH key exchange and login exchange are small, and everything works — then the first full-size data segment (the shell sending a large chunk of output, or a file transfer's first burst) crosses the threshold and vanishes. ## Why nobody is told IPv4 packets carrying TCP normally set the don't-fragment bit, and Linux does this by default so that path MTU discovery can work. A router that must forward a too-large packet with that bit set is required to drop it and return an ICMP destination-unreachable message of the *fragmentation needed* variety, which includes the MTU that would have fit. The sender caches that value for the destination and re-segments accordingly. The mechanism is elegant and, when it works, invisible. It breaks when the ICMP never arrives. Firewalls configured by someone who decided "ICMP is a security risk, drop it all" remove exactly the message the mechanism depends on. Asymmetric routing, an intermediate device that silently drops rather than reporting, or a tunnel endpoint that generates the message from an address the filter rejects all produce the same outcome. The sender has no feedback loop at all: it retransmits the identical oversized segment, the same device drops it again, and the connection hangs until it times out. This is the *path MTU black hole* — the traffic disappears with no error anywhere, which is why it is such a memorable interview scenario. IPv6 sharpens the same trap: routers on the path are never permitted to fragment, so the entire scheme rests on the ICMPv6 *packet too big* message. Filtering ICMPv6 does not merely degrade IPv6, it breaks it. ## Confirming it The probe is a size sweep with fragmentation prohibited. Send an ICMP echo with the don't-fragment behaviour forced and a specified payload size, then binary-search the size until you find the largest that gets a reply. Add 28 bytes — 20 of IPv4 header, 8 of ICMP header — to convert payload size to packet size, and you have the path MTU. If 1472 bytes of payload fails while 1372 succeeds, the path carries roughly 1400 bytes, not 1500, and the deficit is your encapsulation overhead. ```bash ping -M do -s 1472 -c 1 10.20.0.9 # fails: too big for the tunnel ping -M do -s 1372 -c 1 10.20.0.9 # succeeds: fits ``` Two corroborating observations strengthen the diagnosis. First, the failure is direction- and size-correlated, not host-correlated: the same peer works for small exchanges. Second, a transfer that is artificially limited to small writes completes while a bulk transfer does not. ## Fixing it There are three honest fixes, in rough order of preference. **Set the interface MTU to the truth.** If the path can carry 1400 bytes, the tunnel interface should say 1400, so TCP negotiates a segment size that fits from the start. Do it at runtime to test, then put it in the owning network manager's configuration so it survives a reboot — a runtime MTU change is as ephemeral as a runtime address. **Clamp the advertised segment size at the tunnel boundary.** A gateway can rewrite the MSS in passing SYN packets so that endpoints negotiate a size that fits, regardless of what they believe their own MTU to be. This is the standard remedy when you control the tunnel gateway but not the endpoints behind it. **Enable MTU probing on the endpoint.** Linux can detect a black hole and search downward for a working segment size without relying on ICMP, controlled by a sysctl for TCP MTU probing. It is a safety net rather than a design: it costs a stall and some retransmissions before it kicks in. And the standing fix that prevents the whole class of failure: stop filtering the ICMP messages that path MTU discovery requires. Blocking echo requests is a defensible policy choice; blocking fragmentation-needed and packet-too-big is breaking the protocol. ## The reflex to build When small works and large does not, think MTU before you think anything else. It is one of the very few failure modes that is cleanly size-dependent, which makes it fast to confirm and fast to rule out.

  • You control the tunnel gateway but not the machines behind it. What is your fix?
    Clamp the maximum segment size at the gateway. As SYN packets pass through, the advertised MSS is rewritten down to a value that fits the tunnel, so both endpoints negotiate a segment size that survives the path regardless of what their own interfaces claim. It is the standard remedy for exactly this situation, and it costs nothing at steady state because it touches only connection setup. It does not help non-TCP traffic, which still needs a correct MTU.
  • How do you decide what MTU to set on the tunnel interface itself?
    Subtract the encapsulation overhead from the MTU of the underlying path, then verify empirically with a size sweep rather than trusting the arithmetic — nested encapsulation and provider links often take more than the documented figure. Leave a small margin if any further wrapping is possible. Then persist the value in the owning network manager's configuration, because an MTU set with a runtime command disappears on reboot just as an address does.
  • What does enabling TCP MTU probing on a Linux endpoint buy you, and what does it cost?
    It lets TCP detect that a connection has stalled with no ICMP feedback and search downward for a segment size that gets through, so a black hole degrades into a delay rather than a hang. The cost is that the detection is reactive: the connection stalls and retransmits before probing engages, so throughput suffers on every affected flow. Treat it as a safety net for paths you do not control, not as a substitute for setting the MTU correctly.

saying these in an interview costs you the question

  • Blames packet loss or a flaky NIC for a clean size threshold
  • Says all ICMP can be safely blocked at the firewall
  • Thinks routers will just fragment oversized packets anyway
  • Sets the MTU at runtime and never persists it in the config
  • Assumes IPv6 avoids the problem because it has no fragmentation

context