skip to content

An IPv4 TCP connection completes its handshake and small requests, but large responses never arrive; how can Path MTU Discovery cause this, and what fixes it?

level: seniorimportance: should knowfreq 40%

answer

  1. size decides what gets through
  2. DF packets dropped mid-path
  3. the ICMP error never returns
  4. sender retransmits the same size
  5. probe without ICMP: RFC 4821

basics

~20 s

The server sends full-size DF datagrams; a router before a narrower link drops them, but its fragmentation-needed ICMP error is filtered or never sent, so the server never shrinks. Fix the ICMP path, or use packetization-layer PMTUD (RFC 4821, RFC 8899).

solid answer

~50 s

This is a **PMTUD black hole**. Handshake segments and small requests fit the narrowest link, so they pass. The server's large responses go out at its first-hop MTU with `DF` set; the router in front of a smaller link drops them and should return a fragmentation-needed ICMP error, but that error is lost: a firewall near the server drops inbound ICMP, the router never sends it, or there is no route back. The server never lowers its estimate and keeps retransmitting the same size until the connection times out. Confirm it by finding the largest DF datagram that passes and checking whether any too-big error reaches the sender. The real fix is to let that ICMP through; RFC 8900 says operators MUST make PMTUD work. Packetization-layer PMTUD (RFC 4821 for TCP, RFC 8899 for datagram transports) probes without needing ICMP; clamping TCP's MSS at the narrow point is the common workaround.

go deeper

for a junior

Recall that a link with a smaller MTU drops oversized DF packets and that the sender relies on an ICMP error to learn it must send smaller ones.

for a middle

Explain why small exchanges succeed and large ones hang: everything below the path MTU passes, and the sender never hears that its full-size datagrams are being dropped.

for a senior

Demonstrate diagnosis by size and direction, name where the too-big error dies, and rank the fixes: restore the signal, adopt packetization-layer PMTUD, and treat MSS clamping as a workaround.

for a principal

Argue the policy angle: filtering ICMP wholesale trades a small attack surface for silent, hard-to-debug outages, and endpoints that probe for themselves stop depending on every network operator getting it right.

## The symptom, and why size decides it The pattern is distinctive: the TCP handshake completes, a small request gets a small answer, `ping` succeeds, and then a large response, a big page, a file or a TLS certificate chain, stalls until the connection times out. RFC 2923 describes exactly this: pings and some interactive connections work, bulk transfers fail with the first large packet. **Size is the variable.** Everything below the path MTU gets through; everything above it vanishes. That points at the path MTU and at **Path MTU Discovery (PMTUD)**, not at routing, DNS or the application. ## How the black hole forms Take a server on an Ethernet segment (MTU 1,500) and a client somewhere behind a link with a smaller MTU, say 1,400 bytes. 1. The server's TCP sizes its segments for 1,500-byte datagrams and, doing classic PMTUD (RFC 1191), sets **`DF`** on all of them. 2. The router in front of the 1,400-byte link cannot forward a 1,500-byte DF datagram. It **drops** it and should send an ICMP Destination Unreachable, fragmentation needed and DF set, carrying a Next-Hop MTU of 1,400, back to the **server**. 3. That error **never arrives**. 4. The server sees only a missing acknowledgement, retransmits the **same 1,500-byte segment**, and it is dropped again. The connection eventually times out. 5. The client's requests are small and the server's handshake segments are small, so traffic in the other direction looks perfectly healthy. The black hole is **directional**: it exists for whichever side sends large datagrams across the narrow link, and the ICMP error must travel back to that side. ## Where the signal gets lost RFC 2923 and RFC 8900 section 3.8 list the usual causes: | Cause | Why the too-big error does not reach the sender | |---|---| | A firewall drops all inbound ICMP | The error is discarded near the sender, often by a policy written as hardening | | A stateful filter mismatches the error | RFC 8900 notes broken implementations drop it because its source address is not part of an existing flow | | The router does not send it | Misconfiguration, or ICMP rate limiting under load | | Anycast | The error is routed to a different instance of the anycast address than the one that sent the datagram | | Unidirectional routing | The router has a route to the destination but none back to the source | ## Confirming it from the protocol - Reproduce the size dependence: DF-marked datagrams below some size get through; above it, nothing does. The boundary is the real path MTU. - Watch the sender's side: full-size segments are retransmitted unchanged, and **no too-big ICMP error arrives** in response. - Remember what a successful `ping` proves: only that small packets cross the path. Default echo requests are far below any MTU. - Look for links that shrink the MTU: tunnels and encapsulations, PPP over Ethernet, or any segment configured below 1,500. ## Fixes, in order of preference | Fix | What it does | Trade-off | |---|---|---| | Let the error through and make routers send it | Restores classic PMTUD; RFC 8900 says operators MUST ensure PMTUD works | Needs every filter and router on the return path to cooperate | | Packetization-layer PMTUD | The transport probes larger sizes and learns from its own acknowledgements (RFC 4821 for TCP, RFC 8899 for datagram transports) | Must be enabled at the endpoints; a lost probe costs a little time | | Clamp TCP's MSS at the narrow point | A middlebox rewrites the MSS in passing handshakes so segments fit | Helps TCP only, and is a workaround for a broken signal | | Lower the interface MTU on the hosts that send large datagrams | Makes every datagram fit by construction | Static, and wastes capacity on paths that could carry more | **Packetization-layer PMTUD (PLPMTUD)** is the robust answer: RFC 8900 notes the black-hole problem is specific to classic PMTUD and does not occur when the estimate comes from PLPMTUD. RFC 2923 also describes **black-hole detection** in TCP: after several timeouts on full-size segments, the sender tries smaller ones. ## What not to conclude - Clearing `DF` on all of the sender's traffic is not a clean fix: it lets IPv4 routers fragment again, with all the loss and filtering problems that brings. RFC 2923's black-hole detection may drop `DF` for a while as a fallback, and it asks hosts to log detected black holes so the network gets fixed. - Blaming the client because its requests reached the server mixes up the two directions.

  • How does packetization-layer PMTUD find the path MTU without relying on ICMP?
    RFC 4821 has the transport probe the path with progressively larger packets. A probe that is acknowledged proves that size fits; a probe that is lost while ordinary traffic gets through suggests it does not. ICMP errors, when they arrive, are only extra information. RFC 8899 applies the same idea to datagram transports, which must add a way to confirm that probes arrived.
  • Why does ping to the server keep working during an IPv4 PMTUD black hole?
    A default ICMP echo request and its reply are a few dozen bytes, far below any link's MTU, so they cross the narrow link without needing fragmentation or a too-big error. A successful `ping` proves reachability for small packets only; it says nothing about whether full-size datagrams fit the path.
  • Why does the black hole usually show up in only one direction?
    Path MTU Discovery is done by each sender for its own datagrams. The client's requests are small and fit the narrow link, so its direction works. The server's responses are full-size, so they hit the narrow link and need a too-big error to travel back to the server. If that one return path loses ICMP, only the server's direction breaks.

saying these in an interview costs you the question

  • If ping works, the path MTU cannot be the problem.
  • Blocking all inbound ICMP is harmless hardening, since it only carries diagnostics.
  • The router before the narrow link should fragment DF datagrams anyway.
  • Clearing DF on the server is the clean, standard fix.
  • The MSS agreed in the handshake guarantees every segment fits the path.