skip to content

How does IPv4 Path MTU Discovery use the Don't Fragment flag to find the largest datagram a path can carry?

level: middleimportance: must knowfreq 42%

answer

  1. start at the first hop's MTU
  2. set DF on every datagram
  3. too-big router drops and reports
  4. Next-Hop MTU in the error
  5. probe upward only rarely

basics

~20 s

The source assumes the path MTU equals its first link's MTU and sets DF on every datagram. A router that cannot forward one drops it and returns an ICMP error carrying the next hop's MTU; the source lowers its estimate and sends smaller.

solid answer

~50 s

Path MTU Discovery (RFC 1191) turns the `DF` flag into a probe. The source starts by assuming the path MTU is the MTU of its own first-hop link and sets `DF` on every datagram to that destination. A router that cannot forward one without fragmenting must drop it and send back an ICMP Destination Unreachable, fragmentation needed and DF set; RFC 1191 added a **Next-Hop MTU** field to that error so the source learns the exact size. The source lowers its estimate for that destination, never below 68 bytes, and repeats until datagrams get through. It keeps setting `DF` so a later, smaller route is noticed, and may retry a larger size only rarely: no sooner than 5 minutes after a too-big error, 10 recommended. TCP then sizes its segments from the estimate; a UDP application must size its own messages.

go deeper

for a junior

Recall the loop: send with DF set, a router that cannot forward drops and reports back, and the sender sends smaller until datagrams get through.

for a middle

Explain each step of RFC 1191: the first-hop starting estimate, the Next-Hop MTU in the error, the per-destination cache, the 68-byte floor and why DF stays set after convergence.

for a senior

Show that the mechanism's weak point is the returning ICMP error: when it is lost or filtered the sender learns nothing, and you should reason about where that signal can die.

for a principal

Weigh classic PMTUD against packetization-layer probing: one trusts the network to report, the other measures with its own acknowledgements, trading a little probing cost for not depending on ICMP.

## Why discover the path MTU at all Every link has an MTU, and the **path MTU** is the smallest MTU on the route between two hosts. A sender that exceeds it either gets fragmented by an IPv4 router, which multiplies loss and costs the receiver reassembly work, or, with `DF` set, gets dropped. A sender that guesses too low wastes header overhead on every packet. **Path MTU Discovery (PMTUD)**, specified for IPv4 in **RFC 1191** (1990), lets the sender learn the right size from the network instead of guessing. ## The RFC 1191 algorithm 1. **Initial estimate.** The source assumes the path MTU equals the known MTU of its **first hop**, for example 1,500 bytes on Ethernet. 2. **Set DF everywhere.** It sends every datagram on that path with the **Don't Fragment** flag set, so no router may fragment it. 3. **A router that cannot forward drops.** If some link along the way has a smaller MTU, the router in front of it discards the datagram and returns an ICMP Destination Unreachable with the code for "fragmentation needed and DF set"; RFC 1191 calls it a **Datagram Too Big** message. 4. **Lower the estimate.** The source reads the size from the message, reduces its estimate for that destination and resends smaller. 5. **Converge.** Steps 3 and 4 repeat, once per narrower link, until datagrams arrive whole. RFC 1191 requires the host to force the process to converge, and never to lower its estimate below **68 bytes**, the IPv4 minimum link MTU. The estimate is cached **per destination** (or per route), because different paths have different bottlenecks. ## The Datagram Too Big signal Before RFC 1191 the ICMP error said only "too big", not how big was allowed. RFC 1191 redefined an unused 16-bit field in that error as **Next-Hop MTU**: the size, IP header included, of the largest datagram the reporting router could forward on that hop. Two details: - A router that predates RFC 1191 sends **zero** in that field. Hosts must cope; RFC 1191 suggests a **plateau table** of common MTUs and recommends taking the greatest plateau below the `Total Length` of the original datagram quoted in the error. - The error quotes the original IP header plus the first 64 bits of its data, which is how the source matches it to a destination and, for TCP, to a connection. The type and code numbers, and which device sends each Destination Unreachable code, belong to ICMP itself; here what matters is that the signal exists and carries a size. ## Keeping the estimate current A path MTU can **shrink** when routing changes, and the host learns that immediately, because it **keeps setting DF** after converging and the next too-big error arrives. It can also **grow**, but nothing announces that: the host has to try a larger datagram, which usually fails. RFC 1191 therefore rate-limits the attempt: | Event | RFC 1191 minimum before trying larger | Recommended | |---|---|---| | A Datagram Too Big message was received | 5 minutes | 10 minutes | | A previous attempt to increase succeeded | 1 minute | 2 minutes | ## Who uses the estimate - **TCP** turns the estimate into its segment size, so the bytes it puts in each segment shrink when the path MTU does; how the MSS option itself is negotiated in the handshake belongs to the TCP header. - A **UDP application** gets no help from UDP: it must learn the estimate from the IP layer and size its own messages, or accept fragmentation. - **Tunnels** are both users and causes: their encapsulation makes the inner path MTU smaller than the outer one. ## What PMTUD depends on, and the IPv6 version Classic PMTUD has a single point of failure: the **ICMP error must reach the source**. If a router fails to send it, rate-limits it, or a filter drops it on the way back, the source keeps sending datagrams that silently vanish, a **PMTUD black hole**. Packetization-layer PMTUD (RFC 4821 for TCP, RFC 8899 for datagram transports) removes that dependency by probing with the transport's own acknowledgements. IPv6 has its own version, **RFC 8201** (obsoleting RFC 1981), using ICMPv6 Packet Too Big. It matters more there: IPv6 routers never fragment, so a source that ignores the signal simply loses packets.

  • What should an IPv4 host do when a Datagram Too Big message carries a Next-Hop MTU of zero?
    Zero means the router predates RFC 1191 and did not report a size. The host must still lower its estimate; RFC 1191 recommends picking the greatest value in its plateau table of common MTUs that is below the `Total Length` of the original datagram quoted in the error, then continuing discovery from there.
  • Why does an IPv4 host keep setting DF after Path MTU Discovery has converged?
    Because routes change. If traffic moves to a path with a narrower link, a DF-marked datagram makes the router there drop it and report the new size at once. If the host stopped setting DF, that router would fragment silently and the host would never learn the path had shrunk. RFC 1191 notes a host may stop setting DF, but normally continues.
  • Why can't classic IPv4 Path MTU Discovery detect that a path's MTU has grown?
    No message says that a bigger datagram would now fit; ICMP only reports when one did not. The host has to send a larger datagram and see whether a too-big error comes back, and that usually fails because paths rarely grow. RFC 1191 therefore allows such attempts no sooner than 5 minutes after a too-big error, recommending 10.

saying these in an interview costs you the question

  • Classic Path MTU Discovery must send special probes before any data flows.
  • The DF flag asks routers to leave fragmenting to the destination.
  • Once found, the path MTU never needs checking again.
  • Hosts try a larger path MTU every few seconds to regain throughput.
  • The router that drops the datagram reports the sender's own MTU.