skip to content

Fragmentation and MTU

Routers split datagrams bigger than the next link's MTU and only the destination reassembles; DF plus Path MTU Discovery avoids it. A PMTUD black hole is the classic 'small works, large hangs' story.

on this pageshow

questions

6

In IPv4, what is the MTU, and what happens to a datagram that is larger than the next link's MTU?

level: juniorimportance: must knowfreq 52%

answer

  1. a property of each link
  2. counts the IP header, not framing
  3. router splits, destination rejoins
  4. unless Don't Fragment is set

basics

~20 s

The MTU is the largest IP datagram, header included, that a link carries in one frame: 1,500 bytes on Ethernet. An IPv4 router fragments a bigger datagram and only the destination reassembles it; with Don't Fragment set, the router drops it instead.

solid answer

~50 s

The **maximum transmission unit** is a property of each link: the biggest IP datagram, counting the IP header but not the link's own framing, that fits in one frame. On Ethernet it is 1,500 bytes (RFC 894). When an IPv4 router has to forward a datagram onto a link with a smaller MTU, RFC 791 lets it **fragment** the datagram: it cuts the data into pieces, gives each piece its own IPv4 header carrying the original `Identification`, a `Fragment Offset` and the More Fragments (`MF`) flag, and forwards the pieces independently. No router puts them back together (RFC 1812 forbids it); only the **destination** reassembles. If the sender set the `DF` (Don't Fragment) flag, the router discards the datagram instead and sends an ICMP error back to the source, which is the signal Path MTU Discovery is built on.

go deeper

for a junior

Recall that the MTU belongs to a link and counts the IP header, that Ethernet's is 1,500 bytes, and that only the destination host reassembles fragments.

for a middle

Explain what each fragment carries: its own header, the shared Identification, an offset in 8-byte units and the More Fragments flag, and how the DF flag turns fragmenting into dropping with an ICMP error.

for a senior

Show that fragmentation is the fallback, not the plan: production traffic sets DF and sizes itself to the path, because fragments get filtered, multiply loss and cost the receiver reassembly state.

for a principal

Frame the design choice: hop-by-hop fragmentation let IPv4 cross mismatched links without senders knowing the path, while IPv6 moved fragmentation to the source to keep routers simple.

## What the MTU measures The **maximum transmission unit (MTU)** is the size of the largest IP datagram a particular link can carry in a single frame. Two details matter in interviews: - It is a property of **one link**, not of a path or a host. A datagram crossing five links meets five MTUs, and the smallest of them is the **path MTU**. - It counts the **IP datagram** — IP header plus everything inside it — and **not** the link layer's own header and trailer. Ethernet's 1,500 bytes (RFC 894) is the room left for the datagram inside the frame. A few values the specifications record (RFC 1191 lists them in its table of common MTUs): | Link or limit | MTU in bytes | Where it is stated | |---|---|---| | Ethernet | 1,500 | RFC 894 | | IEEE 802.3 with 802.2 framing | 1,492 | RFC 1042 | | X.25 networks | 576 | RFC 1191's table | | Smallest any IPv4 link may have | 68 | RFC 791 | ## What an IPv4 router does with an oversized datagram When the next hop's MTU is smaller than the datagram, RFC 791 gives the router two outcomes, chosen by the **Don't Fragment** (`DF`) flag the sender set: 1. **DF clear — fragment it.** The router splits the data portion into pieces whose size, header included, fits the next link. Every piece except the last carries a multiple of 8 data bytes, because the `Fragment Offset` field counts in 8-byte units. 2. Each piece becomes a complete IPv4 datagram with its **own header**: the same `Identification`, source, destination and `Protocol` as the original, its own `Fragment Offset`, the **More Fragments** (`MF`) flag set on every piece but the last, and a recomputed `Total Length` and `Header Checksum`. 3. The pieces travel on **independently**. They may even take different routes, and a later router with a still smaller MTU may fragment a fragment again. 4. **DF set — drop it.** The router may not fragment, so it discards the datagram and returns an ICMP Destination Unreachable error saying fragmentation was needed. The source is expected to send smaller datagrams; that loop is Path MTU Discovery. ## Why only the destination reassembles RFC 1812, the router requirements, says plainly that a router **MUST NOT** reassemble a datagram before forwarding it. The reasons: - Fragments of one datagram can travel by **different paths**, so no single router is guaranteed to see them all. - Reassembly needs **state and buffer memory** per datagram and a timer, which is a poor fit for a device forwarding millions of unrelated packets. - Putting the datagram back together mid-path would only have to be undone at the next small link. So the **destination host** holds the pieces, keyed by source, destination, protocol and `Identification`, until it has every byte, and then hands the whole datagram to TCP, UDP or whichever protocol the `Protocol` field names. If any piece never arrives, the destination eventually gives up and discards the rest. ## The two floor values: 68 and 576 Two numbers are easy to mix up: - **68 bytes** is the smallest MTU any IPv4 link may have. RFC 791 explains the arithmetic: a header can be 60 bytes long, and the smallest fragment carries 8 data bytes, so a router must always be able to forward 68 bytes without fragmenting further. - **576 bytes** is the smallest **datagram** every destination must be able to **receive**, whole or in fragments (RFC 791, and RFC 1122 says the reassembly limit `EMTU_R` must be at least 576). It is not a link MTU. ## IPv6 does this differently Router fragmentation is an IPv4 behaviour. In IPv6, RFC 8200 says fragmentation is performed **only by the source**; a router that cannot forward a too-big IPv6 packet drops it and reports back. That difference is one reason IPv4 answers about fragmentation should always name the version. ## What to take into the interview The MTU is per link and counts the IP header. An IPv4 router fragments a datagram that does not fit, unless `DF` forbids it, in which case it drops and reports. Reassembly happens only at the destination. Fragmentation is a fallback that keeps IPv4 working across mismatched links, not something well-behaved traffic aims for: most senders set `DF` and size their datagrams to the path instead.

  • Why must every IPv4 link be able to carry at least 68 bytes without fragmenting?
    RFC 791 derives it: an IPv4 header can be up to 60 bytes when options are present, and the smallest fragment carries 8 data bytes, the unit the `Fragment Offset` counts in. A router therefore always needs room for 60 + 8 = 68 bytes to make progress; a link smaller than that could not carry even one minimal fragment.
  • Is 576 bytes the minimum MTU of an IPv4 link?
    No. 576 is the smallest datagram every destination must be able to receive, whole or as fragments to reassemble (RFC 791; RFC 1122 calls the reassembly limit `EMTU_R` and requires at least 576). The smallest link MTU is 68 bytes. A datagram of 576 bytes may well cross a link narrower than that, as fragments.

A removal truck too tall for a low bridge: the load is moved into smaller vans, each labelled with the shipment number, its position in the load and whether more vans follow. Nobody unpacks at the depots on the way; the boxes are reassembled only at the new house, and if one van never arrives the shipment is incomplete.

saying these in an interview costs you the question

  • The MTU counts the link's own frame header, not just the IP datagram.
  • Routers reassemble fragments before forwarding the datagram onward.
  • A router still fragments a datagram that has Don't Fragment set.
  • IPv4 routers never fragment; only the sending host may split a datagram.
  • 576 bytes is the minimum MTU every IPv4 link must support.
open as a page

How does IPv4 Path MTU Discovery use the Don't Fragment flag to find the largest datagram a path can carry?

level: middleimportance: must knowfreq 42%

basics

~20 s

The source assumes the path MTU equals its first link's MTU and sets DF on every datagram. A router that cannot forward one drops it and returns an ICMP error carrying the next hop's MTU; the source lowers its estimate and sends smaller.

open as a page

An IPv4 router must forward a 4,000-byte datagram with a 20-byte header onto a 1,500-byte MTU link; which fragments does it send?

level: middleimportance: should knowfreq 32%

basics

~20 s

Three fragments: 1,480 data bytes at offset 0 with MF set, 1,480 at offset 185 with MF set, and the last 1,020 at offset 370 with MF clear, all sharing one Identification. Only the destination reassembles; losing any fragment loses the datagram.

open as a page

An IPv4 TCP connection completes its handshake and small requests, but large responses never arrive; how can Path MTU Discovery cause this, and what fixes it?

level: seniorimportance: should knowfreq 40%

basics

~20 s

The server sends full-size DF datagrams; a router before a narrower link drops them, but its fragmentation-needed ICMP error is filtered or never sent, so the server never shrinks. Fix the ICMP path, or use packetization-layer PMTUD (RFC 4821, RFC 8899).

open as a page

Why does RFC 8900 call IP fragmentation fragile, and what should a protocol that sends large IPv4 datagrams do instead?

level: seniorimportance: should knowfreq 24%

basics

~20 s

Fragments lack ports, which confuses firewalls, NAT and load balancers; one lost fragment loses the datagram; IPv4's 16-bit Identification wraps at high rates and splices data. RFC 8900 says new protocols should size to the path MTU instead.

open as a page

How do tiny-fragment and overlapping-fragment attacks use IPv4 fragmentation to slip past packet filters, and how are they blocked?

level: seniorimportance: nice to knowfreq 17%

basics

~20 s

A stateless filter judges only the first fragment. A tiny first fragment pushes TCP's flags into the second; an overlapping later fragment rewrites the header at reassembly. Filters drop TCP fragments at offset 1 and too-short first fragments (RFC 1858, RFC 3128).

open as a page