skip to content

Two servers on one Ethernet segment use a 9,000-byte MTU; small requests work but bulk TCP transfers stall through a switch passing only standard frames — what is happening, and why does nothing report it?

level: seniorimportance: should knowfreq 30%

answer

  1. MTU is not frame size
  2. both ends advertise large segments
  3. layer 2 has no error message
  4. probe with known sizes
  5. one size per segment

basics

~20 s

The switch silently discards frames above its maximum frame size. Small packets fit, but both servers derive large TCP segments from their 9,000-byte MTU. No router is on the path, so no ICMP error ever tells the sender to shrink.

solid answer

~50 s

The **MTU** is the largest IP packet a link carries; the frame adds 18 bytes, so a 9,000-byte MTU means frames up to 9,018 bytes untagged. **Jumbo frames** are outside the IEEE 802.3 standard sizes, and 9,000 is a common convention, not a rule. The TCP handshake and small requests fit in standard frames, but each server advertises a maximum segment size derived from its own 9,000-byte MTU, so bulk data leaves in frames of about 9,018 bytes. A switch port limited to standard frames discards them as oversize. Ethernet has no fragmentation and no error message, and Path MTU Discovery depends on an ICMP error from a **router** reducing the MTU: inside one segment there is none. TCP retransmits the same size until it gives up. Confirm it with don't-fragment probes of known sizes. RFC 1042 states the fix: on one network, all hosts must use the same MTU.

go deeper

for a junior

Recall that the MTU is the largest IP packet a link carries, 1,500 on standard Ethernet, and that jumbo frames are a non-standard larger size every device must support.

for a middle

Explain MTU versus frame size, the 18 bytes of header and FCS, and why a switch can only forward or drop a frame. Work the 1,472-byte probe arithmetic.

for a senior

Diagnose the size-dependent stall: the handshake passes, large segments vanish, no ICMP arrives because no router is involved. Show how you would find the hop and why the whole segment must agree.

for a principal

Decide where jumbo frames are worth their operational cost: dedicated storage or replication segments with controlled membership versus general networks, and the router boundaries you need between the two.

## MTU and frame size are different numbers The **maximum transmission unit (MTU)** is the largest IP packet a link carries, header included. The Ethernet **frame** wraps it in 14 bytes of header and a 4-byte frame check sequence, so it is always 18 bytes bigger. | Configuration | MTU (IP packet) | Largest frame, untagged | With one 802.1Q tag | |---|---|---|---| | Standard Ethernet | 1,500 | 1,518 | 1,522 | | A common jumbo setting | 9,000 | 9,018 | 9,022 | RFC 894 fixes Ethernet's data field at 1,500 bytes, and RFC 2464 makes 1,500 the default IPv6 MTU on Ethernet. **Jumbo frames** are a widely supported extension outside the IEEE 802.3 standard frame sizes. RFC 6762 calls 9,000 bytes "the maximum payload size of an Ethernet 'Jumbo' packet" while citing a non-IETF reference; treat 9,000 as a convention. Implementations also disagree about what their configured number counts: some mean the IP MTU, others the whole frame, and some leave room for tags. Two devices both "set to 9,000" can therefore accept different frame sizes. Frames also grow for reasons other than the host's MTU: a VLAN tag adds 4 bytes, and tunnels or overlays add whole outer headers. Each of those must fit the largest frame every hop accepts, which is why mismatches often surface only after a change elsewhere on the path. ## Why the transfer stalls 1. ARP, the TCP handshake and short requests are all well under 1,500 bytes, so they cross the switch. 2. During the handshake each server advertises a TCP maximum segment size taken from its own MTU: 9,000 minus 40 bytes of IPv4 and TCP headers, 8,960. 3. Bulk data now leaves as 9,000-byte IP packets in 9,018-byte frames. 4. The switch port, limited to standard frames, receives an oversize frame and discards it. 5. The sender's TCP retransmits the same segment, which is dropped the same way, until the connection times out. The symptom is size-dependent: everything small works and everything bulky hangs. That pattern is shared with path-MTU black holes, but the mechanism here is different. ## Why nothing reports the drop - **Ethernet has no fragmentation.** A bridge cannot split a frame; it forwards it whole or drops it. - **Ethernet has no error message.** A discarded frame produces nothing on the wire, just as a failed FCS check does. - **ICMP needs a router.** An IPv4 router sends a fragmentation-needed error when a packet with Don't Fragment set must go out on a link with a smaller MTU. A switch forwarding within one segment has no IP role, and a router typically drops an oversize frame on its receiving interface before IP ever sees it. So Path MTU Discovery has nothing to learn from. - **The handshake hides it.** Both ends agree on large segments precisely because both are configured large; nothing in the handshake asks the switch. ## Confirming it 1. Send probes with Don't Fragment set and a chosen payload size between the two servers. 2. Over IPv4 an ICMP echo carries 20 bytes of IP header and 8 of ICMP header, so a 1,472-byte payload makes exactly a 1,500-byte packet, and 8,972 makes 9,000. 3. If 1,472 passes and anything larger disappears without an error, a hop on the segment is passing only standard frames. 4. Check the oversize or giant-frame counters on each switch port along the path to find it. Running the probe tool itself is a diagnostics topic; the reasoning is the sizes. ## Fixing it and what to standardise RFC 1042 is blunt: "on any particular network all hosts must use the same MTU." A segment is one size, so either: - raise the maximum frame size on **every** layer 2 hop between the servers (each switch port, any virtual switch or bridge, any device in bridging mode), or - set the servers back to 1,500 for that segment. Where a jumbo segment meets a standard one, a **router** must sit between them so fragmentation or Path MTU Discovery can work. Jumbo frames pay off mainly for bulk storage and replication traffic: at 9,000 bytes, 9,000 of every 9,038 byte times on the wire carry payload, about 99.6 percent, against about 97.5 percent at 1,500, and the host handles a sixth as many packets.

  • If only one server is set to 9,000 and the other stays at 1,500, does TCP still work?
    Usually, yes. Each side advertises the largest segment it can receive, so the 1,500 server advertises 1,460 and the jumbo server never sends it more. Traffic not governed by TCP's segment size, such as large UDP datagrams, still leaves the jumbo server in big frames that the other side's interface may discard, depending on the implementation. The mismatch is hidden, not gone.
  • Why not enable jumbo frames on every network by default?
    Every layer 2 device in the segment must agree, including virtual switches and bridges, and each new device is a chance to miss one. Traffic leaving the segment drops back to 1,500 at a router, which must fragment or rely on Path MTU Discovery. The gain is fewer packets per byte, which matters for bulk storage traffic and little for typical request-response traffic.

saying these in an interview costs you the question

  • The switch will send an ICMP message telling the server to use smaller packets.
  • A switch fragments oversize frames into standard ones.
  • 9,000 bytes is the jumbo frame size defined by IEEE 802.3.
  • A 9,000-byte MTU means the whole frame is 9,000 bytes.
  • If the TCP handshake succeeds, the path's MTU must be fine.
  • Setting both servers to 9,000 is enough for jumbo frames to work.