When a VXLAN overlay over IPv4 gives tenant hosts a 1,500-byte MTU, what IP MTU must every underlay link carry, and why?
answer
- the tunnel endpoint will not split it
- count what sits inside outer IP
- the inner Ethernet header is payload now
- the outer Ethernet header is not
basics
~20 sAt least 1,550 bytes: the 1,500-byte tenant packet plus its 14-byte inner Ethernet header, 8 bytes of VXLAN, 8 of UDP and 20 of outer IPv4. RFC 7348 forbids VTEPs to fragment, so a smaller link drops large packets.
solid answer
~50 sEvery underlay link on every path between VTEPs needs an **IP MTU of at least 1,550 bytes**. The encapsulated packet is the tenant's 1,500-byte IP packet plus its 14-byte inner Ethernet header, then 8 bytes of VXLAN header, 8 of UDP and 20 of outer IPv4: `1500 + 14 + 8 + 8 + 20 = 1550`. The outer Ethernet header frames the packet on each link and sits outside the IP MTU, so it does not count. The size is a hard requirement because RFC 7348 says VTEPs **MUST NOT fragment** VXLAN packets and lets the receiving VTEP silently discard fragments, so an undersized link drops large packets rather than slowing them. An IPv6 underlay needs 1,570, and a kept inner 802.1Q tag adds 4. In practice operators set one jumbo MTU well above the minimum; its exact size is an implementation choice.
go deeper
Recall that VXLAN wraps a whole Ethernet frame in UDP and IP, so the network underneath must carry bigger packets than the tenant hosts send.
Do the sum aloud: 1,500 plus the inner Ethernet header, VXLAN, UDP and outer IPv4 gives 1,550, and say why the outer Ethernet header is not counted.
Show that the rule covers every link on every path, that devices count MTU differently, and that one missed link fails only the flows ECMP hashes onto it.
Argue for one fabric-wide jumbo value with headroom for IPv6 outer headers, kept tags and future encapsulations, rather than a per-path minimum the next change breaks.
## Why the underlay has to grow A **VXLAN** overlay (RFC 7348) carries a tenant's whole Ethernet frame inside a UDP datagram between two **VTEPs** (VXLAN Tunnel End Points). The tenant hosts believe they sit on an ordinary Ethernet segment with the usual **1,500-byte MTU** - the maximum transmission unit, the largest IP packet a link carries in one piece. The **underlay**, the routed IP network between the VTEPs, has to carry that same packet *plus* the wrapper around it. Two sentences in RFC 7348 section 4.3 turn the size into a hard requirement rather than a performance tweak: - **VTEPs MUST NOT fragment VXLAN packets.** A VTEP that cannot fit the encapsulated packet onto its uplink has nothing legal left to do but drop it. - Intermediate routers may fragment the encapsulated packet, but **the destination VTEP MAY silently discard** such fragments. The RFC therefore RECOMMENDS that MTUs across the physical network be set to accommodate the larger frame, and adds that Path MTU Discovery MAY also be used. ## Counting the bytes inside the outer IP header An **IP MTU** counts the IP header and everything after it, never the link-layer header in front. For an untagged tenant frame over an IPv4 underlay: | Piece | Bytes | Inside the outer IP MTU? | |---|---|---| | Outer Ethernet header | 14 | No - it frames the outer packet on each link | | Outer IPv4 header | 20 | Yes | | Outer UDP header | 8 | Yes | | VXLAN header | 8 | Yes | | Inner Ethernet header | 14 | Yes - it is payload now | | Tenant IP packet | 1,500 | Yes | | Inner FCS | 0 | Not carried - RFC 7348's frame format excludes it | So the outer IPv4 packet is `20 + 8 + 8 + 14 + 1500 = 1550` bytes, and the frame on the wire is `14 + 1550 = 1564` bytes before the new outer FCS. ## Two different 50s The figure usually quoted as VXLAN's overhead is **50 bytes**: outer Ethernet 14 + IPv4 20 + UDP 8 + VXLAN 8. The required IP MTU also grows by 50, but for a different reason: the **outer** Ethernet header drops out of the count and the **inner** Ethernet header, now payload, drops in. The two 14-byte headers trade places. Keeping them apart matters the moment anything changes: 1. **An outer 802.1Q tag on the underlay** makes the frame overhead 54 bytes, but the IP MTU stays 1,550, because the tag lives in the outer Ethernet header. 2. **A kept inner 802.1Q tag** lengthens the inner frame by 4 bytes, so the IP MTU becomes 1,554. RFC 7348 says a VTEP SHOULD strip the inner tag unless configured otherwise, which is why the untagged case is the usual one. 3. **An IPv6 underlay** replaces the 20-byte IPv4 header with a 40-byte IPv6 header, so the IP MTU becomes 1,570 and the frame overhead 70. ## What "MTU" means on a given device Implementations disagree on whether a configured MTU includes the Ethernet header, so one device's 1,564 can be another's 1,550. Compare like with like before declaring a link correct. **Jumbo-frame** sizes such as 9,000 bytes are implementation choices, not values any RFC sets, and both ends of a link must agree: a link whose ends disagree drops whatever fits one end and not the other. ## Every link, not most links The requirement applies to **every link on every path** between any two VTEPs, at both ends of each link, including any interconnect between data centres. An underlay that spreads traffic with equal-cost multipath sends different flows over different links, so one forgotten link breaks only the flows hashed onto it - the classic intermittent failure. Operators therefore usually choose one jumbo value with headroom for IPv6 outer headers, kept tags or a future encapsulation with options, and apply it fabric-wide rather than computing a bare minimum per path. ## Why not simply let the packet fragment? - **Cost at line rate.** Fragmenting at the sender and reassembling at the far VTEP needs buffers and per-packet state that a hardware forwarding path rarely has; RFC 7348 sidesteps this by forbidding VTEP fragmentation and permitting discard. - **Loss amplification.** Losing any one fragment loses the whole packet. - **Load-balancing side effects.** Fragments after the first carry no UDP header, so an underlay that hashes on UDP ports may place them differently from the first fragment. Lowering the tenant MTU instead is the other way to make room; that trade, and what the tenant host hears when a big packet is dropped, are separate questions.
- If VXLAN's overhead is quoted as 50 bytes including the outer Ethernet header, why does the underlay's IP MTU also grow by exactly 50?Because two different 14-byte headers trade places. The 50-byte overhead counts the outer Ethernet header plus IPv4, UDP and VXLAN. An IP MTU leaves out the outer Ethernet header but includes everything inside the outer IP packet, which now contains the inner Ethernet header. For an untagged inner frame over IPv4 both sums are 50; with an outer tag, a kept inner tag or IPv6 they diverge.
- What changes when the VTEPs talk to each other over IPv6 instead of IPv4?The outer header is 40 bytes, so a 1,500-byte tenant MTU needs a 1,570-byte underlay IP MTU. And no router on the path may fragment: RFC 8200 reserves IPv6 fragmentation to source nodes, so an undersized link drops the packet and returns an ICMPv6 Packet Too Big to the source VTEP - never to the tenant host.
saying these in an interview costs you the question
- A VTEP fragments oversized VXLAN packets, so a small underlay MTU only costs speed.
- The outer Ethernet header counts against the underlay's 1,550-byte IP MTU.
- An outer 802.1Q tag on the underlay raises the required IP MTU to 1,554.
- Only the VTEPs' own uplinks need the larger MTU; the spines can stay at 1,500.
- RFC 7348 requires 9,000-byte jumbo frames on every underlay link.