For a VXLAN overlay, what do you trade when you raise the underlay MTU instead of lowering every tenant's MTU to 1,450 bytes?
answer
- make the pipe bigger or the load smaller
- who has to change what
- 1,500 minus the wrapper
- one segment, one MTU
basics
~20 sRaising the underlay MTU leaves tenants at the standard 1,500 bytes but needs every underlay link on every path changed. Lowering tenants to 1,450 fits any 1,500-byte underlay but needs every tenant host changed, and a missed host fails silently.
solid answer
~50 sBoth make room for VXLAN's 50 bytes, which is unavoidable because RFC 7348 forbids VTEPs to fragment. **Raising the underlay MTU** - to at least 1,550 over IPv4, usually a jumbo value - is invisible to tenants: hosts, appliances and images keep 1,500. The cost lands on the network team: every underlay interface on every path, including interconnects you may not own, must be raised and kept that way. **Lowering the tenant MTU** to `1500 - 50 = 1450` works over any standard underlay, including someone else's, but every tenant endpoint must be told, by DHCP or the virtual NIC, and any host you cannot configure becomes a silent failure. Bulk TCP also sends about 3.5% more packets (MSS 1,410 instead of 1,460). Where you own every link, raise the underlay; lowering is the fallback when the underlay is not yours.
go deeper
Recall the two ways to make room for VXLAN's extra bytes: let the underlay carry more, or have the tenants send less.
Explain why 1,500 minus 50 gives 1,450, what each option asks of whom, and why mixed MTUs inside one bridged segment fail silently.
Bring the operational view: missed underlay links, hosts you cannot configure, interconnects you do not own, and MSS clamping as a TCP-only stopgap.
Frame it as ownership: raise where you own every link, lower only where the underlay is someone else's, and never let one segment mix values.
## The constraint both options answer A **VXLAN** overlay (RFC 7348) wraps each tenant Ethernet frame in VXLAN, UDP and outer IP headers between two **VTEPs** (VXLAN Tunnel End Points). For an untagged frame over IPv4 that wrapper adds **50 bytes** inside the outer IP packet, and RFC 7348 says **VTEPs MUST NOT fragment** VXLAN packets. So a full 1,500-byte tenant packet cannot cross a 1,500-byte underlay. There are only two ways to close the gap: make the pipe bigger, or make what goes into it smaller. ## Option 1: raise the underlay MTU - **Transparent to tenants.** Hosts, virtual machines, appliances and images keep the 1,500 bytes RFC 894 gives IP over Ethernet; nobody outside the network team has to know the overlay exists. - **Every link, every path.** The larger MTU (at least 1,550 over IPv4, 1,570 over IPv6) must be set at both ends of every underlay link that any VTEP pair might use. One forgotten link fails only the flows that equal-cost multipath hashes onto it, which is the hardest kind of fault to see. - **Headroom is cheap.** Operators usually pick one jumbo value - its size is an implementation choice, not an RFC value - that also covers an IPv6 outer header, a kept inner 802.1Q tag, or an encapsulation such as Geneve (RFC 8926), whose endpoints must assume the maximum options length when sizing a tunnel. - **The limit is ownership.** A wide-area link or a provider's network between sites may be fixed at 1,500, and no setting of yours can change that. ## Option 2: lower the tenant MTU The arithmetic is the overhead subtracted from the underlay MTU: 1. IPv4 underlay at 1,500: tenant IP MTU `1500 - 50 = 1450`. 2. IPv6 underlay at 1,500: tenant IP MTU `1500 - 70 = 1430`. 3. An inner 802.1Q tag kept on the frame costs 4 bytes more in either case. What it costs: - **Every endpoint must agree.** The smaller MTU has to reach every host in the segment, through DHCP, the hypervisor's virtual NIC settings or the image. A bridged segment has no device that fragments or signals between hosts, so any packet above 1,450 bytes that a host left at 1,500 sends is lost. - **Hosts you do not control.** Appliances, bare-metal servers and images with a hard-coded MTU are where this option usually breaks. - **A little efficiency.** A TCP segment carries 1,410 bytes of data instead of 1,460 (`1450 - 40`), so bulk transfers need about `1460 / 1410 = 1.035` times as many packets - roughly 3.5% more. That is a cost, not a crisis. - **Edges are routers.** Traffic routed into the overlay from a 1,500-byte network crosses a router, the segment's gateway. Like any router whose next link is smaller, it can fragment an IPv4 packet whose DF bit is clear or return Fragmentation Needed when DF is set. ## Side by side | Question | Raise the underlay | Lower the tenants | |---|---|---| | Who must change | Every underlay link | Every tenant endpoint | | Works across an underlay you do not own | No | Yes | | Failure when one place is missed | Flows hashed onto that link lose large packets | That host loses its large packets | | Throughput cost | None | About 3.5% more packets for bulk TCP | | Room for IPv6 outer headers or options | Yes, if the jumbo value has headroom | Only by lowering further | ## How the choice is usually made 1. If you own every underlay link, raise it once, fabric-wide, with headroom. 2. If the overlay crosses a network you cannot change, lower the tenant MTU for the segments that cross it, or terminate the overlay before that network. 3. Never let one bridged segment mix MTUs; pick one value per segment and enforce it. TCP MSS clamping at a tenant router is sometimes offered as a third way. It rewrites the maximum segment size in TCP handshakes passing that router, so it protects only TCP flows that cross it; treat it as a stopgap while the MTU is fixed, not as an answer.
- Why can one host left at 1,500 bytes in a 1,450-byte VXLAN segment fail while its neighbours work?Nothing in a bridged segment fragments or signals: the VTEP must not fragment and, as a bridge, sends the host no error. Its large UDP datagrams vanish, and its TCP sends 1,460-byte segments whenever the peer advertises a 1,460-byte MSS - typically a peer outside the segment - so those transfers stall. TCP to an in-segment peer at 1,450 survives, because that peer advertises 1,410.
- Does TCP MSS clamping at the tenant's router remove the need to choose?No. Clamping rewrites the MSS in TCP handshakes that pass through a router doing it, so it protects only TCP flows that cross that router. UDP, ICMP, tunnels inside the tenant and bridged traffic that never crosses the clamping point can still send full-size packets. It is a stopgap while the MTU is fixed, not a substitute.
saying these in an interview costs you the question
- Lowering the tenant MTU needs no change on the tenant hosts themselves.
- Raising the MTU on the VTEP-facing ports alone is enough; transit links can stay at 1,500.
- Hosts in one VXLAN segment may use different MTUs because the VTEP fragments between them.
- TCP MSS clamping fixes large packets for every protocol, so the MTU choice stops mattering.
- A 1,450-byte tenant MTU costs so much throughput that it is never acceptable.