After a VXLAN overlay goes live, logins and small requests between VMs work but large transfers stall; why doesn't Path MTU Discovery rescue the VMs, and what fixes it?
answer
- small fits, large does not
- who is the outer source?
- the error goes to the wrong host
- fragments the far end may discard
- only some paths are too small
basics
~20 sFull-size packets no longer fit once VXLAN adds 50 bytes, and any ICMP error goes to the VTEP, the outer source, never the VM, so the VM's Path MTU Discovery never learns. Fix the underlay MTU, or lower the tenant MTU.
solid answer
~50 sSmall packets still fit after VXLAN's 50 bytes; full-size ones do not, and where they die decides who hears about it. An **ingress VTEP** whose uplink is too small MUST NOT fragment (RFC 7348), so it drops the packet, and RFC 7348 defines no message back to the tenant. An **underlay router** with a smaller link either drops the packet and sends ICMP **Fragmentation Needed** (or ICMPv6 **Packet Too Big**) to the *outer* source - the VTEP - or, if the outer DF bit is clear, fragments it, and the destination VTEP MAY silently discard the fragments. Either way the VM's Path MTU Discovery never gets a message, so its TCP keeps retransmitting a segment that cannot arrive. Confirm with DF-set probes of growing size, find the undersized link, and fix the MTU: raise the underlay everywhere or lower the tenant MTU. MSS clamping is only a TCP stopgap.
go deeper
Recall that a tunnel makes packets bigger, so large packets can be dropped where small ones pass, and nobody tells the sending host.
Explain why an underlay router's ICMP error goes to the VTEP, the outer source, and so never reaches the VM's Path MTU Discovery.
Walk the diagnosis: DF-set probes from the VM and between VTEPs, every link's MTU at both ends, ECMP explaining partial failure, then fix the MTU.
Treat silent loss as a design defect: one fabric-wide MTU standard, a probe that verifies it after every change, and stopgaps kept as stopgaps.
## The symptom The overlay is up. VMs ping each other, TCP handshakes complete, logins and small API calls succeed - and then a file copy, a database dump or a large HTTP response stalls and times out. Often only some VM pairs or some connections are affected, and a retry sometimes works. This is a **black hole**: packets disappear and the sender is never told why. ## Why small works and large does not **VXLAN** (RFC 7348) wraps each tenant frame in VXLAN, UDP and outer IP headers between two **VTEPs** (VXLAN Tunnel End Points). For a tenant packet of 1,500 bytes over IPv4 the outer packet is 1,550 bytes. Handshakes and small requests are far below any limit; full-size TCP segments carrying bulk data are exactly the packets that no longer fit an underlay link left at 1,500. ## Where the packet dies, and who hears about it Path MTU Discovery works only if the *sender* receives an error when its packet is too big. Inside a tunnel, the error - if there is one - goes to the wrong host: | Where the oversized packet meets a small MTU | What happens | What the tenant VM hears | |---|---|---| | The ingress VTEP's own uplink | RFC 7348: VTEPs MUST NOT fragment, so it is dropped | Nothing RFC 7348 defines | | An underlay router, outer IPv4 DF set | Dropped; ICMP Fragmentation Needed goes to the source VTEP | Nothing - the error is addressed to the VTEP | | An underlay router, outer IPv4 DF clear | Fragmented in transit | Nothing - the destination VTEP MAY silently discard the fragments | | An underlay router, outer IPv6 | Dropped (IPv6 routers never fragment, RFC 8200); Packet Too Big to the source VTEP | Nothing | Whether a VTEP sets DF on the outer IPv4 header is an implementation choice. RFC 7348 also says the VTEPs MAY use Path MTU Discovery between themselves, but that teaches the **VTEP** the tunnel's limit; RFC 7348 gives it no way to pass the lesson to the tenant. The Geneve specification, RFC 8926, goes further and allows a tunnel endpoint that is able to send ICMP to return Fragmentation Needed or Packet Too Big to the tenant system - a useful contrast, not VXLAN behaviour you can count on. The VM's own Path MTU Discovery therefore waits for a message that never arrives, and its TCP retransmits the same full-size segment until the connection gives up - unless the host's stack probes for a working size on its own, as described below. ## Why it looks intermittent - **Equal-cost multipath.** The underlay hashes outer headers, including a UDP source port that varies with each inner flow, so if only one link is undersized, only the flows hashed onto it lose their large packets. A reconnection with new ports may hash elsewhere and succeed. - **Direction.** A request is small and its response large, so one direction works and the other stalls. - **Same-VTEP pairs.** Two VMs behind the same VTEP never use the tunnel, so their traffic is never encapsulated and is unaffected. ## Confirming it 1. From one VM to another, send ICMP echoes with DF set and growing size. A 1,472-byte payload makes a 1,500-byte IP packet (`1472 + 8 ICMP + 20 IP`). If smaller probes pass and 1,472 vanishes without any error, the path is black-holing. 2. Between the two VTEP addresses, send DF-set probes sized to a 1,550-byte IP packet. A failure here points at the underlay, not the tenant. 3. Check the MTU at **both ends of every underlay link** on the paths, remembering that devices differ on whether their MTU setting includes the Ethernet header. 4. Compare with two VMs behind one VTEP; if they pass, the tunnel is the suspect. ## Fixing it - **Raise the underlay MTU** on every link to at least 1,550 for an IPv4 underlay (1,570 for IPv6), usually to a jumbo value with headroom. - **Or lower the tenant MTU** to 1,450 for every host in the segment. - **Stopgaps, not cures:** TCP MSS clamping at a tenant router protects only TCP flows that cross that router, and Packetization Layer Path MTU Discovery (RFC 4821 for TCP, RFC 8899 for datagram transports) lets a host find a working size by probing without ICMP - if its stack has it enabled. - **What does not help:** clearing DF on the tenant's packets. The VTEP bridges the frame and still MUST NOT fragment the VXLAN packet, whatever the inner header says.
- If the VTEPs run Path MTU Discovery between themselves, as RFC 7348 permits, does that fix the VMs' problem?Only half of it. The VTEP learns that the tunnel path carries less than it needs, so it knows which packets will not fit, but RFC 7348 gives it no way to tell the tenant host. Unless the implementation goes beyond the RFC and sends the host an ICMP error, the VM still sees silent loss. The cure is still an MTU that fits.
- Why do some connections between the same two VMs work while others stall?The underlay spreads VXLAN traffic by hashing outer headers, and the outer UDP source port varies with the inner flow. If only one link on one path has a small MTU, only flows hashed onto that link lose their large packets; a reconnection with new ports may hash elsewhere and succeed, which makes the fault look random.
- Why does clearing the DF bit on the VMs' packets not help?The inner DF bit is irrelevant to the tunnel: the VTEP bridges the frame and MUST NOT fragment the VXLAN packet whatever the inner header says. Fragmentation by an underlay router depends on the outer DF bit, set by the VTEP, and even then the destination VTEP MAY drop the fragments.
saying these in an interview costs you the question
- The VM will receive Fragmentation Needed and shrink its packets on its own.
- Small requests succeed, so the overlay's MTU cannot be the problem.
- Clearing the DF bit inside the tenant's packets lets the tunnel fragment them.
- When an underlay router fragments the outer packet, the far VTEP is required to reassemble it.
- An underlay router's ICMP error is addressed to the tenant VM that sent the data.