After an underlay change in a flood-and-learn VXLAN fabric, hosts in one VNI resolve ARP only within a rack while established flows keep working; what is broken?
answer
- two paths, not one tunnel
- established flows prove unicast works
- ARP rides the flood path
- group mapping, tree or peer list
basics
~20 sThe BUM path is broken, not the tunnel: learned unicast still flows between VTEPs, but floods no longer cross racks — typically underlay multicast for the VNI's group, a VNI-to-group mismatch, or a missing ingress-replication peer.
solid answer
~40 sSplit the symptoms by path. Established flows use learned entries and **unicast** VXLAN between VTEP addresses, so underlay unicast routing is fine. A new conversation starts with an ARP **broadcast**, which travels the VNI's flood path — and that path now stops at the rack. In multicast mode, check that every VTEP maps the VNI to the same group, that its IGMP joins are seen, and that the underlay still builds the cross-rack tree — multicast routing on every hop and, for PIM Sparse Mode, a reachable rendezvous point. In ingress-replication mode, check each VTEP's peer list; an asymmetric list gives a tell-tale one-way failure. Fix it quickly: established flows fail too once their entries or ARP caches expire.
go deeper
Recall that ARP requests are broadcasts and travel a VXLAN segment's flood path, while traffic to learned MACs travels unicast between VTEPs.
Explain why working established flows rule out unicast routing and encapsulation, and list what the flood path depends on in multicast and in ingress-replication mode.
Diagnose methodically: group mapping, joins, multicast routing and the rendezvous point, or asymmetric peer lists, and explain why the fault spreads as entries age out.
Use the incident to weigh a design whose first contact depends on a separate, rarely exercised path, and what monitoring of that path would have caught it.
## Read the symptoms as two paths A flood-and-learn VXLAN fabric moves traffic between **VTEPs** (VXLAN Tunnel End Points) on two different paths, and this fault breaks only one of them. | Path | Carries | Underlay mechanism | Status here | |---|---|---|---| | Known unicast | Frames to MACs already learned | Unicast VXLAN to the remote VTEP's IP | Working: established flows survive | | BUM flood | ARP requests, unknown unicast, multicast | The VNI's multicast group, or per-peer unicast copies | Broken between racks | Established flows prove that the VTEPs can reach each other, that the encapsulation is accepted and that the VNI is configured at both ends. ARP requests are **broadcasts**, so they need the flood path, and RFC 7348 learning depends on them: a remote VTEP learns a new host only when that host's flooded request reaches it. "Same rack works, across racks fails" says the flood is replicated locally but never crosses the spine. ## Multicast-mode checks RFC 7348 maps each VNI to an IP multicast group through the management layer; VTEPs send IGMP membership reports and use `(*,G)` joins, and a multicast routing protocol such as PIM Sparse Mode builds the tree. Work through it in order: 1. **Group mapping.** Do all VTEPs map this VNI to the **same** group? A VTEP flooding to a group its peers never joined is invisible to them, and a mismatch on one VNI breaks that VNI only. 2. **Joins.** Does each leaf see IGMP membership for the group from its VTEPs? Without them the router does not know to forward the group toward that rack. 3. **Multicast routing across the spine.** Is multicast routing enabled on every underlay hop the change touched? A leaf can still replicate to receivers on its own ports, which matches "within a rack works". 4. **Rendezvous point.** For PIM Sparse Mode shared trees, can every router reach the rendezvous point for that group? If not, cross-rack trees never form. If **all** VNIs fail across racks, suspect steps 3 and 4; if only **one** does, suspect step 1. ## Ingress-replication checks With ingress replication the underlay sees only unicast, so the multicast checks do not apply. In flood-and-learn, each VTEP's peer list for the VNI is static configuration, and floods go only to the VTEPs on it. An **asymmetric** list produces a one-way failure. Suppose VTEP-A's list lacks VTEP-B, but VTEP-B's list has VTEP-A: - A host behind B broadcasts an ARP request; B copies it to A, and A learns that host's MAC behind B. - The reply from a host behind A is unicast, and A now knows where to send it — so resolution **works when B's side starts**. - A request from a host behind A is flooded to A's list, which omits B — so resolution **fails when A's side starts**. Learning from B's floods fills A's unicast table; it never adds B to A's flood list. ## What it is not - **Not underlay unicast routing.** Established cross-rack flows run over it. - **Not MTU.** An MTU fault drops large encapsulated packets while ARP and small packets pass; here ARP is the casualty. RFC 7348 forbids VTEPs from fragmenting VXLAN packets, which makes size faults look different. - **Not the VLAN-to-VNI mapping at the edge.** A host mapped into the wrong VNI fails within its rack too, and its established flows would not have worked. - **Not ARP itself.** Hosts are sending requests; they are just not arriving. ## How to confirm it - **Look for the flood on a remote rack.** Capture on a VTEP's underlay interface in another rack while a host resolves ARP: in multicast mode, no VXLAN packets to the VNI's group arrive; with ingress replication, no copy arrives from the requester's VTEP. - **Check the learned table.** The remote VTEP holds no entry for the requesting host's MAC in that VNI, although it still holds entries learned before the change. ## Why it gets worse Established flows survive on soft state. VTEPs age out entries that see no traffic, and hosts age out ARP cache entries; each expiry turns a working flow into one that needs the broken flood path. A fabric in this state degrades steadily, which is why the fault is urgent even while dashboards for existing traffic look healthy.
- Why do established flows survive while new conversations fail?Established flows use mappings the VTEPs learned before the change, so they travel as unicast VXLAN between VTEP addresses and need no flood. New conversations start with an ARP broadcast, which needs the broken flood path. Once a VTEP entry or a host's ARP cache entry expires, an established flow needs the flood path too and fails.
- How would an underlay MTU fault in the same fabric look different?ARP and small packets would get through and new conversations would start, but large packets would vanish, because the encapsulation adds overhead and RFC 7348 forbids VTEPs from fragmenting VXLAN packets. The tell is size-dependent loss rather than a failure of first contact.
saying these in an interview costs you the question
- If established flows work, the overlay is healthy and the hosts are at fault.
- ARP failing across racks points to an MTU mismatch on the underlay.
- A VTEP adds peers to its ingress-replication list when it learns from their floods.
- New conversations failing means underlay unicast routing between racks is down.
- Learned VTEP entries never expire, so established flows are safe indefinitely.