skip to content

Why does VXLAN flood-and-learn scale poorly in a large data-centre fabric, and which costs grow as VTEPs, VNIs and hosts are added?

level: seniorimportance: must knowfreq 24%

answer

  1. traffic doubles as signalling
  2. every VTEP pays for every flood
  3. learned only after traffic
  4. moves, silence and ageing
  5. multicast state or N minus one

basics

~20 s

Every VTEP in a VNI carries every BUM frame; ARP and unknown unicast flood because nothing knows mappings in advance; the underlay holds multicast state or sources send N−1 copies; and a moved host is relearned only when it sends.

solid answer

~50 s

Flood-and-learn turns traffic into signalling, so its costs grow with the fabric. **Every** VTEP hosting a VNI receives **every** broadcast, unknown-unicast and tenant multicast frame in it; ARP makes up much of that, and no VTEP can answer an ARP locally because none knows the bindings. Mappings exist only where traffic has passed, so silent hosts and aged-out entries cause unknown-unicast floods, and a host that moves is relearned only when it next sends — until then, or until entries age out, remote VTEPs send its traffic to the old VTEP, which discards it. Delivery costs either underlay multicast state, with groups shared across VNIs once VNIs outnumber them, or N−1 copies per flood with peer lists kept by hand on every VTEP. RFC 7348 also warns that rogue senders could hijack MAC addresses. A control plane such as BGP EVPN replaces the learning but keeps the encapsulation.

go deeper

for a junior

Recall that every VTEP in a segment receives every broadcast and that remote MACs are learned only from traffic, so floods grow with the fabric.

for a middle

Explain each cost separately: ARP and unknown-unicast flooding, multicast state or N minus one copies, and why silent or moved hosts stay unknown.

for a senior

Demonstrate production judgment: black holes after a move, one-way failures from peer lists, shared groups delivering unwanted floods, and why bigger tables fix none of it.

for a principal

Argue the migration case: which costs a control plane removes, which it keeps, and when a fabric is small enough that flood-and-learn's simplicity still wins.

## Where the costs come from In **flood-and-learn** — the data-plane learning model RFC 7348 describes for VXLAN — a **VTEP** (VXLAN Tunnel End Point) discovers which remote VTEP a MAC address lives behind only by decapsulating traffic from it, and it floods every frame it cannot place. Flooding *is* the discovery protocol. A protocol that runs on user traffic scales with user traffic, VTEP count and segment count all at once, and that is the root of every limit below. RFC 7365, the IETF's data-centre virtualization framework, says it directly: multicasting in the underlay for dynamic learning "may lead to significant scalability limitations". ## Cost by dimension | What grows | What it costs | Why | |---|---|---| | Hosts per VNI | More ARP and unknown-unicast floods | Each host's broadcasts and silences reach every VTEP in the segment | | VTEPs per VNI | More copies per flood | Multicast tree fan-out, or N − 1 copies at the source with ingress replication | | VNIs | More underlay multicast state, or shared groups | One group per VNI cannot last; shared groups deliver unwanted floods | | Fabric size | Bigger peer lists to configure | Without a control plane, ingress-replication lists are static on every VTEP | | Host churn | Stale entries, black holes | Moves and new hosts are learned only from traffic | ## The flooding itself - **Every VTEP pays for every flood.** A broadcast in a VNI spanning 200 VTEPs is decapsulated and processed 199 times, whether or not the target is behind any given VTEP. - **ARP is the bulk.** Each host resolves its peers by broadcast, and nobody in the overlay can answer on its behalf, because no VTEP knows IP-to-MAC bindings in advance. RFC 9161 describes the same pattern on large layer 2 peering networks, where roughly all BUM traffic comes from ARP and Neighbor Discovery and reaches every router. - **Unknown unicast keeps returning.** A host that only receives, or whose entry aged out, is unknown at remote VTEPs, so traffic toward it floods until it sends something. - **Tenant multicast is flooded too.** RFC 7348 sends a segment's multicast frames down the same tree as its broadcasts, so every VTEP in the VNI receives them. ## Mobility is learned only from traffic When a virtual machine migrates from VTEP1 to VTEP4: 1. Remote VTEPs still map its MAC to VTEP1 and keep sending unicast there. 2. VTEP1 no longer has the host; RFC 7348 delivers a decapsulated frame only to a local host that owns the destination MAC, so VTEP1 discards it. 3. Nothing corrects the remote VTEPs until the host sends a frame that reaches them, or their entries age out. 4. A **broadcast** from the moved host — commonly an ARP Announcement (a gratuitous ARP) sent after migration — floods to every VTEP and relearns them all at once; a unicast frame relearns only its destination's VTEP. So convergence after a move depends on the host's behaviour and on ageing timers that are an implementation choice, not on anything the overlay guarantees. ## Underlay state and configuration A 24-bit VNI allows 16,777,216 segments, but an underlay holds far fewer multicast groups. Once VNIs outnumber groups, several VNIs share one, and every VTEP on that group receives — and discards — floods for segments it does not host. Moving to ingress replication removes the multicast state but multiplies source bandwidth by N − 1 and turns membership into a list that must be kept consistent on every VTEP; a missing entry can produce one-way failures. ## Security exposure RFC 7348 §7 warns that a MAC-over-IP scheme extends the layer 2 attack surface: rogues can subscribe to the multicast groups that carry a segment's broadcasts, and can source MAC-over-UDP frames into the transport network to inject traffic, "possibly to hijack MAC addresses". A learning VTEP believes whatever inner source MAC arrives from whatever outer source IP. RFC 7348 builds in no protection and points to IPsec, ACLs on the UDP 5-tuple, admission control and a dedicated VLAN for VXLAN traffic. ## What replaces it The fix is to stop learning from the data plane. RFC 8365 contrasts the models: data-plane learning requires flooding unknown unicast and ARP frames, while control-plane learning does not. A control plane such as BGP EVPN advertises MAC-to-VTEP mappings, membership and moves before traffic needs them — its route types and ARP suppression are a separate subject — and it keeps the same VXLAN encapsulation on the wire.

  • After a VM migrates between hosts in a flood-and-learn VXLAN segment, what shortens the black hole?
    Remote VTEPs keep pointing at the old VTEP, which discards the frames, until the VM sends something that reaches them or their entries age out. A broadcast from the moved host, typically an ARP Announcement sent after migration, floods to every VTEP and relearns them all at once. Shorter ageing helps at the cost of more floods; a control plane removes the dependence on traffic.
  • Would larger MAC tables on the VTEPs fix flood-and-learn's flooding?
    No. Table size is not why it floods: broadcasts always flood, and unknown unicast happens because a mapping was never learned or aged out, not because the table was full. Larger tables only help when overflow is evicting entries; they do nothing for silent hosts, moves or ARP volume.
  • What security weakness of data-plane learning does RFC 7348 point out?
    A VTEP learns whatever inner source MAC arrives against whatever outer source IP sent it. RFC 7348 §7 warns that rogues can join the multicast groups carrying a segment's broadcasts and inject MAC-over-UDP frames, possibly hijacking MAC addresses. It offers no built-in defence and points to IPsec, ACLs on the UDP 5-tuple, admission control and a dedicated VLAN for VXLAN traffic.

saying these in an interview costs you the question

  • Flood-and-learn's only scaling problem is the size of the VTEP MAC table.
  • Once a VTEP learns a host, it never floods traffic for that host again.
  • Ingress replication removes the flooding, so it fixes flood-and-learn's scaling.
  • When a VM moves, the old VTEP forwards its traffic on to the new VTEP.
  • With 24-bit VNIs, a multicast underlay gives every segment its own group for free.
  • Adding a control plane means changing the VXLAN header on the wire.