skip to content

EVPN Control Plane

BGP EVPN advertises MAC and IP reachability as routes instead of flooding, adding ARP suppression and anycast gateways. Interviewers treat it as the standard answer to scaling VXLAN.

on this pageshow

questions

6

Why do VXLAN fabrics adopt BGP EVPN as a control plane, and what does it advertise that flood-and-learn must discover by flooding?

level: middleimportance: must knowfreq 30%

answer

  1. control plane versus data plane
  2. who tells a VTEP where a MAC lives
  3. MAC/IP Advertisement over multiprotocol BGP
  4. RFC 7432 applied to VXLAN by RFC 8365

basics

~20 s

BGP EVPN lets each VTEP announce the MAC and IP addresses it learned locally as BGP routes, so remote VTEPs know where hosts live before traffic flows, instead of learning them by flooding unknown and ARP traffic across the fabric.

solid answer

~50 s

Plain VXLAN (RFC 7348, an Informational RFC) learns which VTEP a remote MAC sits behind from the outer source address of the packets it decapsulates, so unknown unicast, broadcast and ARP have to reach every VTEP in the VNI. BGP EVPN (RFC 7432, applied to VXLAN by RFC 8365) moves *remote* learning into the control plane: a VTEP still learns its own hosts locally, then advertises each MAC, optionally with an IP, in a `MAC/IP Advertisement` route whose next hop is the VTEP's address and whose label field carries the VNI. Each VTEP also sends an `Inclusive Multicast Ethernet Tag` route per VNI, so peers learn who needs broadcast and multicast copies. On top of that EVPN adds what flooding never gave: proxy ARP from learned bindings, MAC-move sequence numbers, all-active multihoming and mass withdrawal. The price is a BGP control plane on every leaf.

go deeper

for a junior

Recall that EVPN is a BGP-based control plane: VTEPs announce where MAC and IP addresses live instead of discovering them by flooding traffic across the fabric.

for a middle

Explain local versus remote learning: local stays in the data plane, remote arrives as MAC/IP routes carrying the VNI and the VTEP next hop, plus a per-VNI inclusive multicast route for BUM.

for a senior

Show judgement about the flooding that remains, such as silent hosts and unknown unicast as an administrative choice, and about the route scale every leaf now carries.

for a principal

Frame EVPN as trading a simple data plane for a distributed BGP control plane: multihoming and faster convergence in return for routing-protocol operations and failure modes on every leaf.

## The problem EVPN solves A **VXLAN** fabric carries Ethernet frames between **VTEPs** (VXLAN tunnel endpoints) inside UDP packets across a routed underlay. Each layer-2 segment is identified by a 24-bit **VNI**. The original specification, **RFC 7348** (Informational, 2014), describes one way for a VTEP to find out which remote VTEP a host sits behind: **data-plane learning**. When a VTEP decapsulates a packet, it records "inner source MAC lives behind outer source IP". Until that has happened, traffic for an unknown MAC, every broadcast and every ARP request must be delivered to all VTEPs in the VNI. RFC 7348 itself notes that other schemes are possible, such as a directory the VTEPs query or one that pushes mappings to them. That flooding is what limits a large fabric: ARP storms grow with the number of hosts, the first packet to a new host is flooded, and a moved host is reachable only after traffic teaches every VTEP again. **BGP EVPN** (Ethernet VPN) is the standard answer. **RFC 7432** defined it for MPLS networks; **RFC 8365** applies it to network virtualization overlays, VXLAN among them. Both are Standards Track. ## What a VTEP advertises EVPN runs as a multiprotocol BGP address family whose routes are **typed**. A VXLAN fabric uses: | Route type | Name | What it tells peers | |---|---|---| | 2 | MAC/IP Advertisement | this MAC (and optionally this IP) is behind me, in this VNI | | 3 | Inclusive Multicast Ethernet Tag | send me this VNI's broadcast, unknown-unicast and multicast traffic | | 5 | IP Prefix (RFC 9136) | this IP prefix is reachable through me, with no MAC attached | | 1 and 4 | Ethernet A-D and Ethernet Segment | multihoming: who shares a server link bundle, and who forwards BUM to it | For VXLAN, RFC 8365 reuses the route's MPLS label field to carry the **24-bit VNI**, and sets the BGP next hop to the **VTEP's IP address**. A remote VTEP that installs a type-2 route therefore knows everything it needs to encapsulate: the outer destination (the next hop) and the VNI. ## What changes and what does not 1. **Local learning stays.** RFC 7432 requires a PE (here, a VTEP) to support ordinary data-plane learning on its own access ports, from frames such as DHCP or ARP requests; it may also learn from the control or management plane. 2. **Remote learning moves to BGP.** RFC 7432 requires remote MACs to be learned in the control plane: each VTEP advertises what it learned locally to every other VTEP in that EVPN instance. 3. **Broadcast and multicast still need delivery.** The type-3 routes build each VNI's list of interested VTEPs, using ingress replication or a multicast tree named in the route's PMSI Tunnel attribute. 4. **Unknown unicast flooding becomes optional.** RFC 7432 says its procedures do not require it when MACs are learned through the control plane, but a VTEP may still have to flood a frame for a MAC it has no route for, and whether to do so is an administrative choice. Silent hosts are the usual reason to leave it on. ## What EVPN adds beyond learning - **Proxy ARP / ARP suppression**: a VTEP holding a type-2 route with an IP can answer a local ARP request itself (RFC 7432 section 10, detailed in RFC 9161). - **MAC mobility**: a sequence number on type-2 routes tells every VTEP which advertisement of a moved MAC is current. - **All-active multihoming**: a server dual-attached to two leaves through an Ethernet segment, with designated-forwarder election. - **Mass withdrawal**: one withdrawn route moves every MAC on a failed link at once. - **Constrained distribution**: route targets keep a VNI's routes on the VTEPs that host it. ## What it costs - **A routing protocol on every leaf.** BGP sessions, policies and their failure modes now sit in the forwarding path of layer 2. - **Control-plane state.** Each VTEP holds a route for every host in the VNIs it serves, so table sizes must be planned rather than discovered. - **Interoperability work.** Route types, the encapsulation extended community and route-target derivation must agree across implementations. ## Standards status RFC 7348 (VXLAN) is Informational; RFC 7432, RFC 8365, RFC 9135 (integrated routing and bridging), RFC 9136 (IP prefix routes) and RFC 9161 (proxy ARP/ND) are Standards Track. RFC 8365 also covers NVGRE (RFC 7637), and says it applies to Geneve (RFC 8926) once incremental work in a separate document is done.

  • Does a VTEP running BGP EVPN stop learning MAC addresses in the data plane?
    No. RFC 7432 requires support for local data-plane learning on access ports: the VTEP sees a host's frames and learns the source MAC as any bridge does, or learns it from ARP, DHCP or management integration. What changes is remote learning: addresses behind other VTEPs arrive as MAC/IP Advertisement routes instead of being gleaned from decapsulated traffic.
  • Is unknown-unicast flooding gone once EVPN is running?
    Not necessarily. RFC 7432 says flooding unknown unicast is not required when MACs are learned through the control plane, but a frame for a MAC with no route may still have to be flooded, and doing so is an administrative choice. Silent hosts are the usual reason to keep it. Broadcast and multicast still use the per-VNI list that type-3 routes build.
  • Why is the IP address in a MAC/IP Advertisement route optional?
    A VTEP may know a host's MAC from a data frame before it knows the host's IP. RFC 7432 allows a MAC-only route, and one MAC/IP route per IP address once a binding is learned, so a dual-stack host produces two. The MAC alone is enough for bridging; the IP is what enables proxy ARP and host routing.

Flood-and-learn is a receptionist who shouts a name through every floor and remembers which floor answered; EVPN is a building directory that each desk updates the moment someone sits down, so most lookups need no shouting, though a visitor nobody registered can still trigger one.

saying these in an interview costs you the question

  • EVPN removes MAC learning entirely, even on a VTEP's own access ports.
  • With EVPN no broadcast or ARP frame ever crosses the fabric again.
  • EVPN is a new tunnel encapsulation that replaces the VXLAN header.
  • EVPN cannot work without IP multicast running in the underlay.
  • RFC 7348 defines the BGP EVPN control plane for VXLAN.
open as a page

In BGP EVPN for VXLAN, what do route types 2, 3 and 5 each carry, and when does a VTEP originate each one?

level: seniorimportance: must knowfreq 20%

basics

~20 s

Type 2 (MAC/IP Advertisement) carries a host's MAC and optional IP behind a VTEP; type 3 (Inclusive Multicast Ethernet Tag) enrols a VTEP for a VNI's BUM traffic; type 5 (IP Prefix, RFC 9136) carries a prefix without a MAC.

open as a page

In a BGP EVPN VXLAN fabric, why does every leaf carry the same gateway IP and MAC, and how does EVPN track a host that moves between leaves?

level: seniorimportance: should knowfreq 18%

basics

~20 s

A distributed anycast gateway puts one gateway IP and MAC on every leaf, so hosts are routed at their own leaf and keep a valid gateway ARP entry after moving; EVPN tracks moves with MAC Mobility sequence numbers on type-2 routes.

open as a page

How does ARP suppression on a BGP EVPN VXLAN VTEP answer ARP requests locally, and which requests still get flooded across the fabric?

level: seniorimportance: should knowfreq 16%

basics

~20 s

The VTEP keeps a proxy table of IP-to-MAC bindings, learned by snooping local ARP and from remote type-2 MAC/IP routes, and answers a local ARP request itself on a hit; a request it cannot answer is still flooded to remote VTEPs.

open as a page

How does BGP EVPN multihome a server to two VXLAN leaves through an Ethernet segment, and what do route types 1 and 4 do?

level: seniorimportance: should knowfreq 12%

basics

~20 s

Both leaves give the server's link bundle one Ethernet Segment Identifier; type-4 Ethernet Segment routes let them discover each other and elect a designated forwarder for BUM, and type-1 Ethernet A-D routes provide aliasing and mass withdrawal.

open as a page

You must replace flood-and-learn with BGP EVPN across a 40-rack VXLAN fabric; how would you sequence the move, and what do you trade?

level: principalimportance: nice to knowfreq 8%

basics

~20 s

Bring EVPN up on every leaf first, move VNIs over one at a time so type-3 routes take over BUM membership, then add ARP suppression, anycast gateways and multihoming; you trade flooding for BGP state and control-plane operations on every leaf.

open as a page