skip to content

In BGP EVPN for VXLAN, what do route types 2, 3 and 5 each carry, and when does a VTEP originate each one?

level: seniorimportance: must knowfreq 20%

answer

  1. hosts, flood lists, prefixes
  2. MAC/IP Advertisement: MAC, optional IP, VNI
  3. IMET plus the PMSI Tunnel attribute
  4. RFC 9136 decouples prefixes from MACs

basics

~20 s

Type 2 (MAC/IP Advertisement) carries a host's MAC and optional IP behind a VTEP; type 3 (Inclusive Multicast Ethernet Tag) enrols a VTEP for a VNI's BUM traffic; type 5 (IP Prefix, RFC 9136) carries a prefix without a MAC.

solid answer

~50 s

A **type-2** `MAC/IP Advertisement` route is sent when a VTEP learns a local host: it carries the MAC, optionally one IP (one route per IP), the Ethernet segment identifier, and the VNI in the label field that RFC 8365 reuses, with the VTEP's address as next hop; it is withdrawn when the host ages out and re-sent with a `MAC Mobility` sequence number when the host moves. A **type-3** `Inclusive Multicast Ethernet Tag` route is sent once per VNI a VTEP serves, with its own address and a `PMSI Tunnel` attribute saying whether BUM uses ingress replication or a multicast tree; peers build each VNI's flood list from these. A **type-5** `IP Prefix` route (RFC 9136) is sent for prefixes that are not one host: subnets, external routes at a border leaf, prefixes behind an appliance. Types 1 and 4 serve multihoming.

go deeper

for a junior

Recall the three names and jobs: type 2 for hosts, type 3 for joining a VNI's flood list, type 5 for IP prefixes.

for a middle

Explain the fields that make forwarding possible: the VNI in the label field, the VTEP address as next hop, the optional IP in type 2, the PMSI Tunnel attribute in type 3.

for a senior

Show when each is originated and withdrawn, why prefixes belong in type 5 rather than type 2, and where types 1 and 4 enter for multihoming.

for a principal

Connect route types to scale: per-host type-2 state on every VTEP of a VNI versus aggregated type-5 prefixes, and what that means for leaf table sizing.

## The EVPN route format BGP EVPN carries its routes in a multiprotocol BGP address family. Every EVPN route (its NLRI) starts with a **Route Type** octet and a **Length** octet, followed by fields that depend on the type. RFC 7432 defined types **1 to 4**; RFC 9136 added **type 5**. In a VXLAN fabric, RFC 8365 adds two conventions to every type: the BGP **next hop** is the VTEP's IP address, and the MPLS label field is reused as a 24-bit **VNI field**. A BGP encapsulation extended community says the tunnel is VXLAN. ## Type 2: MAC/IP Advertisement | Field | Size | Role | |---|---|---| | Route Distinguisher | 8 octets | keeps routes from different instances distinct | | Ethernet Segment Identifier | 10 octets | 0 for a single-homed host; the segment's ID when multihomed | | Ethernet Tag ID | 4 octets | 0 for VLAN-based service | | MAC Address Length + MAC | 1 + 6 octets | the host's MAC | | IP Address Length + IP | 1 + 0, 4 or 16 octets | optional host IP | | MPLS Label1 | 3 octets | carries the VNI for VXLAN | | MPLS Label2 | 0 or 3 octets | used by integrated routing and bridging | **When it is sent:** as soon as a VTEP learns a local host. RFC 7432 allows a MAC-only route and requires **one route per IP address** bound to that MAC, so a dual-stack host yields two MAC/IP routes. When a binding goes away, its route is withdrawn. When a host moves, the new VTEP advertises it with a **MAC Mobility** extended community whose sequence number is one higher, and the old VTEP withdraws. **What receivers do:** install the MAC (and the IP binding, if present) pointing at the next-hop VTEP and VNI. The IP binding is what lets a VTEP answer ARP locally. ## Type 3: Inclusive Multicast Ethernet Tag (IMET) Fields: Route Distinguisher, Ethernet Tag ID, and the **Originating Router's IP Address**. RFC 7432 says each PE **MUST** advertise one; for VXLAN that means one per VNI a VTEP serves. The route carries a **PMSI Tunnel attribute** naming how BUM traffic is delivered. RFC 8365 lists the tunnel types usable with VXLAN: 3 PIM-SSM tree, 4 PIM-SM tree, 5 BIDIR-PIM tree, and **6 ingress replication**. **When it is sent:** when a VNI is configured on the VTEP, not per host. With ingress replication, every receiver adds the originator to that VNI's replication list, and the source VTEP sends one unicast copy of each broadcast frame to every VTEP on the list. ## Type 5: IP Prefix (RFC 9136) Fields: Route Distinguisher, ESI, Ethernet Tag ID, **IP Prefix Length** (0-32 or 0-128), **IP Prefix**, **GW IP Address**, and a label (the VNI). The NLRI length is **34** octets for IPv4 and **58** for IPv6. **When it is sent:** for anything that is a prefix rather than a host: a tenant subnet, routes a border leaf learned from outside the fabric, or prefixes behind a firewall or load-balancing appliance. RFC 9136 explains why type 2 is the wrong tool. In its floating-IP example, 1,000 prefixes advertised in type-2 routes tied to one MAC must all be withdrawn and re-sent when the floating IP moves to another device; with type 5 the prefixes point at the floating IP as an **overlay index**, and only the single type-2 route for that IP changes. A type-5 route can also name its next hop directly, with the router's MAC in the **EVPN Router's MAC** extended community (RFC 9135). ## Types 1 and 4, for completeness - **Type 1, Ethernet Auto-Discovery**: per Ethernet segment, for mass withdrawal and split-horizon information; per EVPN instance, for aliasing and backup paths. - **Type 4, Ethernet Segment**: lets the VTEPs attached to the same multihomed segment discover each other and elect a designated forwarder. ## A worked trace A host with MAC `02:00:00:aa:00:20` and IP `10.1.1.20` attaches to a leaf whose VTEP address is `10.0.0.11`, in VNI 10100: 1. When VNI 10100 was configured, the leaf sent a type-3 route with originator `10.0.0.11` and a PMSI Tunnel attribute of ingress replication. Every leaf hosting VNI 10100 added `10.0.0.11` to its replication list. 2. The host sends its first frame; the leaf learns the MAC locally and sends a type-2 route with that MAC, VNI 10100 in Label1, next hop `10.0.0.11`. 3. The host's ARP traffic reveals its IP; the leaf sends a second type-2 route carrying MAC and `10.1.1.20`. 4. A border leaf learns `203.0.113.0/24` from outside and sends a type-5 route for it. Every receiving VTEP now has unicast reachability for the host, a flood list for VNI 10100, and a routed path to the external prefix, without any frame having been flooded to discover them.

  • Why not advertise an external prefix as a type-2 route with a shorter IP length?
    A type-2 route binds an IP to a MAC, and RFC 7432 sets its IP length to 32 or 128 for ARP and ND use. RFC 9136 shows the cost of overloading it: prefixes tied to a MAC must all be withdrawn and re-sent when that MAC changes. Type 5 decouples prefixes from MACs, so a single type-2 update handles a floating IP's move.
  • Which VNI and outer destination does a VTEP use toward a host learned from a type-2 route?
    The VNI comes from the route's MPLS Label1 field, which RFC 8365 reuses as a 24-bit VNI field, and the outer destination IP is the route's BGP next hop, which RFC 8365 sets to the advertising VTEP's address. Those two values are all the encapsulation needs.
  • What happens when a VTEP stops serving a VNI?
    It withdraws its type-3 route for that VNI, and peers drop it from the replication list they built from those routes, so it stops receiving that VNI's BUM copies. Its type-2 routes for hosts in that VNI are withdrawn as well.

saying these in an interview costs you the question

  • Type-3 routes carry the MAC addresses of hosts that joined multicast groups.
  • A type-5 route is just a type-2 route with a shorter IP prefix length.
  • A VTEP sends one type-3 route for every host it learns.
  • Route type 5 is defined in RFC 7432 alongside types 1 to 4.
  • Every type-2 route must carry an IP address.