In a BGP EVPN VXLAN fabric, why does every leaf carry the same gateway IP and MAC, and how does EVPN track a host that moves between leaves?
answer
- route at the first hop
- one gateway address on every leaf
- no gateway re-ARP after a move
- MAC Mobility sequence number
basics
~20 sA distributed anycast gateway puts one gateway IP and MAC on every leaf, so hosts are routed at their own leaf and keep a valid gateway ARP entry after moving; EVPN tracks moves with MAC Mobility sequence numbers on type-2 routes.
solid answer
~50 sRFC 9135 gives two options for the gateway on each leaf's routing interface: every leaf uses the **same anycast IP and MAC**, or the same IP with **each leaf's own MAC**, aliased through the `Default Gateway` extended community on type-2 routes. It recommends the first: nothing extra to advertise, and a host that moves keeps a gateway ARP entry that is still correct. Every leaf answers the gateway's ARP and routes locally, so no traffic hairpins to one elected router as with a virtual-router redundancy protocol. When a host moves, the new leaf advertises its type-2 route with a `MAC Mobility` extended community whose sequence number is one higher; every VTEP prefers the higher number and the old leaf withdraws its route. Repeated moves (by default 5 within 180 seconds, per RFC 7432) flag a duplicate MAC.
go deeper
Recall the idea: every leaf is the default gateway, using the same IP and MAC, so hosts route at their own rack.
Explain how a host's ARP for the gateway is answered locally and why sharing the MAC as well as the IP keeps the host's ARP entry valid after a move.
Walk a move through MAC Mobility sequence numbers, the lowest-IP tie-break and the old leaf's withdrawal, and explain duplicate-MAC detection.
Contrast first-hop routing everywhere with a centralised active gateway: no hairpinning and no election, in return for routing state on every leaf and harder troubleshooting.
## The gateway problem in a stretched subnet A VXLAN segment can span many racks, so the hosts of one IP subnet sit behind many leaves. Every host needs a **default gateway** to reach other subnets. If that gateway lives on one or two devices, as with a virtual-router redundancy protocol such as VRRP (RFC 9568), routed traffic from every rack crosses the fabric to the one **Active** router and back, and that router's tables and links carry the whole subnet's routed load. **Integrated routing and bridging** (IRB, **RFC 9135**) instead places a routing interface for the subnet on every leaf that serves it. The leaf bridges traffic within the subnet and routes traffic leaving it, right at the first hop. The question is what address that interface presents to hosts. ## Two ways to address the gateway | | Option 1: anycast IP and MAC | Option 2: anycast IP, per-leaf MAC | |---|---|---| | Gateway IP | same on every leaf | same on every leaf | | Gateway MAC | same on every leaf | each leaf's own | | Extra signalling | none | Default Gateway extended community on type-2 routes, so leaves alias each other's MACs | | After a host moves | its cached gateway ARP entry still matches | a new leaf must still accept the old MAC, or the host re-ARPs | | RFC 9135 | **recommended** | allowed | RFC 9135 recommends option 1 because provisioning is simple, the Default Gateway extended community and its MAC-aliasing procedure are unnecessary, and "following host mobility, the host does not need to refresh the default GW ARP/ND entry". It also says option 1 should be used if hosts rely on any form of MAC security, so a host receives traffic from the same MAC it sends to. An implementation may auto-derive the anycast MAC from the virtual-router MAC range VRRP defines (`00-00-5E-00-01-{VRID}` for IPv4), but only where it cannot collide with a host's MAC. Leaves can additionally carry a **unique, non-anycast address** on the same interface for ping and traceroute, since the anycast one cannot identify a single leaf. ## What every leaf does 1. A host ARPs for the gateway IP. 2. Its own leaf answers with the anycast MAC; no other leaf hears the request. 3. Frames to the anycast MAC are routed on that leaf, then sent toward the destination's VTEP. 4. Return traffic routed toward the host carries the anycast MAC as its source. How the routed packet crosses the fabric (symmetric or asymmetric IRB, and the VNIs involved) is a separate design choice; the anycast gateway works with either. ## A host moves **MAC mobility** (RFC 7432 section 15) keeps every VTEP agreeing on where a MAC lives: 1. The host was advertised by leaf A in a type-2 route, with no MAC Mobility extended community (treated as sequence 0). 2. The host moves behind leaf B and sends a gratuitous ARP (or just data). 3. Leaf B already holds leaf A's route for that MAC, so it recognises a move and advertises its own type-2 route with a **MAC Mobility** extended community, **sequence number 1**. 4. Every VTEP prefers the route with the higher sequence number. If two advertisements carry equal sequence numbers but different segments, the one from the VTEP with the **lowest IP address** wins. 5. Leaf A sees a higher number for its own MAC and **withdraws** its route; RFC 9135 adds that it probes locally with ARP to confirm the host is gone. Because every leaf presents the same gateway MAC, the host's routed traffic simply continues from leaf B. ## Duplicates and sticky MACs - **Duplicate MAC detection:** a VTEP that sees **N** moves of one MAC within **M** seconds declares a duplicate, alerts the operator and stops sending and processing MAC/IP routes for that MAC until it is fixed. RFC 7432's defaults are **M = 180** and **N = 5**, and both must be configurable. - **Sticky MACs:** a MAC advertised with the static flag in its MAC Mobility extended community is not allowed to move; a VTEP that learns it locally must alert the operator. ## What it costs - Every leaf serving a subnet needs that tenant's routing state, not just its bridging state. - Troubleshooting needs the unique per-leaf addresses, because "the gateway" is everywhere. - MAC moves are cheap but not free: each one is a BGP update to every VTEP in the segment.
- Why not run a virtual-router redundancy protocol across all the leaves instead?A protocol such as VRRP elects one Active router per group, and only that router forwards for the virtual address. Routed traffic from every rack would cross the fabric to it and back, concentrating load on one device. An anycast gateway lets every leaf answer the gateway's ARP and route locally, with no election at all.
- Two VTEPs keep advertising the same MAC with rising sequence numbers; what does EVPN do?RFC 7432 treats it as a probable duplicate MAC: a VTEP that sees N moves within M seconds (defaults 5 and 180, both configurable) alerts the operator and stops sending and processing MAC/IP routes for that MAC until someone corrects it. That stops the sequence number climbing forever.
- How do you ping one particular leaf's gateway interface when they all share an address?RFC 9135 allows each routing interface to carry a unique, non-anycast IP alongside the anycast one for OAM such as ping and traceroute, because the shared anycast address cannot identify a single leaf. Those unique addresses must then be advertised so the rest of the fabric can reach them.
saying these in an interview costs you the question
- The anycast gateway uses VRRP to choose which leaf answers ARP.
- After moving to another leaf, a host must re-ARP for a new gateway MAC.
- A moved MAC goes to whichever type-2 route arrived most recently.
- Each leaf needs a different gateway IP for the same subnet.
- Routed traffic for a subnet goes through one designated leaf.