skip to content

You must replace flood-and-learn with BGP EVPN across a 40-rack VXLAN fabric; how would you sequence the move, and what do you trade?

level: principalimportance: nice to knowfreq 8%

answer

  1. control plane first, features later
  2. per-VNI cutover boundaries
  3. type-3 routes take over BUM lists
  4. route scale on every leaf
  5. new failure modes in BGP

basics

~20 s

Bring EVPN up on every leaf first, move VNIs over one at a time so type-3 routes take over BUM membership, then add ARP suppression, anycast gateways and multihoming; you trade flooding for BGP state and control-plane operations on every leaf.

solid answer

~40 s

There is no single right plan, but a defensible one is staged. First run EVPN on every leaf without changing forwarding, and size each leaf for the type-2 routes of every host in the VNIs it serves. Then cut over **per VNI**: RFC 7432 requires remote MACs to be learned from BGP, so whether one VNI can mix data-plane-learning and EVPN VTEPs is implementation behaviour to test, not something to assume. Decide whether type-3 routes signal ingress replication (no underlay multicast, more copies at the source) or keep a multicast tree. Add ARP suppression, the anycast gateway and segment multihoming afterwards, each with its own rollback. You gain less flooding, fast moves and all-active multihoming; you pay with route scale, interoperability testing and a BGP control plane whose mistakes now blackhole layer 2.

go deeper

for a junior

Recall that EVPN replaces flooding with BGP announcements, and that changing a production fabric is done in small, reversible steps.

for a middle

Explain which route types take over which job during a cutover: type 2 for hosts and type 3 for BUM membership.

for a senior

Plan a per-VNI cutover with verification and rollback, keep unknown-unicast flooding until silent hosts are understood, and size leaves for type-2 state.

for a principal

Own the trade: flooding and data-plane simplicity against route scale, interoperability and a BGP failure domain; sequence features so each risk lands alone.

## What you are actually changing In a **flood-and-learn** VXLAN fabric (RFC 7348), each VTEP learns remote MAC addresses from decapsulated traffic, and broadcast, unknown unicast and multicast (BUM) are delivered through a multicast group or a head-end list per VNI. Moving to **BGP EVPN** (RFC 7432, RFC 8365) changes three things at once: - **Remote learning** moves into BGP: RFC 7432 section 9.2 requires a VTEP to learn remote MACs in the control plane, from type-2 MAC/IP Advertisement routes. - **BUM membership** is signalled by type-3 Inclusive Multicast Ethernet Tag routes, whose PMSI Tunnel attribute says ingress replication or a multicast tree. - **New features** become possible: proxy ARP (RFC 9161), the anycast gateway (RFC 9135), Ethernet-segment multihoming and IP prefix routes (RFC 9136). Each is a separate risk, so the judgement is mostly about what to change together and what to keep apart. ## A phased plan 1. **Inventory and size.** Count hosts per VNI and VNIs per leaf. Under EVPN a leaf holds a type-2 route per MAC, plus one per IP binding, for every host in every VNI it serves; check that against each platform's MAC, host-route and BGP table limits before anything changes. 2. **Build the control plane without using it.** Bring up EVPN peering on every leaf with route targets per VNI, so a leaf imports only routes for VNIs it hosts. The session design itself belongs to BGP planning; here it only needs to exist and be monitored. 3. **Pilot one low-risk VNI.** Move every VTEP serving it to EVPN learning together. Verify type-3 routes from every expected VTEP, type-2 routes for every host, and that traffic follows them. 4. **Cut over VNI by VNI.** Keep each cutover small enough to roll back. Leave unknown-unicast flooding enabled at first, for silent hosts. 5. **Add features one at a time.** ARP suppression once bindings are populated; then the anycast gateway, replacing centralised gateways subnet by subnet; then multihoming to replace proprietary switch pairs. 6. **Tighten.** Turn off flooding you have proven unnecessary, and remove the underlay multicast groups if you moved to ingress replication. ## Decisions with no single right answer | Decision | One way | The other way | What tips it | |---|---|---|---| | BUM delivery | ingress replication: no underlay multicast | multicast tree signalled in type-3 routes | BUM volume and the number of VTEPs per VNI versus the team's multicast skills | | Cutover unit | per VNI, all its VTEPs at once | per leaf, all VNIs at once | whether your implementations can mix models within one VNI, which the RFCs do not promise | | Unknown unicast | keep flooding | drop unknown unicast | how many silent hosts you have; RFC 7432 makes it an administrative choice | | ARP suppression scope | answer hits, flood misses | also suppress misses | RFC 9161 does not recommend suppressing misses where hosts move | | External routes | type-5 prefixes | host routes everywhere | table size on leaves and how often prefixes change | ## What you gain - **Less flooding**: remote MACs arrive before traffic, and ARP for known hosts is answered locally. - **Deterministic moves**: MAC Mobility sequence numbers instead of re-learning from traffic. - **All-active multihoming** with mass withdrawal, from a standard rather than a proprietary pair. - **First-hop routing everywhere** with the anycast gateway. ## What it costs - **Route scale**: per-host state on every leaf of a VNI, which flood-and-learn only built on demand. - **Interoperability**: route types, the encapsulation extended community and route-target derivation must agree. RFC 7606 says a BGP speaker discards an EVPN route type it does not recognise, so an older implementation silently ignores type-5 routes; RFC 9136 requires every VTEP of a tenant that routes between segments to support them. - **A new failure domain**: a wrong route-target import, a reflector outage or a policy mistake now blackholes layer-2 traffic. Duplicate-MAC protection (5 moves in 180 seconds by default, RFC 7432) can freeze a flapping MAC and needs monitoring. - **Skills**: the team operating the fabric must operate BGP. ## Rollback and the failure domain Every phase needs an exit. Per-VNI cutovers make rollback a configuration revert on the VTEPs of one VNI. Features layered later can be disabled without touching learning. The step that is hard to undo is decommissioning the old BUM mechanism, so it comes last, after the new one has carried production traffic through at least one failure test.

  • How do you stop every leaf from holding every host route in the fabric?
    Give each VNI its own route targets so a leaf imports only routes for VNIs it serves; RFC 7432 adds that with RT Constraint a route reaches only speakers that import one of its targets. Use type-5 routes for external and summarised prefixes rather than per-host state, and size leaves for their own VNIs, not the whole fabric.
  • Which failure modes does EVPN introduce that flood-and-learn did not have?
    Learning now depends on BGP: a lost session, a wrong route-target import or a policy error leaves hosts unreachable even though the data plane is healthy. Duplicate-MAC detection can stop advertising a flapping MAC until an operator acts. Exhausting a leaf's tables leaves some hosts' routes uninstalled, with implementation-dependent results, and multihoming adds designated-forwarder mistakes.

saying these in an interview costs you the question

  • Moving to EVPN requires removing multicast from the underlay.
  • Flood-and-learn and EVPN VTEPs can always share a VNI during migration.
  • EVPN removes the need to size leaf MAC and host tables.
  • Once EVPN is running, unknown-unicast flooding can be disabled everywhere at once.
  • An EVPN fabric's control plane cannot fail while its underlay is healthy.