skip to content

Designing the IP underlay for a multi-tenant VXLAN fabric of 40 leaves, what must it give the VTEPs, and which routing, MTU and multicast choices would you make?

level: principalimportance: should knowfreq 14%

answer

  1. keep the underlay boring
  2. only tunnel addresses in its tables
  3. every uplink forwarding
  4. one MTU everywhere
  5. multicast only if flooding needs it

basics

~10 s

A routed leaf-spine with no spanning tree; a routing protocol carrying only VTEP addresses and links, with full-width ECMP and no summarisation; one fabric-wide jumbo MTU; and unicast-only BUM handling unless flood-and-learn needs multicast.

solid answer

~50 s

I would keep the underlay boring: its only job is IP reachability between VTEP addresses, with enough paths and a large enough MTU. **Topology**: a routed leaf-spine, so every uplink forwards and no spanning tree blocks links - RFC 7348 names exactly that as the reason for an IP underlay. **Routing**: eBGP as RFC 7938 describes, or a link-state IGP; either is defensible if it carries only VTEP addresses and fabric links, installs ECMP across every spine, detects failure quickly and never summarises VTEP addresses. **MTU**: one jumbo value on every link, with headroom above 1,550 (IPv4) or 1,570 (IPv6). **BUM**: underlay multicast only if I run flood-and-learn with multicast groups; with an EVPN control plane and ingress replication the underlay stays unicast-only, and the ingress VTEP pays by making the copies. Tenant MACs, VNIs and routes live in the overlay, never in the underlay.

go deeper

for a junior

Recall that the underlay is the routed network carrying tunnel packets between VTEPs, and that it never sees the tenants' own addresses.

for a middle

Explain why a routed leaf-spine keeps every uplink active, and why its MTU must be raised on every link for VXLAN.

for a senior

Show operational detail: ECMP width matching the spines, fast failure detection, unsummarised VTEP addresses, and an MTU probe after every change.

for a principal

Own the trade-offs: eBGP against a link-state IGP, multicast against ingress replication, and keeping tenant state out of the underlay's failure domain.

## The underlay's one job In a **VXLAN** fabric (RFC 7348) every tenant frame travels between two **VTEPs** (VXLAN Tunnel End Points) inside an outer IP packet. The **underlay** - the physical routed network - only has to deliver those outer packets: from one VTEP address to another, over as many paths as possible, without dropping the larger encapsulated sizes. Tenant MAC addresses, VNIs, tenant VRFs and tenant routes all live in the overlay. That separation is the main design principle: the underlay's tables hold the 40 leaves' tunnel addresses and the fabric's own links, so a tenant's churn never touches the underlay's routing. ## Topology: routed, with every uplink forwarding RFC 7348 motivates VXLAN with exactly this choice: a Layer 2 physical network kept loop-free by spanning tree can leave a large number of links disabled, while an IP network can use **ECMP** across all of them. So the underlay is a routed **leaf-spine** fabric: each leaf connects to every spine, and any two leaves are joined by one equal-cost path per spine. Losing a spine removes a share of capacity, not reachability. How to size such a fabric - port counts, oversubscription, how many spines - is a general data-centre topology question; what VXLAN adds is that every one of those paths must carry the encapsulated MTU and must be hashable by the outer UDP source port. ## Routing: two defensible choices | Choice | What it buys | What it costs | |---|---|---| | eBGP on every fabric link, as RFC 7938 (Informational) describes | A simpler state machine riding on TCP; a failure is masked where an alternate path exists instead of flooded fabric-wide | An ASN plan, and tuning to converge as fast as an IGP | | A link-state IGP | Familiar operation and a complete topology view on every router | Every event floods the whole area, which grows with the fabric | Whichever is chosen, the underlay must: - **Install ECMP across every spine**, so the outer source-port entropy has paths to spread over. - **Detect failures fast**, often with BFD (RFC 5880), which RFC 7938 notes can give sub-second failure detection, so tunnels move off a dead link quickly. - **Never summarise VTEP addresses inside the fabric.** RFC 7938 warns that summarisation within a Clos topology black-holes traffic under single link failures: a spine that loses its link to one leaf would keep advertising a covering summary and drop that leaf's share of the tunnels. - **Advertise each VTEP's tunnel address as a host route** that stays up while any uplink survives; how a VTEP picks and uses that address is a VTEP question. ## MTU: one value, everywhere VTEPs MUST NOT fragment VXLAN packets (RFC 7348), so the underlay MTU is a correctness requirement. The minimum is an IP MTU of 1,550 for an IPv4 underlay and 1,570 for IPv6 with a 1,500-byte tenant MTU; a fabric standard usually sets a jumbo value with headroom - its size is an implementation choice - on both ends of every link, and verifies it with a DF-set probe between VTEPs after every change. ## BUM transport: does the underlay need multicast? | Approach | Underlay requirement | Cost | |---|---|---| | Flood-and-learn with multicast groups (RFC 7348 section 4.2) | Multicast routing in the underlay, mapping VNIs to groups | Another protocol and its state to run and debug | | Ingress replication, usually with an EVPN control plane (RFC 8365) | Unicast reachability only | The ingress VTEP sends one copy per remote VTEP | With a control plane advertising which VTEPs serve which VNI, the underlay can stay unicast-only, removing a whole protocol from its failure surface. Multicast earns its place when tenants send heavy multicast, or when the fabric still learns by flooding. ## The judgment calls 1. **Simplicity over cleverness.** Every feature in the underlay - multicast, summarisation, tenant routes - is something that can fail under every tenant at once. 2. **Failure domains.** Tenant churn stays in the overlay; underlay convergence affects only tunnel reachability. 3. **Growth.** Adding spines widens ECMP and must not shrink the MTU or break the hash; adding leaves adds one tunnel address each. 4. **Verification.** MTU, ECMP width and VTEP reachability are checked continuously, because they fail silently.

  • Why not summarise the leaves' VTEP addresses at the spines to shrink the routing tables?
    Each spine reaches each leaf over one direct link. If that link fails, a spine still advertising a covering summary keeps attracting traffic for that leaf's VTEP and drops it, while every other leaf keeps hashing a share of flows toward that spine. RFC 7938 warns that summarisation inside a Clos topology black-holes traffic under single link failures; with 40 tunnel addresses the tables are small anyway.
  • When would you accept multicast in a VXLAN underlay?
    When the fabric runs flood-and-learn with multicast groups for BUM traffic and has no control plane to replace it, or when tenants send heavy multicast that ingress replication would multiply by the number of remote VTEPs. Otherwise an EVPN control plane with ingress replication keeps the underlay unicast-only and removes a protocol from its failure surface.

saying these in an interview costs you the question

  • The underlay should carry the tenants' routes so the VTEPs can reach the VMs.
  • VXLAN always requires IP multicast in the underlay.
  • Summarise the VTEP addresses at the spines; a failed link will simply reroute.
  • eBGP is the only underlay routing protocol the RFCs allow for VXLAN.
  • A spanning-tree Layer 2 underlay gives VXLAN the same multipath as a routed one.