Designing a 2,000-server data centre, would you stretch Layer 2 across a spine-leaf fabric or route at every leaf, and what does each choice trade?
answer
- who needs Layer 2 adjacency
- broadcast domain equals failure domain
- MLAG stops at a pair
- routed underlay, overlay on top
- leaf as VTEP, spine as transit
basics
~20 sRoute at every leaf with ECMP across the spines, adding a VXLAN overlay with BGP EVPN only where workloads need Layer 2 reach. Stretched Layer 2 eases VM mobility but makes the broadcast domain the failure domain.
solid answer
~50 sI would route at every leaf: each rack is its own subnet, links are routed, and ECMP uses every spine. That keeps a storm or loop inside one rack and lets the fabric grow by adding switches; the underlay can be an IGP or eBGP, as RFC 7938 describes. **Stretched Layer 2** buys real things, such as VMs keeping their address across racks, legacy apps that need adjacency and Layer 2 load-balancing tricks, but it makes the broadcast domain the failure domain and relies on spanning tree's active/standby links or on multi-chassis link aggregation, which pairs switches without a standard. Where tenants truly need Layer 2, I would carry it as an overlay: leaves as VXLAN tunnel endpoints with a distributed gateway, spines as plain IP transit that may reflect EVPN routes, and border leaves for the exit. The cost is a second layer to operate.
go deeper
Recall the two options: Layer 2 stretched across racks, or each rack routed as its own subnet, and that a broadcast domain is also a failure domain.
Explain what each option gives: address mobility and adjacency for Layer 2, rack-sized failure domains and ECMP across all spines for a routed fabric.
Size the fabric from port counts, place border leaves for the exit, and explain why spines stay transit while leaves act as tunnel endpoints in an overlay.
Pick by workload, not fashion: start routed, add an overlay only for tenants that need Layer 2, and weigh the operational cost of two layers against the failure domain it removes.
## Size the fabric first The server count fixes the shape before the Layer 2 question arises. With 48 server ports per leaf: | Server attachment | Server-facing ports | Leaves needed | |---|---|---| | Single-homed | 2,000 | 2,000 / 48 = 41.7, so 42 | | Dual-homed to a leaf pair | 4,000 | 4,000 / 48 = 83.3, so 84 (42 pairs) | Each leaf needs a port on every spine, so spines with at least 84 leaf-facing ports keep the dual-homed design in a single three-stage Clos; smaller spines would push it toward a five-stage design. Either way the fabric is spine-leaf. The real question is what runs on it. ## Option A — stretch Layer 2 across the racks VLANs span many leaves, so any server can sit in any rack and keep its address. RFC 7938 lists why operators kept Layer 2: legacy applications that need adjacency or non-IP protocols, virtual machines that must keep their IP address when they move to another rack, fewer subnets to manage, and load balancers using Layer 2 direct server return. The costs, also recorded in RFC 7938: - **The broadcast domain is the failure domain.** A loop, a miscabled port or a broadcast storm floods every switch carrying the VLAN, potentially the whole site. - **Spanning tree is active/standby.** It cannot use six spines in parallel; it blocks all but one path. - **Multi-chassis link aggregation** gives active/active Layer 2 links, but most implementations stop at two switches, it has no standard, and the two switches share state that can itself fail. ## Option B — route at every leaf Each leaf is the default gateway for its rack's subnet, every leaf-to-spine link is a routed point-to-point link, and ECMP spreads flows across all spines. - **Failure domain shrinks to the rack** for data-plane faults: routers do not forward broadcasts. - **Scale-out is clean:** a new rack is a new leaf and a new subnet. - **One routing protocol can run the whole fabric.** RFC 7938, an Informational RFC, describes an eBGP-only design for large fabrics, with private-use ASNs from 64512-65534 (RFC 6996), and an example scheme in which all top-tier spines share one ASN and every top-of-rack switch gets its own. An IGP such as OSPF is the common alternative at smaller scale. - **The cost:** no Layer 2 between racks. A server that moves to another rack changes its address, and applications must work over IP. ## Option C — routed underlay, overlay where needed Most multi-tenant fabrics combine the two: Option B as the **underlay**, and a VXLAN **overlay** that carries Layer 2 segments between racks only for the tenants that need them. With a BGP EVPN control plane (RFC 7432, applied to overlays in RFC 8365), the fabric roles are: - **Leaves** are the tunnel endpoints and the provider-edge devices: they learn local hosts, advertise them in BGP, and act as a distributed anycast gateway for each tenant subnet (RFC 9135). - **Spines** stay plain IP transit for the underlay; they never decapsulate tenant traffic, and they often double as route reflectors for the EVPN sessions. - **Border leaves** connect the fabric to the WAN, the internet and the firewalls. RFC 7938's design puts external connectivity in a dedicated cluster of border routers, the only place a default route is originated. The trade is operational: two layers to troubleshoot, an underlay that must carry the overlay's larger packets, and a control plane with many more routes. ## The recommendation, and when it flips | If the workload is... | Then | |---|---| | Containers or stateless services addressed by name | Option B: route at every leaf, no overlay | | Multi-tenant, or VMs that must keep their address | Option C: routed underlay plus an EVPN overlay | | A small set of legacy hosts needing adjacency | Option B, with those hosts kept in one rack or one leaf pair | | A few racks and a team without overlay experience | Option A can be defensible, kept small | For 2,000 servers I would start from Option B and add the overlay only for the workloads that require it, because the routed fabric is the part every rack depends on, and keeping it simple keeps its failures small. The order of work follows from that: 1. Build the routed underlay first and prove it: every leaf reaches every other leaf over all spines, and a spine can be drained without loss. 2. Inventory which workloads truly need Layer 2 reach or address mobility, and which only assumed it. 3. Add the overlay for that set, on the leaves that host it, with the border leaves as the one exit. 4. Keep any remaining stretched Layer 2 to one leaf pair, with a date to retire it.
- Where do firewalls and the internet edge attach in a spine-leaf fabric?On dedicated border leaves, attached to the spines like any other leaf, not on the spines themselves. That keeps every spine identical and interchangeable, puts north-south policy in one place, and gives the fabric one origin for its default route. RFC 7938 describes the same idea as a dedicated cluster of border routers facing the WAN.
- When is stretched Layer 2 still the right answer?When a few racks host workloads that genuinely need adjacency, such as clustered appliances using Layer 2 heartbeats or VMs that must move without readdressing, and the team cannot yet operate an overlay. Keep the stretched domain as small as possible, ideally one leaf pair, and plan the move to a routed underlay with an overlay.
- Why do spines in an EVPN fabric not need to be tunnel endpoints?Leaves encapsulate tenant frames in VXLAN inside an outer IP packet addressed to the remote leaf, so a spine only routes that outer packet like any other. It never sees tenant MAC or VLAN state, which keeps spines simple and replaceable; at most they reflect EVPN routes between leaves.
saying these in an interview costs you the question
- Routing at every leaf rules out moving a VM between racks in any design.
- Multi-chassis link aggregation removes spanning tree's limits for any number of spines.
- In an EVPN fabric the spines must be tunnel endpoints because all traffic crosses them.
- RFC 7938's eBGP design is a standard every data centre is required to follow.
- Once an overlay is in place, the underlay's design no longer matters.