In a spine-leaf fabric where each leaf has one 100G uplink to each of six spines, what does losing one spine cost, and what keeps that failure small?
answer
- one of N, not one of two
- ECMP group shrinks by one
- no server loses reachability
- routed links stop storms at the leaf
- no summaries inside the fabric
basics
~20 sEach leaf loses one of six uplinks: 600 Gb/s becomes 500 Gb/s, about a sixth of fabric capacity, and a 2:1 leaf becomes 2.4:1. No server loses reachability; width, routed links and ECMP keep the failure small.
solid answer
~50 sEvery leaf drops one next hop from its ECMP group, so each leaf has 500 Gb/s up instead of 600 Gb/s: the fabric loses one sixth, about 17 %, of its capacity between racks. A leaf with 1,200 Gb/s of servers goes from 2:1 to 2.4:1. Flows that were hashed to the dead spine move to the survivors, and nothing becomes unreachable. Three things keep it small: **width** (losing one of six spines costs a sixth, losing one of two would cost half), **routed links** (a Layer 2 storm or loop stops at the leaf instead of spreading across every switch carrying a VLAN) and **fast, local reaction** (each leaf removes one next hop as soon as it detects the failure). The fabric's real single points are elsewhere: a single-homed rack depends on its one leaf.
go deeper
Recall that every leaf connects to every spine, so losing one spine removes one of several parallel paths rather than cutting any rack off.
Compute the cost: one of N spines lost is 1/N of the fabric bandwidth, and work out the new oversubscription ratio for a leaf.
Separate data-plane and control-plane failure domains, explain ECMP shrinkage, rehashing and why summaries are unsafe in a Clos, and find the real single points.
Choose spine count and server dual-homing by what one failure or one maintenance drain may cost, and weigh that against cabling and device count.
## The arithmetic of losing one spine Take the fabric in the question: six spines, and every leaf has one 100 Gb/s uplink to each. Assume each leaf serves 48 x 25 Gb/s of servers, 1,200 Gb/s. | State | Uplinks per leaf | Uplink bandwidth | Leaf oversubscription | Fabric capacity between racks | |---|---|---|---|---| | All six spines up | 6 | 600 Gb/s | 1,200 / 600 = 2:1 | 100 % | | One spine down | 5 | 500 Gb/s | 1,200 / 500 = 2.4:1 | 5/6, about 83 % | The general rule is that losing one of N spines costs 1/N of the bandwidth between racks. Compare a three-tier design with a pair of aggregation switches: losing one of two costs half of the block's uplink capacity, and if the uplinks were Layer 2 and the failed link was the forwarding one, spanning tree must also reconverge before the blocked path is used. ## What actually happens on the wire 1. Each leaf detects that its link to the failed spine is down, from loss of signal or from a fast liveness protocol such as BFD running on the link. 2. Each leaf removes that spine from its ECMP group for every remote prefix. In a routed fabric these prefixes typically share one next-hop group, so a hierarchical forwarding table changes one entry, not thousands. 3. Flows that hashed to the failed spine are rehashed onto the five survivors. Flows already on the survivors may also move unless the switch uses consistent hashing, which RFC 7938 discusses as a way to limit that churn. 4. No server loses reachability, because every leaf still reaches every other leaf through five spines. RFC 7938 analyses this case for its eBGP-based design: when an upper-tier switch fails, the switches directly below it update their ECMP groups, and in its five-stage example the top-of-rack switches take no part in the reconvergence. ## The two kinds of failure domain A **failure domain** is the set of devices and hosts affected by one fault. It has two faces: - **Data plane.** In a stretched Layer 2 design, a loop or broadcast storm floods every switch that carries the VLAN, which can be the whole data centre. RFC 7938 notes that a fully routed design limits data-plane failure domains to the lowest level of the hierarchy: a storm on one rack stops at its leaf, because routers do not forward broadcasts. - **Control plane.** A routed fabric moves the risk into routing. RFC 7938 points out that, because route summarisation is unsafe inside a Clos, a single failure can make every switch update its routes, so the worst control-plane failure scope is the whole fabric. Few prefixes change in that case, but every switch hears about it. ## What keeps the failure small - **Width.** More, smaller spines mean each failure costs less. Six spines lose a sixth; two lose half. The same width lets you drain a spine for maintenance at a known, small cost. - **Routed leaf-to-spine links.** No spanning tree, no shared broadcast domain, every link active. - **Fast detection.** The faster each leaf removes a dead next hop, the fewer packets go into a black hole. - **No summaries inside the fabric.** A spine has exactly one path to each leaf, so if a spine advertised a summary covering a leaf whose link to it had failed, other leaves would keep sending that leaf's traffic to a spine that cannot deliver it. ## Where the real single points are Losing a spine is the easy case. Losing a **leaf** takes its whole rack off the network unless the servers are dual-homed to a pair of leaves, using multi-chassis link aggregation or EVPN multihoming, each with its own costs. Losing the **border** switches cuts north-south traffic for every rack. A design review asks about all three, not only the spines: | Element lost | Effect with single-homed servers | Effect with dual-homed servers | |---|---|---| | One of six spines | One sixth of inter-rack capacity | One sixth of inter-rack capacity | | One leaf | Its whole rack is offline | Half of that rack's server-facing capacity | | The only border pair | All north-south traffic | All north-south traffic | The last row is why border switches always come in pairs, and why their links to the WAN routers are spread across both of them.
- Why does losing one leaf usually hurt more than losing one spine?A single-homed server has exactly one path into the fabric, its leaf, so losing the leaf takes the rack offline, while losing a spine removes one of N parallel paths. Dual-homing servers to a pair of leaves, with multi-chassis link aggregation or EVPN multihoming, turns a leaf failure into lost capacity instead of lost reachability, at the price of more ports and more complex leaves.
- Why not summarise leaf prefixes at the spines to shrink the control-plane failure domain?In a three-stage Clos each spine has exactly one link to each leaf. If that link fails and the spine keeps advertising a summary that covers the leaf's prefixes, other leaves keep sending traffic for that leaf to a spine that cannot deliver it, a black hole. RFC 7938 treats summarisation inside the fabric as unsafe for this reason.
saying these in an interview costs you the question
- Losing one of six spines takes one sixth of the servers offline.
- Losing a spine forces spanning tree to reconverge across the whole fabric.
- Two large spines are as resilient as six smaller ones of the same total capacity.
- A routed fabric has no failure domain larger than one rack.