How does a spine-leaf data-centre fabric differ from a three-tier access, distribution and core design, and why do data centres prefer it?
answer
- scale up versus scale out
- every leaf to every spine
- one spine between any two racks
- paths equal the number of spines
- folded Clos, RFC 7938
basics
~20 sA three-tier design is a tree that grows by buying bigger upper switches; spine-leaf connects every leaf to every spine, so any two racks are one spine apart, with one equal-cost path per spine and growth by adding switches.
solid answer
~40 sA **three-tier** design stacks access, distribution and core switches as a tree. It suits north-south traffic, but traffic between access blocks climbs through shared upper layers, path lengths vary, and growth means replacing upper-tier switches with bigger ones. With Layer 2 uplinks, spanning tree also blocks half the redundant links. A **spine-leaf** fabric (a folded Clos, as RFC 7938 calls it) cables every leaf to every spine and never leaf to leaf or spine to spine. Any rack reaches any other through exactly one spine, so every inter-rack path has the same length, and a routed fabric spreads flows over one equal-cost path per spine. You scale out: add leaves for ports, add spines for bandwidth, and only when spine ports run out move to a bigger design.
go deeper
Recall the cabling rule: every leaf connects to every spine, and nothing else connects to anything, so any two racks are one spine apart.
Explain how the cabling yields equal path lengths, one equal-cost path per spine and scale-out growth, and contrast that with a tree's scale-up and spanning-tree blocking.
Size a fabric from port counts: spine ports cap the leaves, leaf uplinks cap the spines, and know when a single three-stage fabric has to give way to a five-stage design.
Weigh the fabric's cost in cabling, device count and routing complexity against the tree's ceiling, and say which sites still do not need a full fabric.
## The three-tier design The classic hierarchy, used for campuses and for data centres for many years, has three layers: - **Access** — the switches servers or users plug into. - **Distribution** (called aggregation in data centres) — pairs of switches that join a group of access switches into a block, usually where Layer 2 ends and routing begins, and where policy is applied. - **Core** — a small set of high-capacity switches that join the distribution blocks to each other and to the exit. Each layer up carries more traffic, so it gets bigger, faster switches. RFC 7938 describes this as an upside-down tree whose core is the trunk. Three properties matter in an interview: 1. **Growth is scale-up.** More servers mean bigger distribution and core switches, and eventually no device has enough ports. 2. **Paths are uneven.** Two servers in one block are a few switches apart; two servers in different blocks cross distribution, core and distribution again. 3. **Redundancy is often idle.** When access-to-distribution links are Layer 2, spanning tree keeps the topology loop-free by blocking redundant uplinks, an active/standby model. ## The spine-leaf fabric A **spine-leaf** fabric has two layers in the drawing but is a **3-stage folded Clos** in RFC 7938's terms (the leaf stage is counted twice, once on the way up and once on the way down): - every **leaf** (a top-of-rack switch) connects to every **spine**; - leaves never connect to leaves, spines never connect to spines; - servers, storage, firewalls and the exit all attach to leaves; spines carry only leaf-to-leaf transit. The consequences follow from the cabling: - Any leaf reaches any other leaf through exactly one spine, so every inter-rack path is the same length and latency does not depend on which racks you picked. - The number of equal-cost paths between two leaves equals the number of spines. RFC 7938 notes that the design depends on ECMP with a fan-out at least as large as the number of uplinks. - The links are normally routed, so no spanning tree blocks them and every uplink carries traffic. Workloads that need Layer 2 reach across racks get it from an overlay, typically VXLAN with a BGP EVPN control plane, carried over the routed fabric. ## Side by side | Property | Three-tier tree | Spine-leaf (folded Clos) | |---|---|---| | Built for | North-south traffic | East-west traffic | | Path between racks | Varies with placement | Always leaf, spine, leaf | | Parallel paths | Often two, one blocked by spanning tree | One per spine, all active | | Growth | Replace upper-tier switches | Add leaves for ports, spines for bandwidth | | Typical element | Few large, different chassis | Many identical, smaller switches | | Failure of one upper switch | Often half the block's uplink capacity | One spine's share, 1/N of the fabric | ## How a fabric grows Two counts bound a single 3-stage fabric: 1. **Spine port count caps the number of leaves**, because each leaf needs a port on every spine. 2. **Leaf uplink count caps the number of spines**, because each spine needs a port on every leaf. Adding a leaf adds server ports; adding a spine adds bandwidth between racks, up to the leaves' uplink count. When the spines run out of ports, RFC 7938 describes the next step: either higher-density spines, or more stages, a 5-stage Clos in which groups of leaves and their own spines form a cluster and a further spine layer joins the clusters. ## Why data centres moved RFC 7938's stated requirements explain the shift: a topology that scales horizontally by adding more devices of the same type, a narrow feature set supported by many vendors, and small failure domains. A tree fails the first requirement as soon as the core cannot grow. Spine-leaf meets it because every element is the same kind of switch and every addition is the same move. The trade is cabling and device count: a fabric has many more links, and every leaf must reach every spine. In practice that means: - structured cabling planned for the full spine count from day one, because re-cabling every leaf later is the expensive part; - many identical devices to configure, which pushes operators toward automated, templated configuration; - a routing design that has to handle very wide ECMP fan-out and many neighbours per spine.
- How large can a single three-stage spine-leaf fabric grow?Spine port count caps the leaves and leaf uplink count caps the spines. With 64-port spines you can attach 64 leaves; at 48 server ports each that is 3,072 server ports. Beyond that you buy denser spines or, as RFC 7938 describes, add stages: a 5-stage Clos joins several leaf-and-spine clusters through another spine layer.
- Why are spines not cabled to each other?Every leaf already reaches every spine, so a spine-to-spine link only offers a longer leaf-spine-spine-leaf path that equal-cost routing never prefers. It would consume ports, add failure cases and invite detours under failure, where a clean withdrawal of the broken path is easier to reason about.
saying these in an interview costs you the question
- In spine-leaf, leaves link to each other so traffic between racks avoids the spines.
- Spine-leaf is three-tier minus the core, with spanning tree still blocking half the uplinks.
- Adding spines adds server ports, and adding leaves adds bandwidth between racks.
- Spine-leaf gives every pair of servers identical latency, even two servers in one rack.