skip to content

A data-centre leaf switch has 48 x 25G server ports and 6 x 100G spine uplinks; what is its oversubscription ratio, and when does that ratio hurt?

level: seniorimportance: must knowfreq 35%

answer

  1. bandwidth, not port count
  2. down over up
  3. rack-local traffic is free
  4. a bet on simultaneity
  5. ratios multiply across stages

basics

~20 s

Server-facing bandwidth is 48 x 25 = 1,200 Gb/s and uplink bandwidth is 6 x 100 = 600 Gb/s, so the leaf is 2:1 oversubscribed. It hurts when many servers send off-rack at once: replication, shuffles, backups, incast bursts.

solid answer

~50 s

Oversubscription compares what the servers can send with what can leave the rack: `48 x 25 Gb/s = 1,200 Gb/s` down against `6 x 100 Gb/s = 600 Gb/s` up, so **2:1**. It is not a drop rate: if every server sent off-rack at line rate at once, the uplinks could carry only half. Traffic between two servers on the same leaf never touches the uplinks, so it does not count. The ratio hurts when off-rack traffic is sustained and synchronised: storage replication and rebuilds, distributed jobs that shuffle data all-to-all, backup windows, and many-to-one bursts that overflow the leaf's buffers. General-purpose racks often run higher ratios and storage or compute clusters near 1:1, a cost decision rather than a standard. In a three-stage fabric the spines add no ratio of their own, so the leaf is where you set it.

go deeper

for a junior

Recall that oversubscription is server-facing bandwidth divided by uplink bandwidth, and compute it in Gb/s rather than in ports.

for a middle

Explain that the ratio is a worst case, that rack-local traffic does not count, and that one flow is limited to one uplink.

for a senior

Name the workloads that break the bet, read bursts and retransmissions rather than averages, and show how uplink count, speed or server density changes the ratio.

for a principal

Set the ratio per workload as a cost decision, keep it at one layer so applications see one bandwidth pool, and know when ratios would compound.

## The calculation **Oversubscription** is the ratio between the bandwidth a switch's downstream devices can offer and the bandwidth it has toward the rest of the network. For the leaf in the question: | Side | Ports | Speed | Bandwidth | |---|---|---|---| | Server-facing (down) | 48 | 25 Gb/s | 1,200 Gb/s | | Spine-facing (up) | 6 | 100 Gb/s | 600 Gb/s | | **Ratio** | | | **1,200 : 600 = 2:1** | Use bandwidth, not port counts. RFC 7938, the Informational RFC on large data-centre fabrics, states the rule in port counts, a fabric is non-blocking when the uplink count M is at least the downlink count N and oversubscribed by N/M otherwise, but that assumes all ports run at one speed. Counting ports here would give 48 / 6 = 8:1, which is wrong by a factor of four because the uplinks are four times faster. ## What the ratio does and does not mean - **It is a worst case.** 2:1 means that if all 48 servers sent off-rack at full line rate at once, only half of that could leave. It says nothing about average loss. - **Rack-local traffic is free.** Two servers on the same leaf exchange traffic through the leaf alone; the uplinks never see it. - **It is a bet on simultaneity.** Most servers spend most of their time well below line rate and not all at once, so a ratio above 1:1 buys more servers per unit of fabric bandwidth. - **One flow is still bounded by one uplink.** Routing spreads flows across the six uplinks per flow, so a single large flow cannot use more than one 100 Gb/s path no matter how idle the others are. ## When it hurts The bet loses when off-rack traffic is both heavy and synchronised: 1. **Storage replication and rebuilds** — losing a disk or a node triggers bulk copies to and from many racks at once. 2. **All-to-all compute jobs** — distributed jobs that shuffle data between every node keep every uplink busy at the same time. 3. **Backup and migration windows** — scheduled jobs start together and run for hours. 4. **Incast** — many servers answer one requester at the same instant; the burst overflows the leaf's egress buffer even though average utilisation is low, and the senders' TCP stacks back off. The symptom is rarely a clean saturation graph. Five-minute averages can show 20 % utilisation while microsecond bursts drop packets, so latency and retransmission metrics tell you more than averages do. ## Choosing and changing the ratio The ratio is a design choice, not a protocol constant. Common practice is to accept a few to one for general-purpose racks and to build storage and compute clusters near 1:1, but those are conventions and cost decisions. To change it for this leaf: - **More uplinks:** 8 x 100 Gb/s gives 1,200 / 800 = 1.5:1, and 12 x 100 Gb/s gives 1:1, but each uplink needs a port on a spine, so 12 uplinks mean twelve spines or two links per spine. - **Faster uplinks:** 6 x 400 Gb/s gives 2,400 Gb/s up, more than the 1,200 Gb/s down, so the leaf is no longer oversubscribed. - **Fewer servers per leaf:** using 24 of the 48 ports gives 600 / 600 = 1:1 at the cost of more leaves. ## Where else oversubscription can hide In a **three-stage spine-leaf fabric** every spine port faces a leaf, so traffic entering a spine leaves it on another spine port of the same speed and the spine layer adds no ratio of its own. The leaf is where oversubscription lives. RFC 7938 observes the same thing and recommends introducing oversubscription at a single layer, so that applications do not have to reason about separate bandwidth pools within a rack, between racks and between clusters. In a **five-stage Clos**, where clusters of leaves and spines are joined by another spine layer, a cluster can be oversubscribed toward that layer too. Ratios then **multiply**: 3:1 at the leaves and 2:1 at the cluster's uplinks give 6:1 for traffic between clusters. That compounding is the reason to keep the ratio in one place.

  • Do the spines in a three-stage fabric add oversubscription of their own?
    No. Every spine port faces a leaf, so traffic entering a spine leaves on another port of the same speed, and the spine layer is non-blocking as long as the ports match. Oversubscription is introduced at the leaf. In a five-stage design a cluster's uplinks can add a second ratio, and the two multiply.
  • Why can a 2:1 leaf drop packets while its uplinks average 10 % utilisation?
    Averages hide bursts. When many servers answer one requester at once, or a job starts on every node together, the leaf's egress queue toward one uplink or one server overflows for microseconds and drops packets. The five-minute average stays low, so retransmissions and tail latency are the signals to watch.

An airline that sells more tickets than seats is not bumping half its passengers; it is betting that not everyone shows up at once. Oversubscription is the same bet about servers sending off-rack, and it only fails on the day they all do.

saying these in an interview costs you the question

  • A 2:1 oversubscription ratio means the leaf drops half of all traffic.
  • Traffic between two servers on the same leaf uses uplink bandwidth.
  • Count ports, not bandwidth: 48 server ports over 6 uplinks is 8:1.
  • Every fabric must be 1:1; any oversubscription is a design error.
  • Adding spines always fixes oversubscription, whatever the leaves' free uplink ports.