skip to content

Two uplinks in a redundant pair each peak at 65% utilisation; why is this not truly redundant, and how should such links be sized against forecast growth?

level: seniorimportance: should knowfreq 35%

answer

  1. where does the traffic go
  2. the survivor carries both
  3. peaks, not averages
  4. order date before congestion date

basics

~20 s

After one link fails, the survivor must carry both loads, 2 × 65% = 130% of its capacity, so it congests and drops traffic. Keep each member of a pair below about 50% at peak and order upgrades ahead of lead times.

solid answer

~40 s

Redundancy is a capacity claim, not just a topology claim. With two equal links sharing the load, a failure moves everything onto the survivor, so both peaks together must fit one link: each must stay under 50% at peak, and lower to leave room for bursts and uneven hashing. With n equal members and one failure tolerated, each must stay under `(n − 1)/n` — 75% for four. Size from peak percentiles over short intervals, not daily averages, which hide the busy hour. Then forecast: fit the growth rate, compute when the post-failure load crosses your threshold, and subtract the lead time for new circuits or optics — the order date, not the congestion date, is the deadline.

go deeper

for a junior

Remember that a redundant pair must fit both loads on one link, so each link should stay under about half its capacity at peak.

for a middle

Explain the (n − 1)/n ceiling for equal members and why measurement must use short-interval peaks rather than averages.

for a senior

Show the forecast: fit growth, solve for the crossing date, subtract the lead time, and allow for uneven hashing and overlapping maintenance.

for a principal

Decide what degrades under failure and what the headroom costs, and set an upgrade policy that triggers on forecast dates rather than on alarms.

## Redundant topology versus redundant capacity A diagram with two lines between two boxes looks redundant. It is only redundant if, after one line fails, the other can carry **everything** both were carrying. With both uplinks peaking at 65%, the combined peak is 130% of one link. When one fails, the survivor is asked for 130% of its capacity: queues fill, packets drop, and the "redundant" design delivers an outage-shaped brownout precisely when it is needed. The general rule: **the sum of the load the members carry must fit in the capacity left after the failures you design for.** ## The 50% rule and its generalisation For two equal links sharing load, the survivor must carry both, so each link's peak must stay under half its capacity. With more members, a single failure spreads the lost share over the survivors, so the ceiling rises. | Equal members | Ceiling per member to survive one failure | |---|---| | 2 | 50% | | 3 | about 66.7% | | 4 | 75% | | 8 | 87.5% | The ceiling is `(n − 1) / n`. Treat it as an absolute maximum, not a target, because: - **Distribution is uneven.** Bundles and equal-cost groups hash per flow, so one member can run hotter than the average, and after a failure the survivors rarely share the moved traffic evenly. - **Bursts exceed averages.** Short bursts above the measured peak still need queue space and headroom. - **Two failures happen.** If maintenance on one member overlaps a failure on another, the ceiling for surviving that is `(n − 2) / n`. A common working rule for a pair is therefore to plan an upgrade well before either link reaches 50% at peak. ## Measuring the load you size against Sizing is only as good as the measurement behind it. Use **peak** figures over **short intervals** — for example the 95th or 99th percentile of five-minute samples over the busy period — rather than daily or monthly averages, which smooth away the hour that matters. Look at both directions separately; uplinks are often asymmetric, and the direction that fills first sets the deadline. Even five-minute samples flatten sub-second **microbursts**, so a link whose percentile looks comfortable can still drop packets in bursts; drop and queue counters reveal that where utilisation does not. How the counters are collected, by polling or flow export, is a separate subject; what matters here is that the figure reflects the peak. ## Forecasting growth Capacity work is a race against lead time: new circuits, cross-connects and optics can take weeks to months to deliver. The procedure: 1. **Fit the trend.** Take months of peak figures and estimate the growth rate; compound growth is the usual model. 2. **Pick the threshold.** For a pair, choose the combined peak at which a single survivor would run uncomfortably hot — say 80% of one link. 3. **Solve for the crossing date.** With combined peak `L0` growing at rate `g` per year, the threshold `T` is crossed after `t = ln(T / L0) / ln(1 + g)` years. 4. **Subtract the lead time.** That gives the order date, the real deadline. 5. **Add known events.** A new office, a migration or a product launch is a step change the trend will not show. Worked example: two 10 Gb/s uplinks with a combined peak of 6 Gb/s (30% each), growing 40% a year, with a threshold of 8 Gb/s combined (80% of one link after a failure). `t = ln(8/6) / ln(1.4) ≈ 0.288 / 0.336 ≈ 0.86` years, about **10 months**. If new circuits take three months to deliver, the order must go in by about **month 7**. ## What to do when the forecast says "soon" - **Add a member** to the bundle or equal-cost group, which also raises the per-member ceiling. - **Move to faster links**, which resets the clock by a large factor at once. - **Rebalance** traffic that is concentrated on one path. - **Decide what degrades.** Some designs accept that, after a failure, low-priority traffic is dropped so critical traffic fits; that is a deliberate policy choice made with traffic classes, not an accident discovered during the outage.

  • Why can a four-link group fail to absorb a member loss even when the average utilisation is below 75%?
    Traffic is spread by a per-flow hash, so members are rarely equally loaded, and after a failure the moved flows land unevenly too. A member that was already the busiest can tip over 100% while the average still looks safe. Size against the busiest member and leave margin below the (n − 1)/n ceiling.
  • Why size against short-interval peak percentiles rather than daily averages?
    Daily averages blend quiet nights with the busy hour, so a link averaging 30% can peak far higher. Congestion happens at the peak, so the sizing input must be a percentile of short samples over the busy period, taken per direction.

saying these in an interview costs you the question

  • Calling two links redundant because both are up, without checking the combined load
  • Sizing links from daily or monthly average utilisation
  • Assuming per-flow hashing spreads the moved traffic perfectly evenly after a failure
  • Waiting for congestion before ordering capacity that has a long lead time
  • Applying the 50% pair rule to every group size regardless of member count