Two uplinks in a redundant pair each peak at 65% utilisation; why is this not truly redundant, and how should such links be sized against forecast growth?
answer
- where does the traffic go
- the survivor carries both
- peaks, not averages
- order date before congestion date
basics
~20 sAfter one link fails, the survivor must carry both loads, 2 × 65% = 130% of its capacity, so it congests and drops traffic. Keep each member of a pair below about 50% at peak and order upgrades ahead of lead times.
solid answer
~40 sRedundancy is a capacity claim, not just a topology claim. With two equal links sharing the load, a failure moves everything onto the survivor, so both peaks together must fit one link: each must stay under 50% at peak, and lower to leave room for bursts and uneven hashing. With n equal members and one failure tolerated, each must stay under `(n − 1)/n` — 75% for four. Size from peak percentiles over short intervals, not daily averages, which hide the busy hour. Then forecast: fit the growth rate, compute when the post-failure load crosses your threshold, and subtract the lead time for new circuits or optics — the order date, not the congestion date, is the deadline.
go deeper
Remember that a redundant pair must fit both loads on one link, so each link should stay under about half its capacity at peak.
Explain the (n − 1)/n ceiling for equal members and why measurement must use short-interval peaks rather than averages.
Show the forecast: fit growth, solve for the crossing date, subtract the lead time, and allow for uneven hashing and overlapping maintenance.
Decide what degrades under failure and what the headroom costs, and set an upgrade policy that triggers on forecast dates rather than on alarms.
## Redundant topology versus redundant capacity A diagram with two lines between two boxes looks redundant. It is only redundant if, after one line fails, the other can carry **everything** both were carrying. With both uplinks peaking at 65%, the combined peak is 130% of one link. When one fails, the survivor is asked for 130% of its capacity: queues fill, packets drop, and the "redundant" design delivers an outage-shaped brownout precisely when it is needed. The general rule: **the sum of the load the members carry must fit in the capacity left after the failures you design for.** ## The 50% rule and its generalisation For two equal links sharing load, the survivor must carry both, so each link's peak must stay under half its capacity. With more members, a single failure spreads the lost share over the survivors, so the ceiling rises. | Equal members | Ceiling per member to survive one failure | |---|---| | 2 | 50% | | 3 | about 66.7% | | 4 | 75% | | 8 | 87.5% | The ceiling is `(n − 1) / n`. Treat it as an absolute maximum, not a target, because: - **Distribution is uneven.** Bundles and equal-cost groups hash per flow, so one member can run hotter than the average, and after a failure the survivors rarely share the moved traffic evenly. - **Bursts exceed averages.** Short bursts above the measured peak still need queue space and headroom. - **Two failures happen.** If maintenance on one member overlaps a failure on another, the ceiling for surviving that is `(n − 2) / n`. A common working rule for a pair is therefore to plan an upgrade well before either link reaches 50% at peak. ## Measuring the load you size against Sizing is only as good as the measurement behind it. Use **peak** figures over **short intervals** — for example the 95th or 99th percentile of five-minute samples over the busy period — rather than daily or monthly averages, which smooth away the hour that matters. Look at both directions separately; uplinks are often asymmetric, and the direction that fills first sets the deadline. Even five-minute samples flatten sub-second **microbursts**, so a link whose percentile looks comfortable can still drop packets in bursts; drop and queue counters reveal that where utilisation does not. How the counters are collected, by polling or flow export, is a separate subject; what matters here is that the figure reflects the peak. ## Forecasting growth Capacity work is a race against lead time: new circuits, cross-connects and optics can take weeks to months to deliver. The procedure: 1. **Fit the trend.** Take months of peak figures and estimate the growth rate; compound growth is the usual model. 2. **Pick the threshold.** For a pair, choose the combined peak at which a single survivor would run uncomfortably hot — say 80% of one link. 3. **Solve for the crossing date.** With combined peak `L0` growing at rate `g` per year, the threshold `T` is crossed after `t = ln(T / L0) / ln(1 + g)` years. 4. **Subtract the lead time.** That gives the order date, the real deadline. 5. **Add known events.** A new office, a migration or a product launch is a step change the trend will not show. Worked example: two 10 Gb/s uplinks with a combined peak of 6 Gb/s (30% each), growing 40% a year, with a threshold of 8 Gb/s combined (80% of one link after a failure). `t = ln(8/6) / ln(1.4) ≈ 0.288 / 0.336 ≈ 0.86` years, about **10 months**. If new circuits take three months to deliver, the order must go in by about **month 7**. ## What to do when the forecast says "soon" - **Add a member** to the bundle or equal-cost group, which also raises the per-member ceiling. - **Move to faster links**, which resets the clock by a large factor at once. - **Rebalance** traffic that is concentrated on one path. - **Decide what degrades.** Some designs accept that, after a failure, low-priority traffic is dropped so critical traffic fits; that is a deliberate policy choice made with traffic classes, not an accident discovered during the outage.
- Why can a four-link group fail to absorb a member loss even when the average utilisation is below 75%?Traffic is spread by a per-flow hash, so members are rarely equally loaded, and after a failure the moved flows land unevenly too. A member that was already the busiest can tip over 100% while the average still looks safe. Size against the busiest member and leave margin below the (n − 1)/n ceiling.
- Why size against short-interval peak percentiles rather than daily averages?Daily averages blend quiet nights with the busy hour, so a link averaging 30% can peak far higher. Congestion happens at the peak, so the sizing input must be a percentile of short samples over the busy period, taken per direction.
saying these in an interview costs you the question
- Calling two links redundant because both are up, without checking the combined load
- Sizing links from daily or monthly average utilisation
- Assuming per-flow hashing spreads the moved traffic perfectly evenly after a failure
- Waiting for congestion before ordering capacity that has a long lead time
- Applying the 50% pair rule to every group size regardless of member count