skip to content

A design review claims a branch site reaches four nines because it has two WAN circuits at 99.9% each; how do you evaluate that claim and decide what to change?

level: principalimportance: should knowfreq 18%

answer

  1. the pair is not the site
  2. trace the whole path
  3. single elements in series
  4. shared fate breaks the formula

basics

~20 s

Two independent 99.9% circuits do reach 99.9999% as a pair, but the site is the whole path: a single edge router, power feed or shared duct in series caps it far lower. Find those elements, quantify each, fix the largest.

solid answer

~40 s

First, recompute: the circuit pair is `1 − 0.001² = 99.9999%` only if the circuits fail independently. Then model the whole site path, because users experience the path, not the pair. Trace every element in series — the single edge router, the building power, the shared duct or entry, a common upstream exchange, one software release on both devices — and give each a figure. With one 99.95% router and a shared duct at 0.02% unavailability, the site lands near 99.93%, about six hours a year. Rank the terms by downtime and cost to remove: a second router on diverse paths, separate building entries, a second power feed, staggered upgrades, a shorter MTTR. Recompute after each change — and ask whether the business needs four nines at all.

go deeper

for a junior

Recall that a redundant pair only helps the part of the path it covers; anything single in the path still limits the whole.

for a middle

Compute the site path with series and parallel arithmetic and show that a single router caps the result far below the circuit pair's figure.

for a senior

Trace every shared cause — router, power, duct, upstream, software, gateway — and model it explicitly, then verify diversity claims with route records and failure tests.

for a principal

Lead the review to a costed decision: rank fixes by downtime removed per unit of spend, and challenge whether the site needs four nines at all.

## Recompute the claim The arithmetic in the claim is not wrong; its scope is. Two circuits at 99.9% each, failing independently, are down together `0.001 × 0.001 = 0.000001` of the time, so the **pair** is about 99.9999%, roughly 32 seconds a year. But the site's users do not experience the pair; they experience the **whole path** from their access switch to the far end, and that path contains elements that are not duplicated. ## Hunt the series elements Walk the path physically and logically and list everything traffic depends on that exists only once: - **The edge router.** Two circuits into one router make the router a series element. - **Power.** One building feed, one UPS, one rack power strip. - **The physical route.** Both circuits through one duct, one building entry or one street; a single excavation cuts both. - **The upstream.** Two circuits from the same provider exchange, or two carriers who lease the same fibre. - **Software and change.** Two devices on the same release share every defect in it; one change pushed to both can break both. - **The LAN side.** If hosts point at one gateway address with no gateway-redundancy protocol, a second router does nothing for them. ## Model it Give each element an availability — from vendor data, provider history or your own incident records — and combine them: shared causes in series, independent alternatives in parallel. With illustrative figures: | Element | Unavailability | Downtime per year | |---|---|---| | Single edge router at 99.95% | 0.0005 | about 263 min | | Shared duct for both circuits | 0.0002 | about 105 min | | Circuit pair, independent part | 0.000001 | under 1 min | | **Site path** (about 99.93%) | **about 0.0007** | **about 6.1 hours** | The circuits, the subject of the claim, contribute under a minute. The router and the duct contribute almost everything. ## Where the figures come from Every number in such a model is an estimate, and the review should say how good each one is: - **Device figures** come from published MTBF and your own MTTR. The MTTR is usually the weaker number — a branch two hours' drive from the nearest spare has a very different repair time from one with a spare on the shelf. - **Circuit figures** come from the provider's service terms and, better, from your own outage history with that provider at that site type. - **Shared-cause figures** — a duct cut, a building power event, a release defect — are the hardest to estimate and the most important, because they sit in series. Use incident history across the estate, and when in doubt run the model with a pessimistic and an optimistic value. If the conclusion flips between the two values, the design is too close to the target to claim it; if it holds either way, the number is good enough to act on. ## Decide what to change Rank fixes by the downtime they remove per unit of cost. 1. **Second edge router, each on its own circuit and building entry.** Each path is now router × circuit in series, `0.9995 × 0.999 ≈ 0.9985`, and the two paths are in parallel: `1 − 0.0015² ≈ 99.9998%`. Paths must be computed this way — series inside each path, then parallel across paths — unless the routers are cross-connected to both circuits. 2. **Verify the diversity.** Ask for route records showing separate entries, ducts and upstream exchanges; two contracts with two carriers do not prove two physical paths. 3. **Remove the next series term.** With the paths fixed, a single building power feed at an assumed 99.99% becomes the cap: `0.9999 × 0.999998 ≈ 99.9898%`, about 53.7 minutes a year — still just short of four nines' 52.6-minute allowance. A second feed or a battery runtime long enough to cover typical outages is the next spend. 4. **Cut MTTR.** Monitoring, spares on site and rehearsed procedures shrink every remaining term at once. 5. **Separate the software.** Upgrade the two routers at different times, so a defective release shows itself on one router before it reaches the other. ## Question the target Four nines at a branch can cost far more than three nines. The review should end with numbers the business can price: what each option costs, the downtime it removes, and what that downtime costs. Sometimes the right answer is to accept about 99.95% and spend the money on faster repair. Whatever is chosen, test it: fail each element on purpose and measure, because a model with a hidden shared cause is only as good as the trace behind it.

  • With two routers and two circuits, why compute each router-plus-circuit path in series before combining the paths in parallel?
    Without cross-connection, a path works only when both its router and its circuit work, so each path is a series pair. The site is up if either path is up, so the paths combine in parallel. Combining routers and circuits as separate parallel groups would assume any router can use any circuit, which overstates availability.
  • How do you check that two carriers' circuits are physically diverse?
    Ask both carriers for route records showing the building entry, the ducts and the upstream exchanges each circuit uses, and compare them. Carriers often lease the same fibre or share a duct near the building, so separate contracts prove nothing. Re-check after carrier network changes, since diversity can silently disappear.

saying these in an interview costs you the question

  • Quoting the redundant pair's availability as the whole site's
  • Treating two carriers on the contracts as proof of physical diversity
  • Combining routers and circuits as separate parallel groups when they are not cross-connected
  • Adding more circuits while the single edge router still dominates downtime
  • Accepting a four-nines target without pricing what each nine costs