In network availability math, how do you combine components in series versus redundant components in parallel, and what does the parallel formula assume?
answer
- a path needs every hop
- series: every part must work
- parallel: fails only if all fail
- the hidden assumption about failures
basics
~20 sComponents in series must all work, so their availabilities multiply. Redundant components in parallel fail only when all fail, giving 1 − (1 − A)^n for n identical units — but only if their failures are independent.
solid answer
~40 sA path where traffic crosses every element is a series system: `A = A1 × A2 × … × An`, so three 99.9% elements give about 99.70%, worse than any one of them. For small unavailabilities, series unavailabilities roughly add. A parallel group works if any member works, so its unavailability is the product of the members' unavailabilities: for n identical members `A = 1 − (1 − A)^n`, and two independent 99.9% links give 99.9999%. Real designs mix the two: reduce each parallel group to one figure, then multiply the groups in series. The parallel formula assumes independent failures; a shared power feed, fibre duct or software release is a series element hiding inside the redundant pair, and it caps the result.
code
pseudocode · 14 linesfunction series(parts):
a = 1.0
for p in parts:
a = a * p
return a
function parallel(parts):
u = 1.0
for p in parts:
u = u * (1 - p)
return 1 - u
wan = parallel([0.995, 0.995]) // 0.999975
total = series([0.9999, wan, 0.9999]) // about 0.999775go deeper
Recall the two shapes: series availabilities multiply, and a parallel group fails only when every member fails.
Compute a mixed path by collapsing each parallel group to one figure and multiplying the stages, and use the shortcut that small series unavailabilities add.
Name the common causes that break independence — duct, power, software, upstream, change — and model each as a series element so the numbers show what really limits the design.
Use the model to rank spending: compare what each fix removes from the sum of unavailabilities against its cost, and challenge redundancy that only improves an already small term.
## Two building blocks Every availability model of a network path is built from two shapes. - **Series** — traffic must cross every element, so the path is up only when *all* of them are up. A branch router, its access circuit and the provider's edge router form a series chain. - **Parallel** — several elements can each carry the traffic, so the group is up when *at least one* is up. Two uplinks with traffic able to fail over between them form a parallel group. ## Series: availabilities multiply For independent elements in series, the probability that all are working is the product of their availabilities: ``` A_series = A1 × A2 × ... × An ``` Three elements at 99.9% each give `0.999³ ≈ 0.997003`, about **99.70%**, roughly 26 hours of downtime a year — three times the downtime of any one of them. When unavailabilities are small, a useful shortcut is that they **add**: `U_series ≈ U1 + U2 + U3 = 0.3%`. Two consequences follow: 1. A series path is never more available than its weakest element, and is usually less. 2. Every element added in series — another device, another hop, another service — costs availability. ## Parallel: unavailabilities multiply A parallel group fails only if every member is down at once. For independent members the group's unavailability is the product of theirs: ``` U_parallel = U1 × U2 × ... × Un A_parallel = 1 − (1 − A)^n (n identical members) ``` Two independent links at 99.9% give `1 − 0.001² = 0.999999`, about **99.9999%**, about 32 seconds a year. Two 99% links in parallel already reach 99.99%. That is the whole case for redundancy: it multiplies small numbers together. ## Mixed paths: reduce, then multiply Real designs combine both. Work from the inside out: collapse each parallel group into one figure, then multiply the stages in series. | Stage | Shape | Availability | Downtime per year | |---|---|---|---| | Branch router (single) | series | 99.99% | about 52.6 min | | Two WAN circuits, 99.5% each | parallel | 1 − 0.005² = 99.9975% | about 13.1 min | | Core router (single) | series | 99.99% | about 52.6 min | | **Whole path** | product | **about 99.9775%** | **about 118 min (1.97 h)** | The redundant circuits contribute about 13 minutes; the two unprotected routers contribute about 105. The arithmetic points straight at where the next investment belongs: the elements that are still single. ## The independence assumption The parallel formula only holds when the members fail **independently**. Common causes break that: - both circuits run through the **same fibre duct** or enter the building at the same point; - both devices hang off the **same power feed** or the same rack; - both run the **same software release**, so one defect can crash both; - both depend on the **same upstream** element, such as one provider exchange; - one **change** — a mistyped configuration pushed to both — takes both down. Model a shared cause as its own series element: `A = A_shared × (1 − U1 × U2)`. If a shared duct has an unavailability of 0.0002, the "99.9999%" pair of 99.9% links becomes about 99.98%, and the duct dominates: it now accounts for about 105 minutes a year, against under a minute for the two links failing together by chance. ## Common mistakes - **Multiplying a parallel group.** Two 99.9% links "in series" give 99.80%; as a redundant pair they give 99.9999%. Getting the shape wrong moves the answer by orders of magnitude. - **Quoting the group as the path.** A redundant pair's figure says nothing about the single devices in front of and behind it. - **Assuming failover is free.** The formula treats switching to the survivor as instant and certain. In practice a failover takes time (detection and convergence) and can itself fail, so real parallel groups fall somewhat short of the formula. - **Assuming the survivor has room.** A pair counts as parallel only if either member can carry the whole load alone; otherwise one failure still degrades the service. - **Ignoring repair.** These formulas are steady-state figures that already include repair through each element's MTTR; a slow repair process lowers every term at once. ## Using the model The value of this arithmetic is less the final percentage than the ranking it produces. Writing each element's yearly downtime next to it shows where the minutes come from, and the cheapest improvement is almost always to the largest term — usually something that is still single.
- Why does adding a third parallel link usually buy less than fixing a single element elsewhere on the path?A third 99.5% circuit takes the circuit group from about 13 minutes to about 4 seconds of yearly downtime, saving about 13 minutes, while the path still carries about 105 minutes from its two single routers. Doubling one router saves far more. Spend on the largest term in the sum of unavailabilities, which is almost always a series element.
- How do you include a shared power feed in the model of a redundant pair?Treat it as a series element in front of the pair: A = A_power × (1 − U1 × U2). The pair term becomes negligible next to the shared term, which shows that the pair's real availability is roughly the power feed's. The fix is a second, independent feed, not a third link.
A series path is a chain: one broken link drops the load. A parallel pair is two bridges over a river — the crossing closes only if both are out, unless both stand on the same flood-prone bank, where one flood takes both.
saying these in an interview costs you the question
- Saying a series path is as available as its best component
- Multiplying the availabilities of redundant links as if they were in series
- Assuming two links in the same duct fail independently
- Believing redundancy simply halves downtime rather than multiplying unavailabilities
- Spending on a third parallel link while single devices remain in the path