skip to content

In network design, what does an availability of 99.99% mean in downtime per year, and how is availability derived from MTBF and MTTR?

level: juniorimportance: must knowfreq 60%

answer

  1. fraction of time in service
  2. 8,760 hours in a year
  3. each extra nine: ten times less
  4. repair time sits in the denominator

basics

~20 s

Availability is the fraction of time a system is in service: MTBF / (MTBF + MTTR). At 99.99% the allowed downtime is 0.01% of a year, about 52.6 minutes; 99.9% allows about 8.76 hours and 99.999% about 5.26 minutes.

solid answer

~40 s

Availability is the share of time a component or service works. For a repairable component it is `MTBF / (MTBF + MTTR)`: the mean time between failures over the whole fail-and-repair cycle. Yearly downtime is `(1 − A) × 8,760 h`, so 99.9% allows about 8.76 hours, 99.99% about 52.6 minutes and 99.999% about 5.26 minutes — each extra nine is a tenfold cut. The formula exposes two levers: fail less often (raise MTBF) or recover faster (cut MTTR). A router with a 50,000-hour MTBF and a 4-hour MTTR reaches about 99.992%; a spare on site that cuts MTTR to 1 hour lifts it to about 99.998% without touching the hardware.

go deeper

for a junior

Know the formula MTBF / (MTBF + MTTR) and the downtime of the common nines: about 8.76 hours, 52.6 minutes and 5.26 minutes a year for three, four and five nines.

for a middle

Explain why MTTR is as powerful a lever as MTBF, and run a quick example: compute availability and yearly downtime from given MTBF and MTTR values.

for a senior

Show that MTTR is an operational design choice — monitoring, spares, procedures — and say what a device figure leaves out: planned work, partial failure and the rest of the path.

for a principal

Frame each nine as a cost decision: show what an extra nine costs in hardware versus in operations, and push back when a target is set without knowing what it buys.

## What availability measures **Availability** is the fraction of time a component, a path or a service is able to do its job. It is a ratio between 0 and 1, usually written as a percentage, and it says nothing about *how* the downtime is distributed: one eight-hour outage and ninety-six five-minute outages can produce the same figure. Network designers use it because it turns "is this resilient enough?" into a number that can be compared against a target and against the cost of improving it. Two terms carry the arithmetic: - **MTBF (mean time between failures)** — the average operating time between one failure and the next for a repairable component. It is a fleet statistic, not a lifetime guarantee: an MTBF of 50,000 hours (about 5.7 years) means failures arrive at that average rate across many units, not that each unit lasts 5.7 years. - **MTTR (mean time to repair or restore)** — the average time from the failure to service being back. In practice it bundles several steps: noticing the failure, diagnosing it, getting a person and a spare to the device, replacing or reconfiguring it, and confirming traffic flows again. ## Availability from MTBF and MTTR A repairable component cycles between working (on average MTBF hours) and being repaired (on average MTTR hours). The steady-state share of time spent working is: ``` A = MTBF / (MTBF + MTTR) ``` Worked example: a switch with an MTBF of 50,000 hours and an MTTR of 4 hours. 1. `A = 50,000 / 50,004 ≈ 0.99992`, i.e. about **99.992%**. 2. Unavailability `U = 1 − A = 4 / 50,004 ≈ 0.00008`. 3. Yearly downtime `≈ 0.00008 × 8,760 h ≈ 0.70 h`, about **42 minutes**. The formula can also be run backwards. To reach 99.99% with a 4-hour MTTR, the component needs `MTBF ≥ 0.9999 × 4 / 0.0001 ≈ 40,000 hours`. ## The nines as downtime The usual shorthand counts the nines. A year is taken as 365 days, or 8,760 hours (using 365.25 days shifts the figures by well under 0.1%). | Availability | Unavailability | Downtime per year | Downtime per 30-day month | |---|---|---|---| | 99% | 1% | about 87.6 hours (3.65 days) | about 7.2 hours | | 99.9% | 0.1% | about 8.76 hours | about 43.2 minutes | | 99.95% | 0.05% | about 4.38 hours | about 21.6 minutes | | 99.99% | 0.01% | about 52.6 minutes | about 4.3 minutes | | 99.999% | 0.001% | about 5.26 minutes | about 26 seconds | Each additional nine divides the allowed downtime by ten, which is why the cost of each nine tends to grow faster than the benefit: the last minutes are the expensive ones. ## The two levers, worked Because `U = MTTR / (MTBF + MTTR)`, and MTTR is tiny next to MTBF for real equipment, unavailability is close to `MTTR / MTBF`. That makes the two levers symmetrical: - **Raise MTBF** — better hardware, fewer moving parts, fewer risky changes. Doubling MTBF roughly halves downtime. - **Cut MTTR** — monitoring that alerts in seconds, spares on site, documented replacement procedures, remote hands. Halving MTTR also roughly halves downtime, and it is often far cheaper. Applied to the switch above: keeping a spare on site so the repair takes 1 hour instead of 4 gives `A = 50,000 / 50,001 ≈ 99.998%`, about **10.5 minutes** of downtime a year — a fourfold improvement without changing the hardware. ## What the figure leaves out - **Scope.** A device's availability is not a path's or a service's. A path through several devices is combined with series and parallel arithmetic, and the service the user sees also depends on servers, DNS and power. - **Planned work.** Some figures exclude maintenance windows; a user still experiences them. State which convention a number uses. - **Partial failure.** A link that drops 5% of packets is "up" by most definitions but useless to a voice call. - **Choosing the target.** Whether a service *should* aim for 99.9% or 99.99% is a business and reliability-policy decision; this arithmetic tells you what a design can deliver, not what it ought to.

  • Why is cutting MTTR often a cheaper route to an extra nine than raising MTBF?
    Unavailability is roughly MTTR / MTBF, so both levers scale it equally, but MTBF is mostly fixed by the hardware you bought while MTTR is set by operations. Faster alerting, spares on site and rehearsed replacement procedures can cut a 4-hour repair to 1 hour, quartering downtime, for the price of a spare and a runbook rather than new equipment.
  • A vendor quotes an MTBF of 200,000 hours; why should a design not assume each unit lasts over 20 years?
    MTBF is an average failure rate across a population, typically estimated for the useful-life period. Individual units fail earlier or later, and early-life defects and wear-out at end of life sit outside that flat-rate estimate. It is useful for comparing designs and computing expected downtime, not for predicting when a particular box will fail.

saying these in an interview costs you the question

  • Saying 99.99% availability allows about eight hours of downtime a year
  • Treating an MTBF figure as the guaranteed lifetime of every unit
  • Ignoring MTTR, as if availability depended only on how often things fail
  • Assuming one extra nine halves the downtime rather than cutting it tenfold
  • Quoting a device's availability as if it were the whole service's