skip to content

What is the difference between availability and reliability as quality attributes, and how do MTBF and MTTR relate to the "nines"?

level: middleimportance: must knowfreq 62%

answer

  1. A ≈ MTBF / (MTBF + MTTR)
  2. Reliability = fails rarely; availability = back fast
  3. Series multiplies availability; redundancy multiplies unavailability
  4. 99.9% ≈ 43 min/month
  5. SLI/SLO/error budget; RTO ≠ RPO

basics

~20 s

Reliability = how rarely it breaks. Availability = how much of the time it is usable. Availability ≈ MTBF / (MTBF + MTTR), so you can raise it either by failing less often or by recovering faster. "Three nines" = 99.9% ≈ 43 minutes down per month.

solid answer

~60 s

**Reliability** is the probability of running without failure over a period — it is about failure *frequency* and correctness (MTBF, failure rate, data loss). **Availability** is the fraction of time the system is able to serve — it is about failure frequency *and* recovery speed: `A ≈ MTBF / (MTBF + MTTR)`. The key consequence: a system that fails often but recovers in 2 seconds can be more available than one that fails rarely but takes 6 hours to restore. That is why cloud architecture spends so much on **MTTR reduction** — health checks, fast failover, blue/green and canary deploys, automated rollback, circuit breakers, replicas — rather than only on making components perfect. "Nines" are the usual measure: 99% ≈ 7.3 h/month down, 99.9% ≈ 43 min, 99.99% ≈ 4.4 min, 99.999% ≈ 26 s. Each nine typically multiplies cost and operational sophistication, so it is a business decision expressed as an SLO plus an **error budget**. Caveats I'd raise: time-based uptime is a poor proxy — measure the ratio of successful requests; components in series multiply availability while redundant components multiply *unavailability*; and correlated failures (shared dependency, bad config push) defeat naive redundancy math.

code

text · 8 lines
text
# Series: every hop must work
0.999 ^ 5  = 0.995     -> 99.5%   (~3.6 h/month down)

# Redundancy: unavailabilities multiply (IF independent)
1 - (1 - 0.99) ^ 2 = 0.9999  -> 99.99%

# Correlated failure breaks the second formula:
# same build, same config push, same control plane => both replicas fail together

go deeper

for a junior

State the plain difference (breaks rarely vs is usable), give the MTBF/MTTR formula, and translate one or two nines into downtime per month.

for a middle

Add the composition arithmetic (series multiplies availability, redundancy multiplies unavailability), name concrete MTTR-reducing tactics, and distinguish RTO from RPO.

for a senior

Attack the independence assumption: correlated failures, bad deploys, shared control planes, retry storms. Measure availability as request success ratio, not uptime, and frame targets as SLO plus error budget with an explicit cost curve.

for a principal

Treat reliability as an economic and organizational decision: which journeys deserve which nines, what error budgets buy in delivery speed, how dependency ceilings and degraded modes are designed, and how the organization practices recovery (game days, cell isolation, static stability) rather than trusting redundancy math.

## Definitions, precisely - **Fault** — a defect (bad code, failing disk, wrong config). It may lie dormant. - **Error** — the fault activated: internal state is now wrong. - **Failure** — the error became externally visible: the system deviates from its specification. Users see it. **Reliability**: probability that the system performs correctly for a given duration under stated conditions. Concerned with *how often failures occur* and whether results are correct. Measures: MTBF (mean time between failures), failure rate λ = 1/MTBF, defect escape rate, data loss. **Availability**: the proportion of time (or of requests) during which the system is able to deliver service. Concerned with *how often it fails AND how quickly it comes back*. ``` Availability ≈ MTBF / (MTBF + MTTR) ``` where **MTTR** = mean time to *recover/repair* (often decomposed into MTTD detect + MTTA acknowledge + MTTR repair). **Related but distinct**: **Durability** — data survives once committed (measured by RPO, recovery point objective: how much data you may lose). **Resilience** — ability to keep functioning, possibly degraded, while faults are present. **Recoverability** — RTO, how fast full service returns. A system can be reliable but not durable (never crashes, but no backups), or available but not correct (serves stale/wrong data cheerfully). ## The arithmetic that makes the point | System | MTBF | MTTR | Availability | |---|---|---|---| | A: fails weekly, recovers in 5 s | 168 h | 0.0014 h | ≈ 99.999% | | B: fails yearly, recovers in 8 h | 8,760 h | 8 h | ≈ 99.91% | System A is **less reliable** (fails 52× more often) yet **far more available**. This is the single most important insight in the topic and the reason modern practice optimizes MTTR: you cannot drive MTBF to infinity, but you can often drive MTTR toward seconds with automation. ## The nines table (memorize the middle rows) | Availability | Down per year | Down per month | Down per week | |---|---|---|---| | 99% ("two nines") | 3.65 days | 7.3 h | 1.7 h | | 99.9% | 8.77 h | 43.8 min | 10.1 min | | 99.95% | 4.38 h | 21.9 min | 5 min | | 99.99% | 52.6 min | 4.4 min | 1 min | | 99.999% | 5.26 min | 26 s | 6 s | Practical reading: **99.9% means a single 45-minute incident blows the month.** 99.99% means you cannot afford *any* human in the recovery loop — detection and failover must be automatic. 99.999% usually implies multi-region active-active, no maintenance windows, and a large operations investment; very few business problems justify it. ## Composition: series vs redundancy **In series** (a request must traverse all components): availabilities multiply. ``` 5 services @ 99.9% each -> 0.999^5 = 99.5% (~3.6 h/month down) ``` So decomposing a monolith into a long synchronous call chain *lowers* availability unless you add fallbacks, async boundaries, or caching. **In parallel/redundant** (any one suffices): *unavailabilities* multiply. ``` 2 replicas @ 99% each -> unavailability 0.01 * 0.01 = 0.0001 -> 99.99% ``` This is the promise of redundancy — but it assumes **independence**. ### Where the math lies: correlated failure Redundancy only multiplies if failures are independent. In reality: - both replicas share a power zone, a network, a control plane; - both run the same buggy build (a bad deploy takes out all replicas simultaneously); - both depend on the same config service, DNS, certificate, or feature flag; - a retry storm from the failing side saturates the healthy side (metastable failure). Hence: multi-AZ then multi-region, staged rollouts with canaries, cell-based/shuffle-sharded isolation, bulkheads, and static stability (the system keeps working with the control plane down). Most large outages are correlated-failure and control-plane events, not "two independent disks died". ## Measuring it honestly Time-based uptime is a weak proxy: a system "up" but returning 30% errors or 20-second latencies is not available to users. Prefer: - **Request-success ratio**: good events / valid events (SRE style), sometimes latency-qualified ("served in < 500 ms"). - **Per-journey availability**: checkout availability matters more than a marketing page's. - **SLI → SLO → error budget**: an SLI is the measurement, the SLO the target (99.9% of checkout requests succeed in < 500 ms over 28 days), and the error budget (0.1%) is the permitted failure that funds risky work. Burn the budget → freeze feature launches and spend on stability. This makes reliability a *negotiated economic quantity* rather than a virtue. - An **SLA** is the contractual, customer-facing promise with penalties; always set the internal SLO stricter than the SLA. ## Tactics that move each number **Raise MTBF (reliability):** simpler designs, fewer moving parts, removing single points of failure, input validation, idempotency, backpressure and load shedding to avoid overload collapse, exhaustive testing, gradual rollout, dependency hardening (timeouts, circuit breakers, bulkheads) so a partner's failure is not yours. **Lower MTTR (availability):** health checks and automatic failover, hot/warm standbys, leader election, fast automated rollback, immutable infrastructure and re-create-not-repair, good observability (cut MTTD), runbooks and practiced game days, feature flags for instant disablement, graceful degradation so partial failure isn't total. **Protect durability (RPO):** synchronous replication or quorum writes, write-ahead logs, backups *and tested restores*, cross-region copies. ## The trade-offs to name in an interview - **Availability vs consistency** — under a network partition you must choose (CAP); outside partitions you still trade latency vs consistency (PACELC). Synchronous cross-region replication protects RPO but adds latency and a shared failure mode. - **Availability vs cost** — each nine roughly multiplies infrastructure and operational cost; standby capacity is idle money. - **Availability vs velocity** — most incidents are caused by change; error budgets exist to price that tension explicitly instead of banning deploys. - **Availability vs simplicity** — failover machinery is itself a source of faults (a split brain, a flapping health check, a failover that fails). Redundancy adds failure modes even as it removes others. ## Common interview traps - Quoting nines with no window (99.9% over a *year* tolerates a whole 8-hour outage; over a *month* it does not). - Claiming "we're multi-AZ so we're 99.99%" without accounting for correlated failure and the shared control plane. - Confusing RTO (time to restore service) with RPO (data you may lose) — a system can restore in 1 minute yet lose an hour of writes. - Treating scheduled maintenance as free downtime; users do not distinguish planned from unplanned. - Forgetting that dependencies cap you: you cannot promise 99.99% on top of a 99.9% payment provider unless you can degrade without it.

  • Your service depends on a third-party payment provider that promises 99.9%. Can you offer your customers 99.99%?
    Not on the synchronous path — in series you inherit their unavailability, capping you near 99.9% at best. To exceed it you must remove the hard dependency: degrade gracefully (queue the payment and confirm asynchronously), fail over to a second provider, or cache/authorize offline within a risk limit. The general principle: your availability ceiling is set by your critical dependencies unless you design a working mode without them.
  • What is the difference between RTO and RPO, and why can a system have an excellent RTO and a terrible RPO?
    RTO (recovery time objective) is how fast service returns; RPO (recovery point objective) is how much recent data you may lose. A system restoring from an asynchronous replica can come back in 60 seconds (great RTO) while losing the last 5 minutes of writes (bad RPO). Closing the RPO gap requires synchronous or quorum replication, which costs write latency and couples the regions — the classic durability-vs-latency trade.
  • Why do teams set an internal SLO stricter than the customer-facing SLA?
    The SLA carries financial and reputational penalties, so you want to detect and react well before breaching it. A stricter SLO creates buffer: burning the internal error budget triggers a change freeze and stability work while the contractual promise is still intact. It also accounts for measurement differences — customers measure at their edge, including DNS, network, and client effects you don't see.

Two elevators. One breaks once a year and takes a technician eight hours to fix; the other stalls weekly but self-resets in five seconds. The first is more reliable; the second is what you actually want in a hospital, because almost nobody is ever left waiting. Availability is about the waiting, not the breaking.

saying these in an interview costs you the question

  • Using availability and reliability as synonyms, or ignoring MTTR entirely.
  • Quoting a nines figure without stating the measurement window — 99.9% over a year tolerates an 8-hour outage; over a month it does not.
  • Assuming redundancy multiplies availability without checking for correlated failure (same build, same config push, shared control plane).
  • Assuming that splitting a system into more services improves availability, when synchronous series composition lowers it.
  • Confusing RTO with RPO, or claiming a fast restore implies no data loss.
  • Measuring availability as host uptime rather than the ratio of successful user requests.
  • Promising a higher availability than a critical synchronous dependency provides, with no degraded mode.
  • Treating planned maintenance windows as not counting against availability.

context