You're told a service must meet a 99.95% availability NFR. Walk through how you'd translate that number into concrete architectural decisions.
answer
- nines -> downtime budget table
- availability of serial deps multiplies, not adds
- SPOF caps the whole chain
- redundancy + automated failover, not manual
- each extra nine costs non-linearly more
basics
~20 s99.95% means about 4.4 hours of downtime allowed per year. To hit that you typically need redundancy (multiple instances/zones), automatic failover, health checks, no single point of failure, and enough capacity headroom to survive losing a node without falling over.
solid answer
~50 sFirst convert the percentage into an allowed-downtime budget (99.95% ≈ 4.38 hours/year, ≈21.9 minutes/month) so the target is concrete. Then identify every single point of failure in the current design — a single DB instance, a single AZ, a single load balancer — because each one caps your achievable availability regardless of anything else. Add redundancy at each layer: multi-instance app tier behind a load balancer with health checks, multi-AZ (or multi-region, if the budget demands it) database with automatic failover, redundant network paths. Combine the availability of each component in the chain (roughly multiplicatively for serial dependencies) to check the target is even reachable with your current dependency list — an NFR of 99.95% is impossible if you depend on three third-party APIs each only guaranteeing 99.9%. Finally, validate with chaos testing (kill a node, kill a zone) rather than trusting the design on paper.
go deeper
Should know that availability is expressed in 'nines' and correspond roughly to a downtime budget, and that redundancy (more than one instance) is the basic tool for improving it.
Should be able to convert a percentage to a concrete downtime budget, identify obvious single points of failure in a given architecture, and describe standard mitigations (multi-AZ, load balancer health checks, database failover).
Should reason quantitatively about combined availability across a dependency chain, know that automated failover time dominates the achievable number more than replica count, and be able to push back on over-specified targets with a cost argument.
Should be able to design an availability strategy across an entire system landscape (not one service), reconcile differing availability targets across teams/vendors with different SLAs, and make the build-vs-cost call on multi-region active-active architecture for the small set of services that actually warrant five-nines.
## Making the percentage concrete Translating an availability NFR like '99.95%' into architecture starts with making the abstract percentage concrete. Availability targets are conventionally expressed in **nines**, and each allows a downtime budget a year: | Target | Downtime a year | |---|---| | 99% (two nines) | about 3.65 days | | 99.9% (three nines) | about 8.76 hours | | 99.95% | about 4.38 hours | | 99.99% (four nines) | about 52.6 minutes | | 99.999% (five nines) | about 5.26 minutes | The first architectural act is converting the percentage into this downtime budget, usually re-expressed per month or per quarter, because '99.95%' as a bare number gives no engineering intuition, but '21.9 minutes a month' immediately tells you whether a single planned maintenance window blows the whole budget. ## The core mechanism: eliminate the single points of failure The core mechanism for hitting an availability target is eliminating single points of failure (SPOFs) and adding automated recovery, because a system's availability is bounded above by its least-available critical-path component. If the availability of dependent components in series is roughly the product of their individual availabilities (component A at 99.99% AND component B at 99.9% in series gives roughly 99.89% combined, not 99.99%), then any single component sitting below the target availability caps the whole chain — no amount of redundancy elsewhere compensates for one unredundant link in the request path. Practically this means: - run multiple instances of every stateless tier behind a load balancer with **active health checks** so a failed instance is removed from rotation within seconds; - spread those instances across **independent failure domains** (availability zones at minimum, regions if the target and blast-radius tolerance demand it); - make the datastore redundant with **automatic failover** (e.g., a managed database with synchronous or near-synchronous replicas and an orchestrator that promotes a replica on primary failure); - remove **manual steps** from the recovery path, because a human paged at 3am adds minutes-to-tens-of-minutes that a 99.95%-or-tighter budget usually can't absorb. ## Why the number is meaningless until you draw the graph Why this exists as a distinct discipline: business stakeholders state availability as a single NFR number, but that number has no meaning until an architect converts it into a dependency graph and checks whether the graph, as designed, can mathematically achieve it. A common and costly mistake is architecting redundancy for the application tier while leaving a hard dependency on a single external API or a single-region managed service that itself only guarantees a lower number; no amount of internal redundancy raises the composite availability above that external ceiling. ## What each extra nine costs The trade-offs are steep and non-linear. Each additional nine of availability roughly multiplies the engineering and operational cost rather than adding to it linearly: - **99.9%** is achievable with multi-AZ deployment and standard managed-service SLAs; - **99.99%** typically requires multi-region active-active or fast active-passive failover, careful handling of data consistency across regions, and much more sophisticated automated failover tooling; - **99.999%** pushes into needing near-zero-downtime deployments, extremely disciplined change management, and often accepting eventual consistency trade-offs because synchronous cross-region replication at that latency budget is prohibitively slow. The architect's job is to push back on over-specified NFRs — if the business asks for 99.99% on a system whose users would genuinely tolerate an hour of monthly downtime, the extra cost (both dollars and velocity, since highly available systems are harder and slower to change safely) isn't justified by real business impact. ## How it fails in production Failure modes show up in predictable ways in production. 1. Teams add redundant instances but leave a single load balancer or single DNS record as an unnoticed SPOF. 2. They add multi-AZ database replicas but the failover is manual or takes 10+ minutes because it was never tested, so it silently caps effective availability far below the redundancy design intended. 3. They hit the target on paper by summing component SLAs additively rather than multiplying, producing an availability number that's mathematically wrong and optimistic. 4. Or they achieve high availability for reads but not writes (e.g., a single-writer database with fast read-replica failover but slow writer promotion), which matters enormously for a write-heavy service like checkout. ## A worked example on a payments API A concrete worked example: a payments API given a 99.95% NFR (≈21.9 min/month budget) is deployed as three app instances across three AZs behind an Application Load Balancer with 10-second health checks (bounding detection+removal of a bad node to under ~30s), backed by a managed relational database with a synchronous standby in a second AZ and automated failover (typically 30–120s for cloud-managed offerings). The team maps every third-party dependency in the payment-authorization path and confirms each guarantees at least 99.95%, escalating or adding a fallback/queue-and-retry path for any that don't. They validate the design isn't just theoretical by running quarterly chaos tests that kill an AZ and measure actual recovery time against the 21.9-minute monthly budget.
- If component A has 99.99% availability and component B (in the same critical request path, called serially) has 99.5% availability, roughly what's the combined availability, and what does that imply for design?Roughly 99.49% (0.9999 × 0.995), essentially capped by the weaker component B. It implies that improving A further has almost no effect on the overall number — the architectural effort has to go toward either replacing/hardening B or removing it from the critical path (e.g., via caching, async fallback, or a circuit breaker with degraded-but-available behavior).
- Why is a fast, automated failover more valuable for availability than adding more redundant replicas?Because availability is bounded by detection-plus-recovery time during an actual failure, not by the number of standing replicas — three replicas with a 10-minute manual failover process still produce 10 minutes of downtime per incident. Automated failover collapses that detection-and-promotion time from minutes to seconds, which usually matters more to the yearly budget than adding a fourth or fifth replica.
- When would you push back on a stakeholder asking for 99.99% availability rather than just building it?When the cost of getting there — multi-region active-active infrastructure, more complex consistency handling, slower and more cautious deploys — clearly outweighs the business impact of the extra downtime it removes, e.g., an internal admin tool where a 30-minute outage a month has near-zero business cost. The right move is to quantify the cost delta between tiers and let the business make an informed trade-off rather than silently over-building.
Like a bike chain: the chain's overall strength is set by its weakest link, not the average of all links — so buying a stronger frame (redundant app servers) doesn't help if one rusty link (a single-instance database or a lone external API dependency) is what actually snaps.
saying these in an interview costs you the question
- Treats the availability percentage as a target to hit by 'adding more servers' without checking for single points of failure
- Sums or averages component availabilities instead of multiplying them across the critical path
- Designs redundancy but leaves failover as a manual, on-call-triggered step
- Never checks whether third-party/external dependencies can even support the target
- Chases additional nines without discussing the non-linear cost with the business
- Never mentions testing/validating the failover path (e.g., chaos testing) — treats the design as correct on paper