Why do reliability teams deliberately set an availability SLO below 100%, and how would you choose between 99.9% and 99.99% for a specific service?
answer
- perfection forbids shipping
- the path to you is not perfect either
- count the minutes, not the nines
- four nines is faster than a human
- the client's network has a ceiling too
basics
~20 sA 100% target is unachievable and forbids all change, since every deploy and dependency can fail. Choose between 99.9% and 99.99% by what users can actually perceive and what the extra nine costs: 99.9% allows about 43 minutes of downtime a month, 99.99% about 4.
solid answer
~50 s100% is the wrong target for three reasons. It is unreachable — the request crosses DNS, the internet, a load balancer and several dependencies, none of which is perfect. It is unmeasurable, because your own measurement pipeline is not 100% reliable. And it is self-defeating: a 100% target means every change is an unacceptable risk, so the service freezes. Choosing the actual number is an economic decision, not a technical one. Look at what the service has really delivered over the last several months, at what users can perceive given their own connectivity, and at what the next nine costs. Concretely, over a 30-day window 99.9% permits about 43 minutes of unavailability and 99.99% permits about 4.3. Four minutes a month is below human reaction time, so the tighter target is really a commitment to automated failover and multi-zone redundancy — engineering spend, not a slogan.
code
python · 5 linesWINDOW_MINUTES = 30 * 24 * 60 # 43200
for target in (0.99, 0.999, 0.9995, 0.9999, 0.99999):
allowed = (1 - target) * WINDOW_MINUTES
print(f"{target:.5%} -> {allowed:8.2f} minutes of unavailability per 30 days")go deeper
Be able to say plainly that 100% is not a target anyone sets, because every deploy and every dependency can fail. Knowing roughly that 99.9% means tens of minutes of downtime a month is enough at this level.
Convert nines into minutes on the spot for a 30-day window, and explain that the shortfall below 100% is what makes shipping changes legitimate. Be ready to name the cost of the next nine as roughly an order of magnitude.
Argue the choice economically: measured history, what users can perceive given their own connectivity, and who funds the redundancy. Point out that a four-nines target is a commitment to automated recovery because four minutes a month is below human response time.
Own the portfolio view — different journeys deserve different targets, and a uniform company-wide number either overspends or underprotects. Be ready to say why an unfunded target is worse than an honest lower one, since a number nobody believes cannot drive any decision.
## Why 100% is the wrong number Setting an availability target of 100% feels like the responsible choice and is in fact the least serious answer a candidate can give. Three separate arguments kill it. **It is unachievable end to end.** A user request traverses their device, their network, DNS, your edge, a load balancer, your service, and whatever your service depends on. Even if your code never failed, the composition of everything in the path does not reach 100%. You cannot promise a level of service that the path to you cannot deliver. **It is unmeasurable.** The SLI is produced by a measurement pipeline that itself drops data, restarts, and has gaps. At four nines you are already close to the noise floor of most measurement systems; at 100% every measurement artifact is a violation. **It removes the ability to change the system.** This is the argument that matters in practice. Every deployment, config push, schema migration and dependency upgrade carries risk. A target of 100% says that no such risk is ever acceptable, which in a real organisation means either the service stops changing or everybody quietly ignores the target. Deliberately targeting less than 100% is what makes shipping legitimate: the shortfall between the target and perfection is the room the team has to work in. ## What the nines actually cost in time These figures are worth memorising, because interviewers ask for them and because they reframe the choice: ``` 30-day window (43,200 minutes) 99% -> 432 min (7.2 hours) 99.9% -> 43.2 min 99.95% -> 21.6 min 99.99% -> 4.32 min 99.999% -> 26 seconds 365-day window 99.9% -> 8.76 hours 99.99% -> 52.6 minutes ``` Read the 99.99% row operationally. 4.32 minutes across a whole month is less time than it takes a human to be paged, wake up, open a laptop and read a dashboard. A four-nines target is therefore not a promise to respond faster; it is a promise that the common failure modes are handled with no human in the loop — automated failover, multi-zone or multi-region capacity, load balancers that eject a bad instance in seconds, deploys that roll back on their own. That is capital and engineering time, and it is the honest content of the extra nine. ## How to choose the number **Start from measured reality.** Compute the SLI over the last six to twelve months. If the service has been delivering 99.92% with the current architecture, a 99.99% target is a re-architecture project disguised as a number, and setting it changes nothing except how often you are in violation. **Ask what users can perceive.** If a large share of users reach you over mobile networks that are themselves around 99% reliable, the difference between 99.9% and 99.99% is invisible to them — you would be buying reliability that is masked by the client's own connectivity. This is the classic argument for not gold-plating a consumer-facing edge service. **Distinguish the journeys.** Different flows deserve different targets. A checkout path and a recommendations sidebar in the same product should not share a number; targeting the whole product at the strictest journey's level is how teams end up paying four-nines costs for a feature nobody would notice failing. **Ask who pays.** The next nine is roughly an order of magnitude more expensive than the last one. Someone has to fund the redundant capacity, the failover testing and the on-call load that comes with it. If the business will not fund it, the target is fiction and everyone will learn to ignore it — which is worse than a lower honest target, because a target nobody believes cannot be used to make any decision. ## Setting it too low is also a failure The symmetric mistake is a target so slack that the service can be visibly bad while formally compliant. If users complain during periods when the SLO says everything is fine, the target is wrong and needs tightening — the SLO is supposed to be a proxy for user happiness, and a proxy that disagrees with the users has failed. The workable target is the loosest one at which users are still satisfied, and it should be revisited as the service and its traffic change.
- Your SLO says the service is healthy, but users are complaining during those same periods. What do you do?Treat it as a defect in the SLO, not in the users. Either the target is too slack, or the indicator is measuring the wrong thing — a whole-service success rate can look fine while one critical journey is broken, and a server-side measurement can look fine while users time out before reaching you. Investigate the complained-about period first, then fix the definition or the threshold.
- Is there any case for a target of 100%?Only for things that are not continuously served and where a single failure is unacceptable — a durability promise, or a correctness invariant such as never double-charging. Even then it is expressed as a hard constraint with its own controls rather than as an availability SLO, because an availability target implies a budget for failure and these do not have one.
- Should every service in a company share the same reliability target?No. A shared default is useful as a starting point, but the right target follows the user impact of the journey. Uniform targets either overspend on services nobody would notice failing or underprotect the one path that earns the revenue. Differentiating also makes the cost of the strictest tier visible to the people asking for it.
saying these in an interview costs you the question
- "We aim for 100% uptime" offered as an ambitious answer
- Picking the number of nines without computing the allowed minutes
- Assuming a tighter SLO is achieved by trying harder rather than by redundancy
- Applying one target to every user journey in the product
- Setting a target far above what the service has historically delivered