skip to content

A large prospective customer asks for a 99.99% availability SLA with financial penalties. As the engineering owner in that negotiation, how do you decide what to commit to?

level: principalimportance: nice to knowfreq 32%

answer

  1. judge the worst month, not the year
  2. serial dependencies multiply
  3. 4.32 minutes rules out humans
  4. cap the credits
  5. the first bespoke SLA sets a precedent

basics

~20 s

Start from the worst measured month, not the average, then compose the dependency chain: serial dependencies multiply, so three components at 99.99% cap you near 99.97%. Four nines allows 4.32 minutes a month, which is an automated-failover architecture, not a promise you can talk your way into.

solid answer

~50 s

Three checks before any number is agreed. First, measured history — what has the indicator actually done over the last twelve months, judged on the worst month, because that is the month the contract is evaluated on. Second, composition: serial dependencies multiply, so a service sitting on three components each at 99.99% cannot exceed roughly 99.97% without redundancy between them. Third, what the number implies architecturally — 99.99% is 4.32 minutes of unavailability a month, which is shorter than any human response, so it is a commitment to automated failover, multi-zone or multi-region capacity, and the spend that comes with it. If the answer is still yes, the negotiation moves to the clauses that decide what the number means: whose measurement counts, what an outage is, which exclusions apply, and a credit cap proportionate to the contract. And remember the number is a precedent — the next customer will ask for the same terms.

code

python · 11 lines
python
WINDOW_MINUTES = 30 * 24 * 60

def serial_availability(components):
    a = 1.0
    for c in components:
        a *= c
    return a

chain = [0.9999, 0.9999, 0.9999]   # three serial dependencies, no redundancy
a = serial_availability(chain)
print(f"ceiling {a:.5%}, ~{(1 - a) * WINDOW_MINUTES:.1f} min unavailable per 30 days")

go deeper

for a junior

Know that four nines means roughly four minutes of downtime a month and that a contract number should never exceed what the service has actually delivered. The detailed negotiation is above this level, but the arithmetic is not.

for a middle

Be able to compute the allowance and to explain why serial dependencies multiply rather than average. Recognise that an annual average is a weak basis for a per-month contractual commitment.

for a senior

Connect the number to the architecture: 4.32 minutes a month means automated failover and multi-zone capacity, not faster paging. Be ready to list the contract clauses — measurement point, downtime definition, scope, exclusions — that decide what the percentage actually means.

for a principal

Own the commercial framing: cap the remedy at proportionate service credits, treat the first bespoke SLA as a precedent across the customer base, and prefer a priced reliability tier over absorbing the cost in engineering. Be prepared to say no to a number and offer a dated roadmap instead.

## The question underneath the question A request for four nines with penalties is not really a request for a percentage. It is a request for an architecture, an operating model and a bounded financial liability. Answering it well means separating those. ## Step 1: what has the service actually delivered? Compute the indicator over the last six to twelve months and look at the **worst month**, not the mean. A contract is evaluated per period, so an annual average of 99.995% containing one month at 99.7% is a service that would have paid credits. Averages hide exactly the events that trigger remedies, and "we averaged five nines last year" is the most common weak argument in this negotiation. If the history does not exist — a new service, or one that has just been re-architected — say so. Committing a number you have never measured is guessing with the company's money. ## Step 2: compose the dependency chain Availability composes multiplicatively across serial dependencies. A service that must call three components in sequence, each with 99.99% availability and no redundancy between them, is bounded at roughly: ``` 0.9999 ** 3 = 0.99970 -> about 99.97% ``` adding roughly 13 minutes of unavailability a month before your own code has failed once. The ways out are structural rather than motivational: remove a dependency from the critical path, make it redundant so its failures do not compose, or degrade gracefully when it is down so its failure is not your failure. If none of those is on the roadmap, the ceiling is real and the number cannot be promised. ## Step 3: price the nine 99.99% over a 30-day window is 4.32 minutes. That is less than the time it takes to page a human, so at four nines the common failure modes must be handled with no person involved: health-checked automatic failover, capacity in more than one zone or region sized to absorb the loss of one, rollbacks that trigger on their own signals, and a change process that cannot take the whole fleet at once. Every one of those is a funded project. The honest engineering answer to "can we promise four nines?" is usually "yes, after this list of work, and here is what it costs" — which converts an argument about a number into a decision about a budget, where it belongs. ## Step 4: negotiate what the number means Two contracts saying 99.99% can differ by an order of magnitude in difficulty. The clauses that decide it: - **Whose measurement counts.** Your telemetry, the customer's, or a third-party checker. Client-side measurement includes network and DNS failures you do not control. - **What counts as downtime.** Per-request error rate above a threshold, or only sustained full unavailability? Consecutive-minute definitions exclude short spikes entirely. - **Scope.** Fleet-wide, per-region, or per-tenant. A per-tenant definition is far harder, since one customer's bad shard is a breach. - **Exclusions.** Announced maintenance windows, customer misconfiguration, exceeding documented limits, third-party and force-majeure events. - **Claim mechanics.** Whether credits are automatic or must be requested within a window. ## Step 5: bound the downside Remedies should be **capped service credits** — a tiered percentage of the fee for the affected period — rather than open-ended damages or a right to consequential losses. The purpose is to keep the maximum exposure proportionate to what the contract earns, so a bad month costs margin rather than the business. Uncapped or consequential-damage language is the clause to refuse regardless of how confident you feel about the architecture. ## Step 6: remember it is a precedent One bespoke SLA rarely stays bespoke. The next large customer's procurement team will ask for the same terms, and the sales organisation will find it hard to explain why one customer has them and another cannot. The cleaner structure is a **tiered offering**: a standard plan at 99.9%, and a premium or dedicated tier at 99.99% priced to fund the redundancy it requires. That turns reliability into something the customer buys rather than something engineering absorbs, and it keeps the answer consistent across the customer base. ## The answer that sounds reasonable and is wrong "We're already at four nines, so let's sign it." It ignores the worst month, ignores the dependency ceiling, ignores that measurement definitions have not been agreed, and turns an unbounded promise into a precedent. Being right about the number and wrong about all four is how a company ends up paying credits on a service that is working roughly as well as it always did.

  • The customer insists their own monitoring must be the measurement of record. How do you respond?
    Push back on scope rather than on the principle. Client-side measurement includes their network, their DNS and their own client code, none of which you control, so a promise measured there is a promise about the internet. A workable compromise is an agreed third-party checking service, or your telemetry with a defined reconciliation process when the two disagree by more than a stated margin.
  • The architecture cannot support four nines this year. What do you actually offer?
    Offer 99.9% now with the measurement definitions written tightly and honestly, plus a roadmap that names the specific work — redundancy, automated failover, deployment isolation — and a date at which a stronger tier becomes available, ideally as a priced premium plan. A staged commitment you can hold is worth more to both sides than a number that generates credits and a strained relationship.
  • Sales asks why the SLA cannot simply match the internal SLO the team already hits.
    Because the internal target is where the team starts reacting, and a contract set at the same point leaves zero runway before money is owed. The gap is also insurance against measurement disagreement and against a bad month that the internal review absorbs and a contract would not. Matching them converts every internal warning into a customer entitlement.

saying these in an interview costs you the question

  • Quoting an annual average as evidence the target is achievable
  • Ignoring that serial dependencies compose multiplicatively
  • Agreeing a percentage before agreeing how downtime is measured
  • Accepting uncapped or consequential damages as the remedy
  • Treating a bespoke customer SLA as a one-off with no precedent effect

context