Your company publishes a 99.9% availability SLA to customers. What availability SLO should the team run to internally, and why should the two numbers not be the same?
answer
- the gap is reaction time
- internal number is the strict one
- same percentage, different definitions
- missing the SLO without owing credits
- a contract number is hard to walk back
basics
~20 sRun a tighter internal target than the contract — commonly 99.95% behind a 99.9% SLA. The gap is the team's reaction margin: internal alarms fire and remediation starts while the customer is still well inside the promised level and owed nothing.
solid answer
~50 sSet the internal SLO strictly tighter than the contractual number — 99.95% internal against a 99.9% SLA is a typical pairing. If they were equal, the team would only ever discover a problem at the moment it became a breach, with no runway to fix anything. The gap absorbs three things: detection lag, repair time, and disagreement between two measurement pipelines that will never agree exactly. The measurement definitions also differ, and the difference favours the contract: SLA clauses typically count only sustained, broad outages and exclude maintenance and customer-caused failures, whereas the internal indicator counts every failed user request. So even at the same nominal percentage the internal target is the harder one. Sizing the gap is the real judgment call: too small and it buys no reaction time, too large and you are funding reliability nobody purchased — and you can never sell a tighter SLA later, because you have no evidence you could hold it.
go deeper
Know that the internal target is the stricter of the two and that the contract number is the looser one. Being able to say the gap exists so the team hears about problems before customers are owed anything is enough here.
Explain what the margin absorbs — detection lag, repair time, and disagreement between measurement pipelines — and give a concrete pairing such as 99.95% internal behind a 99.9% SLA.
Show that the definitions differ, not just the thresholds: sustained-outage clauses, maintenance and customer-fault exclusions, and claim windows all make the contractual number easier to hit. Be ready to explain a month that missed the SLO without owing credits.
Own the sizing argument in both directions and the one-way ratchet on published commitments. Be ready to say why an internal target the team permanently violates is worse than no target, and to hold the line on what reaches a contract.
## Why two numbers rather than one A published SLA is a promise with a price attached. An internal SLO is a decision-making tool. Collapsing them into one number destroys the tool, because the instant it tells you something is wrong you already owe money. The standard arrangement is a **tighter internal target behind a looser external one**: 99.95% internally against 99.9% contractually, or 99.9% internally against 99.5% contractually. The distance between them is a margin, and it is worth being specific about what it buys. **Detection and repair runway.** Between the first failed request and a fix being live there is time to notice, page, diagnose and act. The margin is the reliability you are willing to spend on that runway before the customer's entitlement begins. **Measurement disagreement.** Your indicator and the customer's perception are produced by different pipelines with different vantage points. Two honest measurements of the same hour will differ, and you do not want a contractual verdict to hinge on which side of a rounding boundary your collector landed. **Room to be wrong about the target.** A new service's real reliability is not well known. The margin means an over-optimistic internal target costs an uncomfortable review rather than a refund. ## The definitions differ too, and they differ in your favour A point that separates candidates who have written an SLA from those who have only read about them: the same percentage means different things in the two documents. Internal SLIs are typically harsh — every failed user request counts, measured as close to the user as the team can manage, across every region and every customer. SLA definitions are typically generous, and each clause is negotiated: - Downtime often has to be *sustained* — a defined number of consecutive minutes — so brief error spikes never register. - Impact often has to be broad, affecting the whole service rather than one tenant or one region. - Scheduled maintenance announced in advance is excluded. - Failures caused by the customer's own configuration or by their exceeding documented limits are excluded. - Third-party and force-majeure events are excluded. - The customer usually has to *file a claim* within a window to receive anything. The practical consequence: it is entirely normal to miss a 99.95% internal SLO in a month while comfortably satisfying a 99.9% SLA. That is the margin doing its job, and it is the right answer to "how can you be missing your target and still not owe anything?" ## Sizing the gap This is the decision with a cost, and it should be argued from the service, not from convention. **Too narrow.** If the internal target is 99.92% behind a 99.9% SLA, the warning arrives minutes before the breach. There is no time to act, so the margin is decorative. **Too wide.** If the internal target is 99.99% behind a 99.9% SLA, engineering is funding two nines of redundancy that nothing commercial requires. Worse, the team lives in permanent violation of its own target, which is how an SLO becomes something people learn to ignore — and an ignored target cannot inform any decision. Useful inputs to the sizing: how long detection plus mitigation actually takes for this service based on past incidents; how much of the difference between the two nominal numbers is already given back by the SLA's looser definitions; and whether the architecture can even distinguish the two levels, since a single-zone service cannot meaningfully hold 99.99% no matter what number is written down. ## The ratchet only turns one way An SLO can be tightened as soon as the service can hold it, and loosened with an internal discussion. A published SLA is close to permanent: raising it is a commercial commitment that is hard to withdraw, and lowering it is a conversation with every customer who signed the old one. This asymmetry is the reason the engineering owner must be in the room before a number reaches a contract. The rule to state plainly: never commit externally to a level you are not already comfortably beating internally, measured over months and judged on the worst of them rather than the average.
- A service met its 99.9% SLA every month last quarter but missed its 99.95% internal SLO twice. How do you read that?As the design working. The internal target fired, the team investigated, and no customer became entitled to anything. It is only a problem if the internal misses are becoming routine, which suggests either the target is set above what the architecture can hold or reliability is genuinely eroding. Look at whether the misses share a cause before touching either number.
- Can the gap between the SLO and the SLA ever be zero in practice?Effectively it can be, and often is, because the SLA's looser definitions already create a margin even when the two percentages match. A 99.9% SLA that only counts sustained fleet-wide outages and excludes maintenance is a materially weaker promise than a 99.9% internal target counting every failed request. Relying on that implicit margin is risky though, since it disappears the moment someone tightens the contractual definitions.
- Marketing wants to advertise the internal SLO because it is a better number. What is your objection?Publishing it turns a revisable engineering artifact into a commitment. The whole value of the internal target is that it can be tightened experimentally or relaxed after an honest review; once customers have seen it, both moves become commercial events. If a stronger public number is genuinely wanted, negotiate a stronger SLA deliberately and re-size the internal margin above it.
saying these in an interview costs you the question
- Setting the internal SLO equal to the contractual SLA number
- Assuming an SLA breach and an SLO breach mean the same thing
- Believing the SLA is stricter because customers are paying
- Treating a missed internal target as a customer-facing incident by default
- Publishing an SLA number based on an average month rather than the worst