In service reliability practice, define SLI, SLO and SLA, explain how the three relate to one another, and say which one a financial penalty attaches to.
answer
- measure, target, promise
- one of the three costs money
- good events over valid events
- the contract number is the loosest
- a target with no window is not an SLO
basics
~20 sAn SLI is a measured number, usually good events divided by valid events. An SLO is the internal target that number must meet over a defined window. An SLA is the customer-facing contract, set looser than the SLO, and it is the only one carrying penalties.
solid answer
~50 sThe three form a chain from measurement to promise. The **SLI** is the indicator: an actual measurement of service behaviour, conventionally expressed as good events over valid events — for example successful requests divided by all valid requests over the last 28 days. The **SLO** is the objective: a target the SLI must hold, such as 99.9% availability over a rolling 28 days. It is internal, owned by engineering, and it is what the team is actually run against. The **SLA** is the agreement: a contractual promise to a customer, with consequences — normally service credits — when it is missed. Only the SLA has teeth in the legal sense. In a healthy setup the SLA number is deliberately weaker than the SLO number, so the team is already reacting to a missed internal target long before a customer is entitled to anything.
go deeper
Be able to give the three definitions cleanly and in order — measurement, target, contract — and to say that only the SLA carries a penalty. Naming a concrete example of each, such as request success rate, 99.9% over 28 days, and service credits, is what a screener is listening for.
Explain why a complete SLO names an indicator, a threshold and a window, and why a bare percentage is not an SLO. Be ready to say that the SLA number is deliberately weaker than the SLO and to explain what that gap buys the team.
Show that you know the two are measured differently, not just set differently: internal SLIs count every failed request close to the user, while SLA definitions carry exclusions and coarser downtime definitions. Talk about how you would keep a number you cannot hold out of a contract.
Own the asymmetry: SLOs are revisable engineering artifacts, published SLAs are effectively permanent commercial commitments. Be ready to argue why internal services should get SLOs rather than penalty-bearing internal SLAs, and what that does to incident honesty across teams.
## The chain: measure, target, promise These three acronyms are asked together because they describe one pipeline at three levels of commitment: something you measure, something you aim at, and something you owe. **SLI — Service Level Indicator.** A number describing how the service actually behaved. The standard shape is a ratio: ``` SLI = good events / valid events ``` For availability, that is successful responses over all valid requests. For latency, it is requests served faster than a stated threshold over all valid requests — note that a latency SLI is still a ratio, not an average, because "the fraction of requests under 300 ms" is a statement you can hold a target against and "mean latency" is not. The SLI is a fact. It has no opinion about whether the number is acceptable. **SLO — Service Level Objective.** The target the SLI is expected to meet over a stated window. A complete SLO always names three things: the indicator, the threshold, and the window. "99.9% availability" is not an SLO; "99.9% of valid requests succeed over a rolling 28-day window" is. The SLO is internal. Nobody outside the company needs to see it, and the team should be free to tighten it whenever they can afford to. **SLA — Service Level Agreement.** A contract clause promising a customer a level of service, with a defined remedy when it is missed. The remedy is nearly always **service credits** — a percentage of the monthly fee refunded — rather than damages, and it is usually tiered: miss 99.9% and you refund 10%, miss 99.0% and you refund 25%. The SLA also carries definitions the SLO does not need: who measures, what counts as downtime, what is excluded (scheduled maintenance, customer misconfiguration, force majeure), and how the customer must file a claim. ## Why the SLA is looser than the SLO If your SLA promises 99.9% and your SLO also targets 99.9%, then the moment your team notices a problem you are already in breach. The gap between the two is the team's reaction time. A common arrangement is a 99.95% internal SLO behind a 99.9% contractual SLA: engineering treats the tighter number as real and customers are only ever entitled to money after the internal alarm has been ringing for a while. The two are usually also *measured* differently. Internal SLIs tend to be harsh — every failed request counts, measured close to the user. SLA definitions tend to be generous — often only sustained, fleet-wide outages count, with maintenance windows excluded. So even at the same nominal percentage the contractual number is easier to satisfy than the internal one. ## Ownership and direction of travel SLIs and SLOs are engineering artifacts; the team that runs the service owns them and can revise them. SLAs are negotiated by sales and legal with engineering sign-off, and they are effectively one-way: you can always give a customer *more* reliability than promised, but renegotiating a published SLA downward is a commercial event. That asymmetry is why an engineer should never let a number they cannot hold reach a contract. ## Not every service needs all three Internal services normally have SLIs and SLOs and no SLA at all — there is no one to pay credits to, and inventing internal penalties tends to make teams hide incidents rather than fix them. What internal consumers actually need is a published SLO they can design against, so that a caller knows whether it is building on a 99.9% or a 99.99% foundation. The reverse case is worse and common: a company that has signed SLAs but never defined SLOs. It has a contractual promise with nothing measuring it, so the first time it learns of a breach is when a customer files a claim. ## The frequent confusions - Calling the target the SLI ("our SLI is 99.9%") — the SLI is the measurement, the 99.9% is the objective. - Treating the SLA as the engineering goal, which leaves no reaction margin. - Quoting an SLO without a window, which makes it unfalsifiable. - Assuming an SLA breach automatically means an SLO breach, or vice versa — different definitions, different exclusions, different measurement points.
- Can a service have an SLA but no SLO? What goes wrong?Technically yes, and it is a common failure. Without an SLO there is nothing measuring the promise continuously, so the first signal of a breach is a customer claim rather than an internal alert. There is also no margin: the contractual threshold becomes the operational one, leaving the team zero reaction time before money is owed.
- Who should own each of the three inside a company?The team running the service owns the SLI definition and the SLO, because both are engineering statements they must be able to revise. The SLA is owned jointly by sales and legal, but must not be signed without engineering sign-off — it is the only one of the three that cannot easily be walked back once published.
- Should internal platform services publish SLAs to their internal consumers?Usually not. Penalties between internal teams create incentives to under-report incidents and argue about attribution rather than fix things. What internal consumers genuinely need is a published SLO, so a caller can compose its own target knowing whether it is building on a 99.9% or a 99.99% dependency.
The SLI is the speedometer reading, the SLO is the speed you have decided to hold yourself to, and the SLA is the promise you made to the insurer — only the last one costs you money when it is broken.
saying these in an interview costs you the question
- Saying "our SLI is 99.9%" — confusing the measurement with the target
- Claiming the SLA is the number engineers work against
- Stating an SLO with no measurement window
- Believing an SLA breach and an SLO breach always coincide
- Assuming every internal service needs an SLA with penalties