skip to content

You own reliability standards for an organisation where every team has invented its own SLIs and none of them are comparable. How would you standardise SLI definitions across teams, and where would you deliberately let them diverge?

level: principalimportance: nice to knowfreq 28%

answer

  1. standardise shape, not targets
  2. one vantage point buys comparability
  3. compute a default for everyone centrally
  4. the vanity indicator never moves
  5. did it degrade during the last outage?

basics

~20 s

Standardise the shape and the measurement point, not the targets: compute one default availability and latency indicator centrally from a shared ingress for every service, and let teams add journey-specific indicators — freshness, correctness — that only they can define.

solid answer

~50 s

I would separate what must be uniform from what must be local. Uniform: the definition shape, the measurement vantage point, the window, and the classification rules, all computed centrally from ingress or mesh telemetry so that every service gets a baseline availability and latency indicator with no per-team work and comparable semantics. Local: the targets, which depend on what the product needs, and any indicator the platform cannot see — data freshness, correctness, per-tenant journeys, mobile client experience. The failure I would design against is the vanity indicator: a team choosing what is easy to measure or what already reads five nines, so the number never moves and nobody learns anything. The audit is behavioural, not bureaucratic — when a service's indicator is healthy through a week of user complaints, its definition is wrong regardless of how well it is documented. And I would tier: hand-designed journey indicators for the services that matter, the platform default for the long tail.

go deeper

for a junior

Understand that indicators defined independently by each team are not comparable, and that an organisation usually publishes a default definition rather than letting everyone invent one.

for a middle

Be able to name what should be uniform — definition shape, measurement point, classification rules, window — and what should not, namely the target, which follows from what the product needs.

for a senior

Show how you would get coverage without team-by-team work: compute a baseline indicator centrally from shared ingress telemetry, then let teams correct it. Name the classes of indicator the platform structurally cannot see.

for a principal

Own the failure modes and the economics: the vanity indicator that never moves, exclusions accumulating into an unfalsifiable number, cardinality cost, and tiering design effort so the critical journeys get hand-built indicators while the long tail gets the default. Define success as indicators that degraded during real incidents.

## The problem is comparability, not correctness In an estate of a few hundred services, each team's indicator may be individually defensible while the set is useless. One team counts 4xx as failures; another counts only 5xx. One measures inside the process; another at the load balancer. One uses a rolling 28-day window; another a calendar month. Leadership then reads a list of percentages that cannot be compared, summed, or ranked, and the natural conclusion — "just make everyone use the same number" — is only half right. The useful split is between **the parts that must be uniform to be comparable** and **the parts that must be local to be true**. ## Standardise the shape, the point and the rules Four things are worth mandating: 1. **Definition shape.** Every indicator is a ratio of good events to valid events over a stated window, or a proportion-of-time indicator for the data cases. No team reports a raw average or a bare percentile as its service level. 2. **Measurement vantage point.** Pick one — typically the shared ingress or service mesh — and compute from it. This single decision removes most of the incomparability, because it fixes what class of failure is in scope for everyone. 3. **Classification rules.** A published default: 5xx bad, 2xx/3xx good, health checks and internal probes excluded, 4xx out of the availability indicator, shed-under-pressure counted as bad. Teams may deviate, but a deviation is a documented exception with a reason, not a local invention. 4. **The window.** A common window makes numbers arithmetically comparable and makes tooling possible. The leverage here is enormous, because if the ingress or mesh already carries the telemetry, the platform can compute a baseline availability and latency indicator for **every service in the estate on day one**, with zero work from the owning team. That inverts the usual adoption problem: instead of asking three hundred teams to define an indicator, you hand them one and ask them to correct it where it is wrong. ## Do not standardise targets Targets are product decisions. A checkout path and an internal batch reporting API do not deserve the same reliability, and forcing a common target either over-invests in the second or under-protects the first. What you standardise is the *process* for setting one: derived from what the consuming journey needs, reviewed, and revisited when the service's role changes. ## Where divergence is correct The central pipeline is structurally blind to several things, and pretending otherwise is worse than admitting it: - **Data and async systems.** Freshness, coverage and correctness cannot be derived from ingress traffic. These teams must define their own, and the standard should give them a template rather than a metric. - **Correctness inside a 200.** A service returning a degraded fallback with a success status looks perfect from the outside. Only the team knows what "right" means. - **Client-side experience.** Mobile and browser experience includes hops the platform never sees; teams shipping clients need their own indicator. - **Per-tenant obligations.** A multi-tenant platform with per-customer commitments needs per-tenant indicators; a fleet-wide ratio can be healthy while a named account is down. ## The failure modes to design against **The vanity indicator.** The strongest gravitational pull in this exercise is toward the metric that is easy to compute and already reads 99.99%. It never moves, so it never causes an argument, so nobody notices it teaches nothing. The detection is empirical: an indicator that has not left its healthy band in six months, across incidents that generated user complaints, is not measuring the service. **Definition drift by exclusion.** Each individual exclusion is defensible; the accumulation makes the indicator unfalsifiable. Require excluded volume to be visible next to the indicator, and treat a growing exclusion as a change that needs review. **Cardinality and cost.** Per-journey, per-region, per-tenant indicators multiply. An estate that defines indicators per endpoint per tenant can generate a metric bill larger than the services it monitors. Budget for it explicitly: a fixed number of indicators per service tier, with additions requiring a case. **Standardising the wrong layer.** Mandating a single indicator *definition* for products with genuinely different user journeys produces compliance without meaning — teams report the mandated number and privately watch something else. If you discover a team running a shadow dashboard, the mandated definition is the thing that is broken. ## Tier the effort Not every service deserves the same design investment. A workable structure: - **Tier 1**, the handful of revenue-critical journeys: hand-designed indicators per journey, including correctness, reviewed with the product owner, measured at more than one vantage point. - **Tier 2**: the platform default plus any journey indicator the team argues for. - **Tier 3**, the long tail of internal services: the platform default only, with no per-team work at all, on the honest grounds that an indicator nobody reads is not worth a team-week to design. ## How you know it worked The test is not adoption percentage. It is whether the indicators moved during the last several incidents — whether the number a team watches degraded when its users suffered. Run that check retrospectively after each significant incident: if the service level looked fine throughout, the definition is the finding, and fixing it matters more than any dashboard rollout.

  • How do you detect that a team's SLI is not measuring their service, without auditing every definition?
    Correlate retrospectively. After each significant incident, check whether the service's indicator actually degraded during it; and flag any indicator that has never left its healthy band over several months of real incidents. Both checks are automatable and behavioural, so they surface broken definitions without a review board reading three hundred documents.
  • A team insists their service needs a per-tenant SLI for two thousand tenants. How do you respond?
    Ask what decision changes per tenant. If there are per-customer commitments, per-tenant indicators are legitimate but should be tiered — full indicators for the accounts with contractual obligations, a worst-tenant or proportion-of-healthy-tenants roll-up for the rest. That keeps the failure of a single named account visible without paying for two thousand series across every journey.
  • What is the risk of mandating one universal indicator definition for every product in the company?
    Compliance without meaning. Teams report the mandated number and quietly watch something else that reflects their users, so you have created reporting overhead and learned nothing. The mandate should cover the shape, the vantage point, the classification defaults and the window — the parts that create comparability — while leaving the journey and the target to the people who know the product.

saying these in an interview costs you the question

  • One universal SLI definition for every team
  • Standardise the targets so services are comparable
  • Adoption percentage proves the programme worked
  • Per-tenant, per-endpoint indicators are free to add
  • A documented definition cannot be a wrong definition

context