skip to content

NFR Design

Turning non-functional requirements into actual architectural decisions rather than aspirations in a document. You will cover scalability, availability, latency, security and compliance budgets, and expressing them as an SLA/SLO/SLI triad that can be measured.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

What is the difference between an SLA, an SLO, and an SLI, and how do the three fit together when you design a system's reliability target?

level: juniorimportance: must knowfreq 85%

answer

  1. I=Indicator (measured)
  2. O=Objective (internal target)
  3. A=Agreement (external, contractual)
  4. SLA looser than SLO
  5. error budget = 1 − SLO

basics

~20 s

SLI is what you actually measure (e.g., % of successful requests). SLO is the internal goal for that measurement (e.g., 99.9%). SLA is the external, often contractual, promise to customers — usually set looser than the SLO so you have margin.

solid answer

~50 s

The triad separates measurement from target from promise. An SLI (Service Level Indicator) is a concrete, measured metric — request success rate, p99 latency, uptime minutes. An SLO (Service Level Objective) is the internal target you set for that SLI, e.g. 'p99 latency < 300ms, 99.9% of the time.' An SLA (Service Level Agreement) is the external, often contractual, commitment made to customers, typically with financial penalties for breach. The SLA should always be looser than the SLO — if your SLO is 99.9% and your SLA promises 99.9%, you have zero margin and will breach the SLA the moment you're merely meeting your internal target. Practically: pick SLIs that reflect user-perceived quality, set SLOs a notch tighter than the SLA, and use the gap (the 'error budget') to decide how much risk — deploys, experiments, chaos testing — you can absorb without breaching either.

go deeper

for a junior

Should be able to state the three definitions correctly and explain, at a high level, that SLA is the external promise and SLO/SLI are internal engineering concepts used to hit it.

for a middle

Should be able to compute an error budget from a given SLO and window (e.g., minutes of allowed downtime per month), and explain why the SLA needs a buffer under the SLO.

for a senior

Should be able to choose appropriate SLIs for a given service (not just default to uptime), justify a specific SLO target against business/cost trade-offs, and describe an error-budget policy that actually changes team behavior.

for a principal

Should be able to design an SLO hierarchy across a multi-service system (differentiated targets by criticality), reconcile conflicting SLA commitments across teams/products, and reason about the cost curve of chasing additional nines at the architecture level.

## The vocabulary at a glance The SLA/SLO/SLI triad is the vocabulary solution architects use to turn a vague reliability aspiration ('the system should be reliable') into something you can measure, target, and contractually promise. Understanding the mechanism requires separating three distinct artifacts that sit at different altitudes. | Artifact | What it is | |---|---| | **SLI** — Service Level Indicator | the lowest-altitude artifact: a directly measured, quantitative signal of the service's behavior | | **SLO** — Service Level Objective | the target value or range you commit an SLI to hit, over a rolling window | | **SLA** — Service Level Agreement | the outward-facing, often contractual commitment made to customers or business partners | ## The indicator you measure An SLI (Service Level Indicator) is the lowest-altitude artifact: a directly measured, quantitative signal of some aspect of the service's behavior, expressed as a ratio or a metric over a time window. Common SLIs: - 'proportion of HTTP requests returning non-5xx within the last 5 minutes,' - 'p99 read latency,' - 'proportion of successful writes.' An SLI is **not a target** — it is a number the monitoring system produces continuously by dividing good events by total events (or measuring a distribution). ## The objective you set, and the budget it implies An SLO (Service Level Objective) is the target value or range you commit an SLI to hit, over a rolling window, e.g., '99.9% of requests succeed, measured over a rolling 30 days.' The SLO is **internal**: it's the number engineering teams design and operate against. It is the anchor from which an **error budget** is derived: `error budget = 1 − SLO` (as a fraction of allowed bad events over the window). A 99.9% SLO over 30 days permits roughly 43 minutes of full downtime (or an equivalent amount of partial degradation) — that 43 minutes is the budget the team can spend on risky deploys, infra maintenance, or absorbing incidents before they're 'out of budget.' ## The promise you sign, and why it sits looser An SLA (Service Level Agreement) is the outward-facing, often contractual, commitment made to customers or business partners, typically with a remedy (service credit, refund, termination right) attached to a breach. Because breaching an SLA has financial and reputational consequences, the SLA number is deliberately set **looser** than the internal SLO — e.g., SLO 99.95%, SLA 99.9%. That buffer is not slack for its own sake; it exists because SLOs get missed occasionally (incidents happen), and you want 'we missed our internal target this month' to not automatically mean 'we owe customers money.' AWS's published EC2 SLA (99.99% for certain configurations) is a well-known public example of this outward commitment; the internal SLOs AWS actually engineers to are typically tighter and are not published. ## What the triad buys you Why this exists: without the triad, teams either over-invest in reliability chasing an undefined 'as reliable as possible' (burning enormous engineering time for diminishing returns past a certain nine), or under-invest and get blindsided by customer-visible outages. The triad turns a fuzzy NFR into an operating model: 1. Pick the SLI that best proxies user happiness. 2. Set an SLO tight enough to leave SLA margin. 3. Use the resulting error budget as a shared currency between product (which wants velocity) and SRE/ops (which wants stability). This is the mechanism popularized by Google's SRE book and now standard practice at most mid-to-large engineering organizations. ## Too tight and too loose Trade-offs run in both directions. - **Too tight.** Setting an SLO too tight (e.g., 99.999%, 'five nines') forces architectural costs that grow non-linearly: multi-region active-active deployments, synchronous cross-region replication or careful eventual-consistency handling, extensive chaos engineering, on-call staffing for sub-minute response — costs that can dwarf the revenue the extra nine protects. - **Too loose.** Setting an SLO too loose erodes trust: teams stop taking the number seriously, incidents become 'normal,' and the SLA is at constant risk of breach because there's no cushion. The real skill is choosing an SLO that matches business impact — a payments API and an internal analytics dashboard should not share a target. ## Where it goes wrong Failure modes in production are recognizable. - Teams that pick the **wrong SLI** (e.g., server-side uptime instead of client-perceived success rate) can show 100% SLO compliance while users experience errors from a broken CDN or client library — the indicator didn't capture what mattered. - Teams that don't leave an SLA/SLO gap breach contractual SLAs from routine operational noise. - Teams that don't operationalize the error budget (no policy for what happens when it's exhausted) end up with the number as decoration — deploys continue at the same pace regardless of budget burn, defeating the entire point. ## A month in the life of a checkout service A worked scenario: a checkout service sets an SLI of 'proportion of checkout requests completing under 2s with a 2xx response.' The SLO is 99.9% over a rolling 28 days (≈40 minutes of budget). The public SLA quoted to enterprise customers is 99.5% with service credits. Mid-month, a bad deploy burns 25 of the 40 minutes of budget in one incident. Per policy, the team halts non-essential deploys and prioritizes a reliability fix until the budget partially replenishes on the rolling window — all without ever coming close to breaching the looser, customer-facing 99.5% SLA.

  • Why should the SLA published to customers be looser than the internal SLO rather than equal to it?
    Because SLOs get missed periodically due to real incidents, and if the SLA equals the SLO, every routine miss becomes a contractual breach with financial penalties. The gap is a deliberate buffer so operational noise doesn't automatically trigger service credits or reputational damage; teams calibrate the gap based on how often they expect to graze the SLO.
  • What is an error budget and how is it used operationally?
    It's the allowed amount of unreliability under the SLO — 1 minus the SLO, expressed as downtime or bad-event minutes over the measurement window. Teams spend it on risk: deploys, experiments, planned maintenance. When it's exhausted, a good error-budget policy pauses feature releases and shifts the team to reliability work until the budget recovers, turning reliability into a shared, numeric negotiation instead of a political one.
  • How do you choose a good SLI for a customer-facing API?
    Measure what the user actually experiences, not what's easiest to instrument server-side — success rate and latency measured as close to the client as possible (e.g., at a CDN/edge or via real-user monitoring), not just origin-server 200s. A bad SLI can read green while users see errors from a layer the SLI doesn't cover, like a CDN or client SDK.
  • Should every service in a system share the same SLO?
    No — SLOs should be set per service based on business impact and what depends on it. A payments write path might warrant 99.95%, while an internal reporting dashboard might be fine at 99.5%; uniform targets either over-engineer low-value paths or under-protect critical ones.

Like a car's speedometer (SLI, the actual measured speed), your personal rule to never exceed 65 in a 70 zone (SLO, tighter internal target), and the legal speed limit of 70 with a ticket if you break it (SLA, external commitment with a penalty) — you keep a buffer between your own rule and the law so an honest mistake doesn't turn into a fine.

saying these in an interview costs you the question

  • Treats SLA and SLO as interchangeable terms
  • Sets the SLA equal to or tighter than the SLO
  • Picks an SLI that measures server uptime instead of user-perceived success
  • Has no policy for what happens when the error budget is exhausted
  • Chases 'as many nines as possible' without tying the target to business impact or cost
  • Cannot explain how the error budget is derived from the SLO

context

open as a page

You're told a service must meet a 99.95% availability NFR. Walk through how you'd translate that number into concrete architectural decisions.

level: middleimportance: must knowfreq 80%

basics

~20 s

99.95% means about 4.4 hours of downtime allowed per year. To hit that you typically need redundancy (multiple instances/zones), automatic failover, health checks, no single point of failure, and enough capacity headroom to survive losing a node without falling over.

open as a page

A business NFR states the system must scale to support 10x current traffic within 12 months without a redesign. What concrete architectural decisions does that requirement drive, and how do you validate the system will actually get there?

level: seniorimportance: must knowfreq 75%

basics

~20 s

It pushes you toward stateless services that can scale horizontally, autoscaling on real metrics, partitioning/sharding data stores so no single DB instance is the ceiling, and caching/async processing to cut load on the bottleneck. You validate it with load testing at the target scale, not by assuming the design works.

open as a page

A product NFR states an API must respond in under 300ms at p99. The request fans out to a database call, an auth check, and two downstream microservices. How do you turn that single latency number into a design constraint across the call chain?

level: middleimportance: should knowfreq 65%

basics

~20 s

Split the 300ms budget across every hop in the request's path (network, auth, each downstream call, your own processing), leaving slack for tail variance, not just averages. If the sum of the pieces plus slack doesn't fit, redesign — e.g., call downstreams in parallel instead of one after another, cache, or drop something from the critical path.

open as a page

A solution architecture must satisfy a security NFR such as 'all sensitive data encrypted at rest and in transit, with defense in depth against a compromised application server.' What architectural decisions does this drive, and what does it cost elsewhere in the design?

level: seniorimportance: should knowfreq 60%

basics

~20 s

It means encrypting data both while stored (at rest) and while moving over the network (in transit, e.g. TLS), and not relying on just one layer of defense — network segmentation, least-privilege access, and secrets management so that if one layer (like the app server) is breached, the attacker still can't reach everything. The cost is extra latency, operational complexity, and key-management overhead.

open as a page

A solution architecture must satisfy a compliance NFR requiring EU customer data to stay within EU data centers and be retained for a fixed audit period, while the product also has a global latency target and a cost ceiling. How do you architect for a compliance NFR that actively conflicts with other NFRs, and how do you decide the trade-off?

level: principalimportance: should knowfreq 45%

basics

~20 s

You typically split data by residency requirement (data-partitioning by region), keeping EU customer data physically in EU infrastructure while allowing non-regulated data or metadata to flow globally. That usually means non-EU users hitting EU-hosted data pay a latency penalty, or you replicate read-only, non-sensitive views elsewhere — the trade-off is decided by ranking compliance as a hard constraint (non-negotiable, legal risk) versus everything else as tunable.

open as a page