SRE Practices
How teams keep production services reliable: SLOs and error budgets, incident response, on-call, postmortems, toil reduction, capacity planning, and chaos engineering. Interviewers for SRE, DevOps, and senior backend roles probe these practices to see whether you can run systems, not just build them.
on this pageshowhide
explore
- Service Levels & Error Budgets22 questions
- SLI vs SLO vs SLA6 questions
- Choosing Good SLIs6 questions
- Error Budgets & Budget Policy5 questions
- Burn Rates & SLO Alerting5 questions
- Alerting Philosophy16 questions
- Symptom-Based vs Cause-Based Alerts5 questions
- Actionable Alerts & Paging Discipline5 questions
- Alert Noise Reduction6 questions
- Incident Management22 questions
- Severity Levels & Incident Declaration5 questions
- Incident Command & Roles5 questions
- Incident Communication & Escalation6 questions
- Mitigation-First Response6 questions
- On-Call Practice16 questions
- Rotations & Handoffs6 questions
- Escalation Policies & Paging Chains5 questions
- Pager Load & On-Call Health5 questions
- Postmortems & Learning16 questions
- Blameless Culture5 questions
- Root Causes & Contributing Factors6 questions
- Templates, Action Items & Follow-Through5 questions
- Toil, Runbooks & Automation16 questions
- Defining & Measuring Toil5 questions
- Runbooks & Playbooks6 questions
- Automation Ladder & Self-Healing5 questions
- Capacity Planning & Load Management15 questions
- Demand Forecasting & Provisioning5 questions
- Overload Protection & Load Shedding5 questions
- Load Testing & Saturation Signals5 questions
- Chaos Engineering & Resilience Testing17 questions
- Principles of Chaos Experiments6 questions
- Fault Injection Techniques5 questions
- Game Days & Disaster Recovery Exercises6 questions
- Release Safety17 questions
- Progressive Rollouts & Canary Analysis6 questions
- Feature Flags as Operational Controls5 questions
- Rollback Readiness6 questions
- DevOps / SRE Engineerroleanchors this topic
- MLOps Engineerroleanchors this topic
- PostgreSQL DBAroleanchors this topic
- Software Architectroleanchors this topic
- Backend Developerrole
- DevSecOps Engineerrole
- Forward Deployed Engineerrole
- Full Stack Developerrole
- Java Backend Developerrole
- Java SDETrole
- Kotlin Backend Developerrole
- QA Engineerrole
questions
157 · 9 sectionsYour team proposes CPU utilization and cache hit rate as the SLIs for a user-facing API. Why are those poor choices for a service level indicator, and what test does a candidate SLI have to pass before you adopt it?
basics
~20 sA good SLI tracks what a user experiences — success, latency, freshness — and moves with user happiness. CPU and cache hit rate are internal resource metrics: they spike while users are fine, and stay flat while users suffer.
A service has a 99.9% availability SLO measured over a 30-day rolling window. How much error budget does that give you, and what does it mean for a team to "spend" it?
basics
~20 sAn error budget is one minus the SLO target over the window. A 99.9% target across 30 days allows about 43 minutes of unavailability, or 0.1% of requests, before the objective is missed. Spending it means shipping change against that allowance.
In service reliability practice, define SLI, SLO and SLA, explain how the three relate to one another, and say which one a financial penalty attaches to.
basics
~20 sAn SLI is a measured number, usually good events divided by valid events. An SLO is the internal target that number must meet over a defined window. An SLA is the customer-facing contract, set looser than the SLO, and it is the only one carrying penalties.
A service has a 99.9% availability SLO measured over a rolling 30-day window. Explain what error-budget burn rate means, what a burn rate of exactly 1 signifies, and how you compute it from an observed error ratio.
basics
~20 sBurn rate expresses error-budget consumption as a multiple of the sustainable pace: rate 1 spends the whole budget exactly at the window's end. Compute it as the observed bad-event ratio divided by (1 minus the SLO target).
How would you define a latency SLI for an HTTP API, and why is "average response time under 300 ms" the wrong shape for one?
basics
~20 sDefine latency as the proportion of valid requests served faster than an explicit threshold — for example 99% of checkout requests under 400 ms. An average hides the slow tail entirely: a handful of thirty-second requests barely move it while those users are the ones leaving.
Name the four golden signals from Google's SRE monitoring guidance, and say which of them make good paging alerts for a request-serving service and which one does not.
basics
~20 sLatency, traffic, errors and saturation. Latency and errors are what users feel and are the natural paging signals; a traffic collapse is also user-visible. Saturation measures resource fullness — a cause, so it belongs on dashboards and tickets.
Your team is adding a new alert for a production service and must decide whether it pages a human immediately, files a ticket, or is only recorded for later investigation. What test decides the tier, and what does each of the three tiers commit the team to?
basics
~20 sPage only when a human must act within minutes and can actually change the outcome. Ticket when the work is real but can wait days. Log when nobody will act. Each tier is a response-time promise.
A single database failover causes 200 alert notifications — one per application instance, differing only in the instance label — to hit the pager inside a minute. Explain how alert deduplication and grouping would collapse that into one notification, and what you give up as you group more aggressively.
basics
~20 sDeduplication collapses byte-identical alerts from redundant senders into one. Grouping batches alerts that share a chosen set of labels into a single notification after a short wait window. Grouping harder costs detection delay and can bury an unrelated failure inside an already-acknowledged page.
A service pages the on-call engineer whenever a host's CPU stays above 85% for five minutes, and users have never noticed anything during those pages. In alerting design, what is the difference between a symptom-based and a cause-based alert, and which of the two belongs on the pager?
basics
~20 sSymptom alerts fire on what users experience — failed requests, slow responses, a service level burning down. Cause alerts fire on machine state such as CPU or disk. Page on symptoms; send cause signals to tickets and dashboards.
You join a team whose pager delivers roughly 50 alerts a week, and most are acknowledged with no action taken. Describe the process you would run over the next month to reduce that, and how you would decide the fate of each rule.
basics
~20 sMeasure before changing: per rule, count firings over the last quarter and what share led to a human action. Then apply a disposition ladder — retire, demote off the paging path, retune thresholds and durations, group or suppress the fan-out, or fix the underlying instability — and make the review recurring at every handoff.
Why do incident response practices call for a dedicated channel or bridge per incident rather than discussing the outage in the team's usual chat channel, and what belongs on a voice bridge versus in writing?
basics
~20 sA per-incident channel gives one authoritative place for the response, so anyone joining reads the scrollback instead of interrupting responders, and it becomes the raw material for the postmortem timeline. A voice bridge is faster for fast-moving coordination but leaves no record, so decisions and actions still get typed into the channel.
You are on call. Five minutes after a routine deploy, your service's error rate jumps from 0.1% to 12% and users are seeing failures. What is your first action, and why is "open the logs and find the bug" the wrong one?
basics
~20 sRoll back to the last known-good release first, then investigate. Restoring users is the goal during an incident, the deploy timing is strong enough evidence to act on, and reading logs leaves users broken for however long the debugging takes.
You are handling communications for an ongoing SEV1 outage where customer logins are failing. What goes into each stakeholder update, how often do you send one, and what should you never promise in it?
basics
~20 sIncident updates go out on a fixed cadence tied to severity — commonly every 30 minutes at the top severity — and each one states user-visible impact, what the response is doing now, and the time of the next update. Never promise a restoration ETA.
A major outage pulls a dozen engineers onto the bridge. Your incident process defines an Incident Commander, an Operations (Tech) Lead, a Communications Lead and a Scribe. What does each role own, and why is the Incident Commander explicitly kept out of the debugging?
basics
~20 sThe Incident Commander steers, the Operations Lead is the only one changing production, the Communications Lead handles everyone outside the response, and the Scribe timestamps decisions and actions. The IC stays out of debugging because attention is single-threaded — in a stack trace, nobody is steering.
Your team is writing a SEV1–SEV4 severity matrix for its production services. Which dimensions should decide an incident's severity, and why must the matrix be written in terms of user impact rather than which component broke?
basics
~20 sGrade on observable user impact: what share of users or requests is affected, how critical the blocked journey is, whether a workaround exists, and whether data or money is at risk. Which component failed predicts none of those.
In an on-call paging tool, what does acknowledging a page actually commit you to, and what happens if you acknowledge it and then do nothing?
basics
~20 sAcknowledging a page stops the escalation chain and claims ownership: it tells the system a human is now working the problem. Acknowledging and then going back to sleep is worse than never answering, because nobody else will be notified.
You are configuring the escalation policy for a new customer-facing service in a paging tool such as PagerDuty or Opsgenie. What rungs would you define, how long would you set the acknowledgement timeout, and what must the final rung do?
basics
~20 sPage the service's primary on-call first, escalate to the secondary after an unacknowledged timeout of roughly 5-15 minutes, then to a manager or a named fallback team. The final rung must reach a guaranteed-reachable human and never dead-end.
How do you measure whether an on-call rotation is healthy, and what pager-load numbers would tell you it is not?
basics
~20 sMeasure pager load per shift, not per month: pages per shift, the share that arrive outside business hours, the share that turned out to be actionable, and hours lost to interrupts. Google's SRE book suggests at most two incidents per 12-hour shift.
You are setting up 24/7 on-call for a team that owns a production service. How many engineers does a sustainable rotation need, and what are your options if the team is smaller than that?
basics
~20 sA single-site 24/7 primary rotation needs roughly six to eight engineers, so each person carries the pager about one week in six. With fewer people, narrow what pages overnight, merge the pager with another team, or add a second site.
What has to transfer at an on-call shift handoff, and what goes wrong when the handoff is just a message saying "quiet week"?
basics
~20 sA handoff transfers state, not a summary: open incidents and their next step, every alert silenced during the shift with its expiry, in-flight changes and migrations, temporary manual fixes with an owner, and risky events scheduled in the next shift.
In incident analysis, what is the difference between the trigger of an outage and its underlying cause, and why does the distinction change what you fix?
basics
~20 sThe trigger is the event that started this outage; the underlying cause is the latent condition that made the trigger dangerous. Removing the trigger prevents this instance, fixing the latent condition prevents the whole class — and they cost very different amounts.
What separates a good postmortem action item from a bad one?
basics
~20 sA good action item names one concrete change, has a single named human owner, a priority and a due date, lives in the team's normal tracker, and has an unambiguous definition of done. Bad items are open-ended investigations or exhortations to be careful.
During a postmortem your 5 Whys chain ends at "the config parser didn't validate the field." Why do experienced reviewers push back on stopping there, and what analysis do you do instead?
basics
~20 s5 Whys follows one causal line and stops at the first fixable defect, so it names the bug but not the review, test, rollout and detection gaps that let the bug reach production. Ask "what else" at each step and record multiple contributing factors.
In a blameless postmortem, what does "blameless" actually mean, and how does a team still hold anyone accountable?
basics
~20 sBlameless means the postmortem asks how the system let a reasonable action cause harm, not who to punish. Accountability survives as named owners for the fixes and an obligation to explain honestly; deliberate recklessness is a management matter handled outside the document.
A draft postmortem lists the root cause as "human error — the operator ran the wrong command." Why would an SRE reject that, and what should the document say instead?
basics
~20 s"Human error" is where an investigation stopped, not an explanation. Treat it as a symptom and record what made the wrong command reachable, plausible and undetected — no confirmation prompt, identical staging and production shells, a stale runbook, time pressure — and fix those.
In SRE, what makes operational work count as "toil", and name two kinds of manual work that are not toil?
basics
~20 sToil is manual, repetitive, automatable, tactical work with no enduring value that scales with the service. Manual work failing those tests — a one-off migration, or novel debugging of a new failure — is engineering or overhead, not toil.
A recurring production fix currently runs as a script an on-call engineer executes by hand after being paged. What do you gain, and what do you take on, by promoting it to closed-loop automation that runs with no human in the loop?
basics
~20 sClosing the loop buys speed and consistency: the fix runs in seconds, identically, without waking anyone. In exchange you take on a new production actor that changes systems unsupervised, so it needs guardrails, its own telemetry, and a named owner.
You are writing a runbook for a failure that pages the on-call engineer at 3am. What sections should it contain, and what does each section have to do for someone who did not build the service?
basics
~20 sA usable runbook states its trigger and the user impact, then gives diagnostics with expected output, remediation with preconditions and blast radius, a verification signal that proves the fix worked, an escalation path with a time-box, and metadata naming the owner and last validation date.
A watchdog restarts your API process whenever its health check fails. It has been quietly restarting the process about three times a night for two months; last night the restarts could not keep up and the service was down for 40 minutes. What went wrong in the way that automation was built?
basics
~20 sThe watchdog suppressed the symptom without reporting it, so a slowly worsening fault stayed invisible until it outran the remediation. Auto-remediation must count every action, escalate when the action rate climbs, and refuse to keep repairing indefinitely.
Your team's runbooks have gone stale — commands reference a decommissioned host and several document alerts that no longer exist. How do you make runbook freshness a property of the process rather than a one-off cleanup project?
basics
~20 sTie maintenance to use: whoever follows a runbook during a page repairs it in the same shift. Bind each runbook to the alert that links it so orphans are visible, review them at on-call handoff, stamp owner and last-validated date, and delete more than you write.
In performance testing, what is the difference between a load test, a stress test, and a soak (endurance) test, and what does each one find that the others miss?
basics
~20 sA load test applies expected peak traffic to confirm latency and error targets hold. A stress test pushes past that limit to find where and how the service breaks. A soak test runs moderate load for hours to expose slow leaks.
Your capacity plan has to cover both steady user growth and a marketing launch that goes live next month. How does forecasting organic growth differ from forecasting launch-driven demand, and how do you provision for each?
basics
~20 sOrganic growth is extrapolated from your own traffic history and arrives smoothly. Launch-driven demand has no history to extrapolate, so its size and timing must come from business inputs and be provisioned before the event, not corrected after it.
During a ramp load test you plot achieved throughput and p99 latency against the offered request rate. How do you identify the knee of the latency curve, and what number do you take away as the service's usable limit?
basics
~20 sThe knee is where achieved throughput stops tracking the offered rate and p99 latency starts climbing super-linearly — queues are no longer draining. The usable limit is the highest sustained rate that still meets the latency target, which sits below the knee, not at peak throughput.
Your retail service hits its yearly peak on Black Friday, eight weeks away, and the forecast is roughly 4x your normal daily peak. Which parts of that capacity have lead times you cannot compress, and how would you sequence the eight weeks?
basics
~20 sCompute is usually the easy part. The long poles are cloud quota increases and capacity reservations for specific instance types and zones, stateful work like resharding or index builds, and third-party rate limits. Start those first and leave the final weeks for verification and a freeze.
Your service is over capacity and must reject some fraction of its traffic. Why is shedding a random 10% of requests worse than shedding a deliberately chosen 10%, and what does a system need in place to be able to choose?
basics
~20 sRandom shedding drops checkout traffic and background prefetches at the same rate, so it damages revenue-bearing work to save cheap work. Choosing requires a criticality label attached at the entry point and propagated to every downstream call, with each server shedding lowest tier first.
In SRE practice, what is a game day, and what does rehearsing an outage with the whole team test that runbooks and automated monitoring alone do not?
basics
~20 sA game day is a scheduled exercise where a team induces or simulates a failure and responds to it as if it were a real incident. It tests the human path — detection, paging, runbooks, comms — which no amount of monitoring configuration proves on its own.
In a chaos experiment, what is the steady-state hypothesis, and how do you choose the metric it is stated in?
basics
~20 sThe steady-state hypothesis is a falsifiable claim that a measured, user-visible output of the system stays inside a defined band while a fault is present. State it on output the customer feels, not on internal resource metrics.
You are starting a chaos program for a payment API that runs on Kubernetes across three availability zones and calls twelve downstream services. Which faults would you inject first, and why not start by killing random pods?
basics
~20 sStart with dependency faults — blackhole each downstream one at a time to check that every call you call optional really is — then latency on the busiest ones, then losing one whole zone. Random pod kills come last because rolling deploys and node scale-down already kill pods daily.
Your disaster-recovery plan claims a 30-minute recovery time objective for failing a service over to a secondary region. How would you verify that claim in practice, and what typically turns out to be wrong the first time you actually try it?
basics
~20 sAn untested recovery time objective is an estimate, not a number. Verify it by executing a real failover with a clock running from detection to service restored, in a low-traffic window with a tested way back — then treat the measured time, not the plan's, as the truth.
You are scoping the first production chaos experiment for a payment service. Along which dimensions do you minimize the blast radius, and how do you decide when to widen it?
basics
~20 sMinimize along population, duration, severity and reversibility: smallest slice of traffic or one instance, a hard time cap, the mildest fault that still tests the claim, and an undo. Widen one dimension at a time only after the hypothesis holds.
Your service can only disable a bad new code path by redeploying the previous build, which takes about 25 minutes end to end. What does putting that code path behind a runtime feature flag change about your recovery, and what has to be true of the flag for that improvement to be real?
basics
~20 sA runtime feature flag turns recovery from a redeploy into a configuration change, cutting mitigation from tens of minutes to seconds. That only holds if the flag is evaluated per request, the old path still works, and the change reaches every instance quickly.
You roll a change out to 1% of production traffic, the canary's error rate and latency stay clean for an hour, and then the change breaks one enterprise customer's integration the moment it goes to 100%. Why can a 1% traffic canary contain 0% of the users a change affects, and how would you choose the canary population instead?
basics
~20 sA traffic percentage is a random sample, not a representative one. A rare code path, a single tenant, or one old client version can send so few requests that the canary receives none of them. Choose the canary population deliberately, not just its size.
A service runs from an immutable image identified by digest, but its environment configuration is applied separately and always from the config repository's HEAD. Redeploying the previous image digest did not restore the previous behaviour. Why, and what would you change so that one action reverts the whole release?
basics
~20 sOnly the code was versioned. Configuration applied from a moving HEAD stays at its new value, so half the release is still live after the rollback. Bind the artifact digest and the config revision into one versioned release record you revert as a unit.
A change passed every canary stage, reached 100% of the fleet, and the failure surfaced two days later. Which classes of failure does a progressive rollout structurally fail to catch, and what controls would you add for them?
basics
~20 sProgressive rollouts miss anything that needs time, scale, a specific event, or a specific population: slow leaks and disk fill, dependencies that only saturate at full traffic, monthly batches, and one tenant or old client. Each needs its own control.
A release ships application code together with a database migration that renames a column, and the release turns out to be bad. Why can you no longer simply redeploy the previous version, and how should that migration have been sequenced so the code stays rollback-safe?
basics
~20 sA renamed column leaves rolled-back code reading a column that no longer exists, and the reverse migration destroys data. Sequence it expand-contract: add the new column, dual-write, backfill, switch reads, and drop the old one only after the rollback window closes.