skip to content

Chaos Engineering & Resilience Testing

Deliberately injecting failure to prove systems degrade the way you think they do — steady-state hypotheses, controlled blast radius, and game days. Interviewers raise it to test whether you validate resilience instead of assuming it.

on this pageshow

questions

17

In SRE practice, what is a game day, and what does rehearsing an outage with the whole team test that runbooks and automated monitoring alone do not?

level: juniorimportance: must knowfreq 65%

answer

  1. a fire drill for production
  2. incidents test people, not just code
  3. the pager path is never unit-tested
  4. runbook accuracy under a stranger's hands
  5. findings, not a clean run

basics

~20 s

A game day is a scheduled exercise where a team induces or simulates a failure and responds to it as if it were a real incident. It tests the human path — detection, paging, runbooks, comms — which no amount of monitoring configuration proves on its own.

solid answer

~50 s

A game day is a planned rehearsal of an outage: you pick a failure, cause or simulate it in a realistic environment, and let the on-call team respond using the same alerts, runbooks and chat channels they would use at 3am. The point is that an incident tests people and process at least as hard as it tests software. Monitoring config proves a rule exists, not that the page reaches a human who knows they are on point. A runbook proves someone wrote steps down, not that a person who did not write them can execute them under pressure with the console they actually have. A game day is the only cheap way to find out that the alert never fired, the escalation went to someone who left, the runbook's first command needs an access role nobody on call holds, or that three people all assumed someone else was updating the status page.

go deeper

for a junior

Be able to define a game day in one sentence — a rehearsed outage handled like a real one — and name two things it tests that monitoring does not, such as whether the page reaches a human and whether the runbook is still accurate.

for a middle

Explain the mechanics: who facilitates, who observes, what timeline gets recorded, and why the drill uses the real alerting and comms channels rather than a simulated one. Be ready to describe how a finding turns into an owned action item.

for a senior

Show judgment about scenario selection and risk: which failures justify a live drill, what safety controls you put around one, and how you keep the exercise from causing the outage it was meant to prevent. Bring a concrete drill you ran and what it exposed.

for a principal

Own the tradeoff between the cost of drilling and the cost of learning during a real outage, and be able to argue for the engineering time. Talk about how drill findings compete with feature work for capacity, and how you keep a program from degrading into a compliance ritual.

## What a game day is A **game day** is a scheduled exercise in which a team deliberately creates or simulates a failure and then handles it exactly as it would handle a real incident — same alerts, same pager, same incident channel, same runbooks, same escalation path. It usually runs for a fixed block of time (an hour or two), has a named facilitator who introduces the failure and answers questions about the fiction, and ends with a review of what was learned. It is not a code review, a load test, or a design discussion. The distinguishing feature is that **the responders respond**: they get paged, they open the runbook, they declare (or fail to declare) an incident, they post updates. ## What the automated layer already covers, and where it stops Automated testing and monitoring cover the machine half of reliability well. Continuous integration proves the code behaves. A configured alert rule proves that a threshold exists in a config file. A health check proves the platform can tell a dead process from a live one. None of that exercises the chain that actually decides how long an outage lasts: - Does the signal that fires actually correspond to something a human can act on, or does it arrive as one line in a channel nobody watches? - Does the page reach a phone that is on, unmuted, and held by someone who is genuinely available this week? - If the primary does not acknowledge, does anything escalate — and to whom? - Can a responder who did not write the runbook execute it? Do they have the access it assumes? Does the command in step 4 still exist? - Does anyone tell customers, and who decides that? Every one of those links is a piece of *process* that quietly rots between incidents. People leave, tools get replaced, permissions get tightened, a runbook's cloud console gets redesigned. Nothing in CI notices. The only thing that notices is an incident — either a real one, which is expensive and unscheduled, or a rehearsed one, which is cheap and happens on a Tuesday afternoon with senior engineers in the room. ## The decision a game day forces A game day is a deliberate trade: you spend a few engineer-hours, and accept a small amount of controlled risk, in order to discover the gaps at a time you choose rather than at 3am. The choice is *what to rehearse*, and the answer follows from consequence, not from novelty: the failures worth a game day are the ones that are plausible and whose response path is long or rarely used. Losing a region, losing the primary database, losing the identity provider you log into your own tooling with, losing the third-party payment provider. A failure the team handles twice a month does not need a drill — they are already drilling it for free. ## What a game day produces The output is not "we passed". The output is a **timeline plus a list of findings**. A good game day generates observations like: the alert took eleven minutes to fire because the rule averaged over a long window; the second responder was never paged because the escalation policy pointed at a disbanded team; the runbook's failover step required a role only one person holds; nobody updated the status page for forty minutes because it was unclear whose job that was. Those become owned, dated action items, tracked the way real incident actions are. A finding that nobody owns is the normal way a drill's value evaporates. ## Common shapes Game days come in escalating degrees of realism: a **tabletop**, where the team talks through a scenario without touching anything; a **live drill**, where a real fault is introduced in a controlled way; and an **unannounced** drill, where the responders do not know it is coming. Teams typically start at the cheap end and earn their way up. ## What to say in an interview Say that a game day rehearses the *response*, not just the system; name at least two people/process gaps it finds that tooling cannot (paging path, runbook accuracy, comms ownership, unclear roles); and be able to describe one you have participated in, including one finding that surprised you. Interviewers ask this to find out whether you have actually stood in a war room or only read about it.

  • How often should a team run a game day, and what drives the cadence?
    Cadence follows risk and turnover, not the calendar for its own sake. A common rhythm is quarterly per team, plus a drill whenever something material changes: a new region, a new critical dependency, a rewritten failover path, or enough staff turnover that most of the on-call rotation has never executed the plan. If your last three drills produced no findings, the scenarios are too easy, not the cadence too high.
  • Is a game day the same thing as chaos engineering?
    They overlap but are aimed at different targets. Chaos experimentation is primarily about the system: inject a fault, see whether the software's steady state holds. A game day is primarily about the humans and the process around the system: the page, the roles, the runbook, the comms. A game day often uses fault injection as its mechanism, but its success criterion is what the team learned about responding, not just whether the service stayed up.
  • How would you run a meaningful game day for a service that has no staging environment good enough to break?
    Start with a tabletop against the real production topology — you can walk the paging path, read the runbook aloud step by step, and check access without touching anything. Then take the safe live subset: expire a credential in a controlled way, fail over one replica, or blackhole a single non-critical dependency during a low-traffic window with a tested way back. Realism is a ladder, not a switch.

saying these in an interview costs you the question

  • We have runbooks and alerts, so a drill adds nothing
  • A game day is just a team-building exercise
  • The goal is to finish the drill with no problems found
  • Only the on-call engineer needs to take part
  • Drills are too risky to ever run near production

context

open as a page

In a chaos experiment, what is the steady-state hypothesis, and how do you choose the metric it is stated in?

level: middleimportance: must knowfreq 70%

basics

~20 s

The steady-state hypothesis is a falsifiable claim that a measured, user-visible output of the system stays inside a defined band while a fault is present. State it on output the customer feels, not on internal resource metrics.

open as a page

You are starting a chaos program for a payment API that runs on Kubernetes across three availability zones and calls twelve downstream services. Which faults would you inject first, and why not start by killing random pods?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Start with dependency faults — blackhole each downstream one at a time to check that every call you call optional really is — then latency on the busiest ones, then losing one whole zone. Random pod kills come last because rolling deploys and node scale-down already kill pods daily.

open as a page

Your disaster-recovery plan claims a 30-minute recovery time objective for failing a service over to a secondary region. How would you verify that claim in practice, and what typically turns out to be wrong the first time you actually try it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

An untested recovery time objective is an estimate, not a number. Verify it by executing a real failover with a clock running from detection to service restored, in a low-traffic window with a tested way back — then treat the measured time, not the plan's, as the truth.

open as a page

You are scoping the first production chaos experiment for a payment service. Along which dimensions do you minimize the blast radius, and how do you decide when to widen it?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Minimize along population, duration, severity and reversibility: smallest slice of traffic or one instance, a hard time cap, the mildest fault that still tests the claim, and an undo. Widen one dimension at a time only after the hypothesis holds.

open as a page

Netflix's Chaos Monkey randomly terminates running instances in production. What reliability property does that specifically verify, and which common failure modes does randomly killing an instance never exercise?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Killing a random instance verifies only that losing one replica is a non-event: it drops out of the load balancer, traffic reroutes, capacity is replaced. It never exercises partial failure such as slow dependencies, resource exhaustion, or a whole zone going away.

open as a page

A colleague proposes to "kill a random production server on Friday afternoon and see what breaks" and calls it chaos engineering. What does a real chaos experiment have that this proposal is missing?

level: juniorimportance: should knowfreq 62%

basics

~20 s

A chaos experiment adds four things to breaking something: a measured steady state, a falsifiable hypothesis about it, one fault scoped to the smallest population that can still test it, and pre-agreed abort conditions with a rollback.

open as a page

When injecting a fault into a service's call to a downstream dependency, what is the practical difference between injecting added latency and injecting error responses, and why does latency injection usually uncover more bugs?

level: middleimportance: should knowfreq 50%

basics

~20 s

Error injection returns a failure fast, so it tests the error-handling branch. Latency injection holds each call open, so it consumes threads, connections, and memory across the whole service. Latency finds more bugs because slowness spreads to unrelated work while a fast error stays local.

open as a page

Beyond terminating processes and injecting latency, resource-exhaustion faults deliberately starve a host or container of CPU, memory, disk space, or file descriptors. What does each of those prove that a process kill does not?

level: middleimportance: should knowfreq 40%

basics

~20 s

Each starves a different resource and produces a different degradation. CPU starvation causes queuing and timeouts while the process stays healthy; memory pressure triggers GC thrash or an out-of-memory kill; a full disk breaks writes and logging; exhausted file descriptors block new connections. All keep the process alive and misbehaving.

open as a page

During a game day, what should you be measuring and recording, and how do you tell the difference between a drill that succeeded and one that merely went smoothly?

level: middleimportance: should knowfreq 38%

basics

~20 s

Record a timestamped timeline — fault, first signal, page, acknowledgement, declaration, mitigation, recovery — plus every moment someone was blocked. A drill succeeds when it produces owned findings; a smooth run with zero findings usually means the scenario was too easy or nobody was watching properly.

open as a page

Your team needs to validate its response to a regional database failover. What is the difference between rehearsing that as a tabletop exercise and running it as a live drill, and how would you decide which to run first?

level: middleimportance: should knowfreq 50%

basics

~20 s

A tabletop walks the scenario verbally, costs an hour and risks nothing, but proves only that people know the plan. A live drill actually performs the failover, and it is the only one that proves the tooling, access and timings still work.

open as a page

When designing a game day, what changes if you run it unannounced — no advance notice to the on-call responders — versus announcing the scenario and time in advance, and which would you choose for a team that has never drilled before?

level: seniorimportance: should knowfreq 42%

basics

~20 s

An announced drill teaches the response and validates the plan with everyone available. Only an unannounced drill tests detection, paging and mobilization, because an expectant team is already watching. A team that has never drilled gets the announced version first.

open as a page

What makes a chaos experiment's abort condition a real control rather than a formality, and what belongs in its rollback plan?

level: seniorimportance: should knowfreq 48%

basics

~20 s

An abort condition is real when it is a numeric threshold agreed before the run and wired to an automatic halt, backed by a one-action rollback that was tested first and works even if the fault broke the path you would normally use to undo it.

open as a page

Chaos engineering's principles call for running experiments in production. What does chaos testing only in a staging environment fail to catch, and when is staying out of production the right call?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Staging lacks production's traffic volume and mix, data size, real dependency versions, scaling settings and — above all — its configuration, which is where most failures live. A green staging run is evidence about staging. Stay out of production when the fault is irreversible or you cannot yet measure harm.

open as a page

To make a downstream dependency fail during a chaos experiment, you can inject the fault in the caller's own client code, at a sidecar or mesh proxy, or at the host network layer with firewall or traffic-control rules. How do you choose, and what does each level change about what the experiment can prove?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Choose the lowest layer that can produce the fault your hypothesis is about, and the highest that still lets you target and revoke it safely. Client-code injection targets precisely but skips DNS, connect, and TLS. Proxy injection is realistic on the wire and centrally revocable. Network rules are the most faithful and the hardest to undo.

open as a page

You are asked to stand up a company-wide disaster-testing program across dozens of teams, in the style of Google's DiRT exercises. How would you structure it, and what keeps it from degrading into theatre?

level: principalimportance: nice to knowfreq 25%

basics

~20 s

Centralize facilitation, safety rules and the scenario library; leave execution and findings with the owning teams. Measure the program by findings produced and findings closed, never by exercises run — the moment attendance is the metric, teams stage safe rehearsals that prove nothing.

open as a page

When does a one-off chaos experiment deserve to be automated into one that runs continuously, and what does that automation cost you?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Automate when the property being tested is one that silently regresses as the system changes, the hypothesis already passes at the target scope, and the abort path runs without a human. The cost is a new production-critical system with its own on-call, maintenance and error-budget spend.

open as a page