In SRE practice, what is a game day, and what does rehearsing an outage with the whole team test that runbooks and automated monitoring alone do not?
answer
- a fire drill for production
- incidents test people, not just code
- the pager path is never unit-tested
- runbook accuracy under a stranger's hands
- findings, not a clean run
basics
~20 sA game day is a scheduled exercise where a team induces or simulates a failure and responds to it as if it were a real incident. It tests the human path — detection, paging, runbooks, comms — which no amount of monitoring configuration proves on its own.
solid answer
~50 sA game day is a planned rehearsal of an outage: you pick a failure, cause or simulate it in a realistic environment, and let the on-call team respond using the same alerts, runbooks and chat channels they would use at 3am. The point is that an incident tests people and process at least as hard as it tests software. Monitoring config proves a rule exists, not that the page reaches a human who knows they are on point. A runbook proves someone wrote steps down, not that a person who did not write them can execute them under pressure with the console they actually have. A game day is the only cheap way to find out that the alert never fired, the escalation went to someone who left, the runbook's first command needs an access role nobody on call holds, or that three people all assumed someone else was updating the status page.
go deeper
Be able to define a game day in one sentence — a rehearsed outage handled like a real one — and name two things it tests that monitoring does not, such as whether the page reaches a human and whether the runbook is still accurate.
Explain the mechanics: who facilitates, who observes, what timeline gets recorded, and why the drill uses the real alerting and comms channels rather than a simulated one. Be ready to describe how a finding turns into an owned action item.
Show judgment about scenario selection and risk: which failures justify a live drill, what safety controls you put around one, and how you keep the exercise from causing the outage it was meant to prevent. Bring a concrete drill you ran and what it exposed.
Own the tradeoff between the cost of drilling and the cost of learning during a real outage, and be able to argue for the engineering time. Talk about how drill findings compete with feature work for capacity, and how you keep a program from degrading into a compliance ritual.
## What a game day is A **game day** is a scheduled exercise in which a team deliberately creates or simulates a failure and then handles it exactly as it would handle a real incident — same alerts, same pager, same incident channel, same runbooks, same escalation path. It usually runs for a fixed block of time (an hour or two), has a named facilitator who introduces the failure and answers questions about the fiction, and ends with a review of what was learned. It is not a code review, a load test, or a design discussion. The distinguishing feature is that **the responders respond**: they get paged, they open the runbook, they declare (or fail to declare) an incident, they post updates. ## What the automated layer already covers, and where it stops Automated testing and monitoring cover the machine half of reliability well. Continuous integration proves the code behaves. A configured alert rule proves that a threshold exists in a config file. A health check proves the platform can tell a dead process from a live one. None of that exercises the chain that actually decides how long an outage lasts: - Does the signal that fires actually correspond to something a human can act on, or does it arrive as one line in a channel nobody watches? - Does the page reach a phone that is on, unmuted, and held by someone who is genuinely available this week? - If the primary does not acknowledge, does anything escalate — and to whom? - Can a responder who did not write the runbook execute it? Do they have the access it assumes? Does the command in step 4 still exist? - Does anyone tell customers, and who decides that? Every one of those links is a piece of *process* that quietly rots between incidents. People leave, tools get replaced, permissions get tightened, a runbook's cloud console gets redesigned. Nothing in CI notices. The only thing that notices is an incident — either a real one, which is expensive and unscheduled, or a rehearsed one, which is cheap and happens on a Tuesday afternoon with senior engineers in the room. ## The decision a game day forces A game day is a deliberate trade: you spend a few engineer-hours, and accept a small amount of controlled risk, in order to discover the gaps at a time you choose rather than at 3am. The choice is *what to rehearse*, and the answer follows from consequence, not from novelty: the failures worth a game day are the ones that are plausible and whose response path is long or rarely used. Losing a region, losing the primary database, losing the identity provider you log into your own tooling with, losing the third-party payment provider. A failure the team handles twice a month does not need a drill — they are already drilling it for free. ## What a game day produces The output is not "we passed". The output is a **timeline plus a list of findings**. A good game day generates observations like: the alert took eleven minutes to fire because the rule averaged over a long window; the second responder was never paged because the escalation policy pointed at a disbanded team; the runbook's failover step required a role only one person holds; nobody updated the status page for forty minutes because it was unclear whose job that was. Those become owned, dated action items, tracked the way real incident actions are. A finding that nobody owns is the normal way a drill's value evaporates. ## Common shapes Game days come in escalating degrees of realism: a **tabletop**, where the team talks through a scenario without touching anything; a **live drill**, where a real fault is introduced in a controlled way; and an **unannounced** drill, where the responders do not know it is coming. Teams typically start at the cheap end and earn their way up. ## What to say in an interview Say that a game day rehearses the *response*, not just the system; name at least two people/process gaps it finds that tooling cannot (paging path, runbook accuracy, comms ownership, unclear roles); and be able to describe one you have participated in, including one finding that surprised you. Interviewers ask this to find out whether you have actually stood in a war room or only read about it.
- How often should a team run a game day, and what drives the cadence?Cadence follows risk and turnover, not the calendar for its own sake. A common rhythm is quarterly per team, plus a drill whenever something material changes: a new region, a new critical dependency, a rewritten failover path, or enough staff turnover that most of the on-call rotation has never executed the plan. If your last three drills produced no findings, the scenarios are too easy, not the cadence too high.
- Is a game day the same thing as chaos engineering?They overlap but are aimed at different targets. Chaos experimentation is primarily about the system: inject a fault, see whether the software's steady state holds. A game day is primarily about the humans and the process around the system: the page, the roles, the runbook, the comms. A game day often uses fault injection as its mechanism, but its success criterion is what the team learned about responding, not just whether the service stayed up.
- How would you run a meaningful game day for a service that has no staging environment good enough to break?Start with a tabletop against the real production topology — you can walk the paging path, read the runbook aloud step by step, and check access without touching anything. Then take the safe live subset: expire a credential in a controlled way, fail over one replica, or blackhole a single non-critical dependency during a low-traffic window with a tested way back. Realism is a ladder, not a switch.
saying these in an interview costs you the question
- We have runbooks and alerts, so a drill adds nothing
- A game day is just a team-building exercise
- The goal is to finish the drill with no problems found
- Only the on-call engineer needs to take part
- Drills are too risky to ever run near production