You are asked to stand up a company-wide disaster-testing program across dozens of teams, in the style of Google's DiRT exercises. How would you structure it, and what keeps it from degrading into theatre?
answer
- central rules, distributed execution
- default participation with recorded exceptions
- count findings, not attendance
- safety owner and a stop word
- unfunded fixes kill the program
basics
~20 sCentralize facilitation, safety rules and the scenario library; leave execution and findings with the owning teams. Measure the program by findings produced and findings closed, never by exercises run — the moment attendance is the metric, teams stage safe rehearsals that prove nothing.
solid answer
~50 sStructure it as a small central function that owns the scenario library, the safety rules, the calendar and the facilitation, while each team executes against its own service and owns the resulting findings. Set non-negotiable safety controls: a named safety owner per exercise with authority to stop it, drill-labelled comms, a global stop word, a rule that a real incident preempts any drill, and frozen windows around peak business events. Participation should be a default with a documented opt-out rather than a mandate nobody can refuse, because a coerced team games the exercise. The thing that turns a program into theatre is metrics: if leadership counts exercises completed, teams will stage a pre-briefed run of something they already handle. Count findings per exercise and the close rate of those findings instead, publish the interesting failures across teams, and make sure the fixes get funded — findings that recur every year are the clearest sign the program is decorative.
go deeper
Know that large organizations run coordinated disaster-testing exercises across many teams, and that they need shared safety rules such as a stop word and clearly labelled drill communications.
Explain the split between a central function owning scenarios, safety and calendar, and teams owning execution and findings, and why a central team running drills on others produces findings nobody fixes.
Show the metric argument: counting exercises produces safe theatre, counting findings and their close rate produces learning. Be ready to describe blast-radius review for a drill on a shared dependency.
Own the whole design: participation policy and its exceptions, how fix capacity is funded against feature work, how unclosed critical findings feed the organization's reliability decisions, and how you keep the program producing new information rather than attendance.
## What the program is for A company-wide disaster-testing program exists to answer one question at organizational scale: **when a serious failure happens, will the response work?** Google's DiRT (Disaster Recovery Testing) exercises are the best-known published example — a coordinated, company-wide programme of deliberately induced failures, described in Google's SRE writing. Treat its specific cadence and scale as Google's, not as an industry standard; what generalizes is the structure, not the calendar. At one team, drills are a practice. At fifty teams, they are a **program**, and programs fail in ways a single team's practice does not: they get captured by metrics, they get owned by one enthusiast who then leaves, and they gradually select for scenarios that are safe rather than likely. ## Structure: central where it must be, distributed where it should be **Central** owns the things that must be uniform: - The **safety rules**, which are not negotiable per team. - The **scenario library**, graded by risk and by what each scenario tests, so a team can pick something appropriate instead of inventing from scratch. - The **calendar and freeze windows**, so twelve teams do not drill simultaneously into the same shared dependency, and nobody drills during the peak sales period. - **Facilitation** for cross-team scenarios — anything spanning multiple services needs someone whose loyalty is to the exercise, not to one participant. - **Aggregate reporting**: what the program has found, and what remains open. **Distributed** stays with the teams: - Choosing which scenarios matter for their service, because they know their failure modes. - Running the exercise and taking the findings. - Owning and closing the action items, in their own backlog, funded from their own capacity. Centralizing execution is a common and fatal design: a central team running drills *on* other teams produces defensiveness, and the findings never get fixed because they belong to nobody who can fix them. ## Participation without coercion A hard mandate produces compliance behaviour: pre-briefed teams staging a rehearsed run of a failure they already handle weekly. The better design is a **default with a documented exception**: every team of a given tier participates each cycle, a team may decline a specific scenario, and the decline is written down with a reason. That record is itself valuable — a team that declines the same scenario three cycles running is telling you something real about its confidence in that path, and it surfaces as a reliability risk rather than disappearing. ## Safety, at scale Everything a single drill needs, a program needs formalized: - A **named safety owner** per exercise, with unambiguous authority to stop it, distinct from the facilitator. - A **stop word** everyone knows, whose use is never treated as failure. - **Drill-labelled comms** in every channel, so a genuine incident during a drill is instantly distinguishable. - **Real incidents preempt drills**, always, immediately. - **Blast-radius review** before anything touching a shared dependency, because the payments team's drill is the checkout team's outage. - **Freeze windows** for peak business events and major launches. - A rule that if a drill causes real customer impact, it is written up as a real incident, publicly. Hiding one drill-caused outage ends the program's credibility permanently. ## The metrics that decide whether it is real This is the crux. **What you count is what you will get.** - Counting *exercises completed* gets you exercises completed — safe, pre-briefed, uninformative. - Counting **findings per exercise** rewards ambitious scenarios and honest observation. - Counting **close rate and age of findings** is the one that has teeth, because it asks whether anything changed. - Tracking **repeat findings** — the same gap appearing across cycles or across teams — points at systemic problems: a platform-level paging defect, a runbook tool nobody can use, an access model that blocks responders. - Tracking **measured recovery times** for critical paths against their published objectives converts the program's output into a number the business already understands. A program with a 95% participation rate and a 20% finding-close rate is theatre with good attendance. ## Funding the fixes The most common way these programs die is not that drills stop; it is that drills continue while nothing gets fixed. Findings compete with feature work, and without an explicit allocation they lose every quarter. Two levers help: tie unclosed critical findings to the same reliability decision-making the organization already uses — a service with a serious open drill finding on its recovery path is carrying reliability risk, and that should influence what it is allowed to ship — and give the fixes named owners with dates in the team's normal backlog rather than a program-owned list. ## Sustaining it Rotate scenario ownership and facilitation so the program is not one champion's hobby. Publish the good failures widely, anonymized where necessary — a team learning that another team's standby was undersized will check their own. Retire scenarios that have stopped producing findings and add new ones as the architecture changes. And measure the program itself the way you would measure a service: if it produced no new information this cycle, that is a defect in the program, not a clean bill of health for the company.
- Leadership wants a single number to report the program's health. What do you give them?Give the close rate and age of critical findings, not the participation rate. Participation counts attendance and quietly rewards teams for choosing trivial scenarios; close rate asks whether the organization actually changed anything it learned. If they want a second number, use measured recovery times for tier-one services against their published objectives — that converts drill output into a business-legible reliability statement.
- A team's drill on a shared payments dependency risks affecting six other teams. How do you handle it?Blast-radius review before the calendar entry is confirmed. Identify the dependents, decide whether the exercise can be scoped to a slice of traffic or a single consumer, notify the owners of affected services so they can staff and can distinguish drill from reality, and give the exercise a safety owner drawn from outside the running team. If the scenario genuinely cannot be contained, it becomes a scheduled multi-team exercise with central facilitation, not a single team's drill.
- How do you tell the difference between a program that is mature and one that has quietly become a ritual?Look at whether it still produces new information. A mature program shows a mix of scenario ages, findings that surprise people, and repeat findings decreasing over time. A ritual shows a stable roster of comfortable scenarios, near-perfect completion, few or no findings, and the same gaps reappearing each cycle because nothing was funded. The strongest single tell is a drill that has never once been stopped by a safety owner — it means nothing risky has been attempted.
- What role should a central team play once individual teams are drilling competently on their own?Shift from running exercises to owning the parts individual teams structurally cannot: cross-team and shared-dependency scenarios, the safety rules and freeze calendar, the scenario library, aggregate reporting on open findings, and the systemic patterns visible only across teams — a paging platform defect, an access model that blocks responders, a runbook tool nobody uses. It becomes an enabling function with a small amount of genuine authority over safety and scheduling.
saying these in an interview costs you the question
- Mandate participation and report the completion rate
- A central team should run the drills for everyone
- Findings go in a program backlog outside team planning
- Bigger, riskier drills are always better drills
- Hide a drill-caused outage to protect the program