During a game day, what should you be measuring and recording, and how do you tell the difference between a drill that succeeded and one that merely went smoothly?
answer
- responders cannot take their own notes
- timestamps, not recollections
- record friction, not just outcome
- zero findings means too easy
- re-run the scenario to prove the fix
basics
~20 sRecord a timestamped timeline — fault, first signal, page, acknowledgement, declaration, mitigation, recovery — plus every moment someone was blocked. A drill succeeds when it produces owned findings; a smooth run with zero findings usually means the scenario was too easy or nobody was watching properly.
solid answer
~50 sAssign a dedicated observer or scribe, because responders cannot take notes while responding. What the observer records is a timeline with real timestamps — fault injected, first signal visible, alert fired, human acknowledged, incident declared, mitigation started, service recovered — which gives you measured detection, acknowledgement and mitigation times rather than remembered ones. Alongside that, they record friction: every time someone asked who owns this, opened a runbook and found it wrong, lacked an access role, or posted in the wrong channel. Those are the findings, and findings are the product. A drill where nothing went wrong has told you either that the scenario was too easy or that observation was too shallow — so the response is to harden the next scenario, not to celebrate. The findings then become owned, dated items alongside real incident actions, and you re-run the same scenario later to prove the fixes work.
go deeper
Know that someone other than the responders must take notes, and that a drill's value is the list of problems it finds rather than whether the team recovered quickly.
Be able to list the timeline points you record and the times you derive from them — detect, acknowledge, mitigate — and explain why a drill with no findings signals an easy scenario rather than a healthy system.
Show how you use the numbers: comparing drill timings to real incident timings, baselining a scenario so a re-run can prove the fixes, and routing findings by category to the owners who can actually close them.
Own the measurement policy across teams: what makes drill results comparable, how findings compete with feature work for capacity, and how you prevent the program being scored on smooth runs, which quietly incentivizes trivial scenarios.
## Somebody has to be watching The first rule of instrumenting a game day is that **responders cannot observe themselves**. A person diagnosing a failure is fully loaded; they will not note that it took nine minutes for the second engineer to be paged, and afterwards they will remember it as "a couple of minutes". Every drill needs at least one person whose only job is to watch and write, ideally two for a large exercise: one on the timeline, one on the human friction. The observer does not help. The moment they answer "the runbook is in the other wiki", they have destroyed the finding they were there to record. ## The timeline Record wall-clock timestamps for a fixed set of moments, so drills are comparable to each other and to real incidents: - fault introduced - first signal observable in telemetry - alert fired - page delivered - human acknowledged - incident declared, and at what severity - roles assigned - first mitigation action attempted - customer-facing impact ended - full recovery From those you derive the numbers worth arguing about: **time to detect** (fault to alert), **time to acknowledge** (page to human), and **time to mitigate** (acknowledgement to impact ended). Compare them against the same measures from your real incidents — a drill whose detection time is far better than your real incidents' usually means the drill was announced, and you should read the number accordingly. ## The friction log The timeline tells you how long. The friction log tells you why. Record verbatim, with the time: - "Who owns the payments dashboard?" — nobody answered for four minutes. - Runbook step 3 references a console page that has been redesigned. - Responder lacked the role required to run the mitigation and had to find someone with it. - Two people both started the same mitigation, unaware of each other. - The status page was never updated; when asked, three people each thought another owned it. - The alert that fired named a cause rather than the user-visible symptom, so the first ten minutes were spent on the wrong subsystem. Each of those is a defect in the response system, and each is invisible in a chat transcript read a week later. ## What success actually means The temptation is to score a drill by whether the team recovered. That is the wrong scoreboard: it rewards easy scenarios and punishes honest ones. **The product of a drill is findings.** Judge it by: - **How many findings, and how severe.** Zero findings is a red flag, not a gold star. It means the scenario stayed inside what the team already does routinely, or the observation was too shallow to catch the friction. - **Whether the numbers moved.** If the same scenario was drilled two quarters ago with a 22-minute detection time and it is now 6, the fixes worked. That comparison only exists if you recorded numbers the first time. - **Whether the previous drill's findings are actually fixed.** Re-running an old scenario is the cheapest verification you will ever get, and the most uncomfortable. If a drill genuinely produces nothing, the correct response is to make the next scenario harder — remove the person who always knows the answer, combine two failures, run it during a deploy, or move it to a shift that has never handled it. ## Turning findings into change A finding that leaves the room as a shared feeling is worth nothing. Each one leaves with a named owner and a date, tracked in the same place real incident actions are tracked so it competes for the same capacity rather than living in a drill-only backlog that nobody reads. Findings usually sort into a few buckets — a missing or badly targeted alert, a runbook defect, an access or tooling gap, an unclear role or decision authority, and a genuine system weakness — and the buckets matter because they route to different owners. The closing move is the one most teams skip: **schedule the re-run**. The fix for "the escalation target was a disbanded team" is itself untested until a drill pages it. ## What to say in an interview Name the dedicated observer, list the timeline points you record, explain the derived detection and acknowledgement times, and then make the counter-intuitive point clearly: a drill with no findings is a failed drill. That last sentence is what tells an interviewer you have actually run one.
- Why should a game day's findings go into the same tracker as real incident action items rather than a drill-specific list?Because a separate list is a backlog nobody prioritizes. Drill findings describe the same defects real incidents would surface — missing alerts, stale runbooks, unclear ownership — and they should compete for the same engineering capacity on the same terms. A dedicated drill backlog also quietly signals that these are hypothetical problems, which is exactly the belief that lets them survive until they are real ones.
- Your observer's timeline shows detection took eleven minutes. What do you do with that number?Treat it as a finding with a cause, not a score. Ask what happened in those eleven minutes: was there no rule for that symptom at all, did the rule average over a long window, or did it fire quickly into a channel nobody watches? Each cause has a different owner and a different fix. Then record the number so the re-run has something to beat — an unbaselined improvement claim is unfalsifiable.
- How do you record findings about people without making the drill feel like a performance review?Write the system's defect, not the person's. "Only one engineer can execute the failover" is a finding about a single point of human knowledge; "Sam did not know how to fail over" is a finding about Sam and will cost you every honest answer in future drills. The same discipline applies in postmortems, and in a drill it matters more because the failure was scheduled by you.
saying these in an interview costs you the question
- The drill went smoothly, so we're in good shape
- Responders can write up what happened afterwards
- Findings can be tracked in a separate drill backlog
- Skip the timestamps — the chat log has everything
- Zero findings proves the system is resilient