skip to content

Your team needs to validate its response to a regional database failover. What is the difference between rehearsing that as a tabletop exercise and running it as a live drill, and how would you decide which to run first?

level: middleimportance: should knowfreq 50%

answer

  1. talking versus actually doing it
  2. one costs an hour, one costs risk
  3. beliefs versus measured truth
  4. cheap format finds the plan gaps
  5. expensive format finds the rot

basics

~20 s

A tabletop walks the scenario verbally, costs an hour and risks nothing, but proves only that people know the plan. A live drill actually performs the failover, and it is the only one that proves the tooling, access and timings still work.

solid answer

~50 s

A tabletop is a discussion: the facilitator narrates the failure, and responders say out loud what they would do, in what order, and who they would call. Nothing is touched, so it costs an hour and carries no production risk, and it is very good at exposing plan-level gaps — unclear ownership, a missing decision-maker, two people who believe different things about the failover criteria. A live drill actually executes the failover in a real environment with a clock running. It is the only version that catches the things a discussion cannot: the credential that expired, the runbook command that no longer exists, the standby that was scaled down to a tenth of production, the DNS record whose TTL adds ten minutes nobody budgeted for. The usual order is tabletop first, because it is cheap and it surfaces the embarrassing gaps before you spend real risk. You go live once the plan survives being said out loud.

go deeper

for a junior

Know that a tabletop is a spoken walkthrough with nothing touched and a live drill actually performs the failure, and be able to say why a team would start with the cheaper one.

for a middle

Explain what each format catches: plan-level gaps like unclear ownership and decision criteria versus mechanical rot like expired access, undersized standbys and DNS caching. Give the escalating ladder from tabletop to production drill.

for a senior

Demonstrate the risk call: when the fidelity is worth the exposure, what safety controls you put around a production drill, and how you exercise everything up to an irreversible step when a full failover is off the table.

for a principal

Own the policy: which critical paths must be live-drilled on a schedule, who signs off, how drill risk is budgeted against the risk of an untested recovery plan, and how you defend that time when a delivery deadline is pressing.

## Two formats, two different kinds of truth A **tabletop exercise** is a facilitated conversation. Everyone sits in a room (or a call), the facilitator says "at 09:14 the primary database in the main region stops accepting writes," and the participants describe their next actions in turn. The facilitator plays the world: when someone says "I check the replication dashboard," the facilitator says what they see. Nothing in production changes. A **live drill** introduces the real condition — you actually fail the database over to the secondary, or you actually cut the application's path to the primary — and the team responds with the real tools, on the real clock. The distinction matters because they find *categorically different* defects. ## What a tabletop reliably finds Tabletops are excellent at exposing gaps in shared understanding, which are surprisingly common and surprisingly expensive: - **Ownership.** Who decides to fail over? A tabletop makes the silence audible when you ask "who says go?" - **Criteria.** Two engineers state different thresholds for pulling the trigger. In a real incident this becomes twenty minutes of debate. - **Sequencing.** Someone reveals they would drain traffic before promoting the replica; someone else would do the reverse. - **Missing participants.** The plan turns out to require the data team, who are not in the rotation and have no pager. - **Comms.** Nobody has said who writes the customer-facing update or who authorizes it. It is also the only format that is safe for the truly catastrophic scenario. You cannot live-drill "the identity provider we log in with is gone" casually. You can absolutely tabletop it, and you will learn that half the runbook assumes you can already log in. ## What only a live drill finds A tabletop measures what people *believe*. A live drill measures what is *true*. In a real failover the things that break are usually not the parts anyone would have described wrongly in a meeting: - The failover script needs a role that was removed in a permissions cleanup six weeks ago. - The standby region runs at a fraction of primary capacity, so it accepts the traffic and then falls over — the plan said "fail over", it never said "and scale first". - Client-side DNS caching and connection pools keep pointing at the dead endpoint long after the record changed. - Replication lag means the promotion loses the last few seconds of writes, and nobody had a decision path for that. - Caches are cold in the secondary, so the first minutes after recovery look like a second outage. - The actual elapsed time is 70 minutes against a plan that claimed 30. None of that is discoverable by talking. ## The decision, and its cost The honest framing is that this is a **risk-versus-fidelity ladder**, and you climb it: 1. **Tabletop** — no risk, plan-level findings, an hour of calendar. 2. **Live drill in a pre-production or isolated environment** — some fidelity, no customer exposure, but only as truthful as the environment resembles production. 3. **Live drill in production, announced, low-traffic window, with a tested way back** — high fidelity, real risk, real value. Start at the top of the ladder when the team has never drilled this scenario, when the plan is new or has changed materially, or when the failure is one you cannot safely cause. Go straight to a live drill when the plan is already well understood and the open question is mechanical — "we all agree what to do, does it still work?" Running a live drill first on a plan nobody has read aloud is how a drill becomes an incident: you spend real risk to discover a gap a free conversation would have found. The cost of stopping at the tabletop is that you will believe a number you have never measured. A stated recovery objective that has only ever been discussed is a guess, and it will be wrong in the expensive direction. ## Practical notes Give a tabletop the same discipline as a live drill: a facilitator, a scribe, timestamps in the notes, and findings that become owned action items. A tabletop that ends with "good discussion" and no written findings produced nothing. For the live drill, agree in advance on the abort condition and the way back, run it when a senior engineer who knows the system is available rather than when they are on vacation, and prefix every message in the incident channel so that a real incident arriving mid-drill is not mistaken for the fiction. A reasonable pattern for a critical failover path is: tabletop when the plan changes, live drill in production at least once or twice a year, and re-run the same live drill after you fix what the last one found — because the fix is also untested until you exercise it.

  • You cannot safely fail over the real production database. What is the most valuable live element you can still exercise?
    Exercise everything up to the irreversible step. Page the on-call for real, have them assemble, open the runbook, and actually run the read-only and reversible parts: check replication lag, verify access to the failover tooling, confirm the standby's current capacity, dry-run the promotion command where the tool supports it. You measure the detection and mobilization time truthfully, and you find expired access and stale commands, which is most of what a full drill would have caught.
  • How do you keep a tabletop from turning into a design discussion?
    The facilitator enforces the clock and the fiction. Every answer must be an action a named person takes now, and the facilitator responds with what the world does in return. When someone starts redesigning the system, the scribe records it as a finding and the facilitator moves the clock forward. Design work is a valid output of the exercise, but doing it in the room replaces the rehearsal you came for.
  • What should a live drill in production have in place before you start?
    An agreed stop condition and a rehearsed way back, a low-risk window, a named person with authority to call it off, a small circle who knows the drill is happening, drill-prefixed comms so a genuine incident cannot be confused with the exercise, and a rule that a real incident preempts the drill immediately. Without those, you are not drilling, you are gambling.

saying these in an interview costs you the question

  • A tabletop is enough — no need to ever go live
  • Live drills are too risky, so we never run them
  • Tabletop is only a meeting, so it finds nothing
  • Run the live drill first, it's the realistic one
  • A drill with no rollback plan is fine if we're careful

context