A nightly database backup job has reported success every night for six months. What does that success signal actually prove, and what would you do before trusting those backups in a real outage?
answer
- green job = producer, not artifact
- untested backup = Schrodinger's backup
- restore from the off-site copy
- start the engine, then assert on data
- time the drill; that is your real recovery time
basics
~20 sIt proves the job ran and wrote bytes somewhere. It does not prove the backup is complete, readable, or restorable. The only proof is restoring it onto a separate instance, starting the database, and querying the data.
solid answer
~50 sA green backup job is a statement about the **producer**, not the **artifact**. It means a process exited zero. It does not mean the file is complete, that it is uncorrupted, that the encryption key still exists, that the restore procedure still works, or that a restore finishes inside the time budget. Before trusting it I run a restore drill: fetch the backup from the same place a real recovery would fetch it (the off-site copy, not a local scratch directory), restore onto a clean isolated instance, start the database, and run application-level checks - row counts on the biggest tables, the newest timestamp in an append-only table, object counts versus production, and a couple of known business queries. The drill has to be timed and written down: how long it took, how much recent data was missing, and every manual step someone had to improvise. Those numbers are the real recovery capability; the runbook's numbers are only a claim.
go deeper
Say plainly that a successful job only proves the job ran, and that the real test is restoring the backup onto a separate machine and checking the data is there and current.
Add the mechanics: restore from the off-site copy, start the engine, run row-count and freshness assertions, and time the whole thing. Name concrete silent-failure modes such as backing up a stalled replica or a missing schema.
Frame it as evidence generation: cheap verification on every backup, automated restore-and-assert on a schedule, periodic full game days with a rotating operator, and trended timings that feed the recovery-time objective.
Talk about it as an organisational control - restorability as a measured, reported property of every data store, with ownership, sampling strategy across a fleet, and an escalation path when a cluster's last successful proof-of-restore ages out.
## The job versus the artifact A backup pipeline has two separable things: the **process** that produces a backup, and the **artifact** that must survive to be used. Monitoring almost always watches the process - exit code, log line, scheduler status - because that is trivially instrumentable. The artifact is what you actually need, and nothing about a zero exit code speaks to it. Real-world failures that coexist happily with six months of green jobs: - The dump ran against a replica that stopped replicating in March, so every backup since then is a snapshot of stale data. - Only one schema was included because the tool's include-list was never updated when a new schema was added. - The archive is written to a volume that silently filled; the tool truncated and still exited zero (or the exit code was swallowed by a shell pipeline). - Bit rot or a bad disk corrupted blocks in the middle of the archive; nothing reads those blocks until a restore does. - The backups are encrypted with a key that lives only on the database host being backed up - the host you just lost. - The restore works but takes 14 hours, and the business promised 1 hour. - Large objects, sequences, or extensions were not captured, so the restored database will not run the application. None of these are exotic. Each is a routine post-incident finding. ## What a restore drill is A restore drill is an end-to-end rehearsal that produces evidence. The essential properties: 1. **Use the real source.** Pull from the off-site or object-store copy through the same credentials and network path a disaster recovery would use. Restoring from a file that is still sitting on the primary proves nothing about the copy that survives the primary. 2. **Restore to an isolated target.** A scratch instance, container, or throwaway cloud host - never anything sharing storage, credentials, or a connection string with production. A restore drill that can overwrite production is a bigger risk than no drill. 3. **Actually start the database.** Recovering files is not recovery. The engine must open the data, replay whatever log it needs, and reach a consistent, writable state. 4. **Validate at the application layer.** The engine starting is a low bar. Check that the newest row in an append-only table is as recent as the backup claims, compare table and index counts against production, verify a few referential-integrity spots, and run two or three queries the application actually issues. 5. **Time everything and record it.** Fetch time, restore time, replay time, validation time, plus every step a human had to figure out live. The gap between the wall-clock total and the target recovery time is your risk. 6. **Have a non-expert drive it.** If only one engineer can execute the runbook, the runbook fails whenever that engineer is asleep or gone. Rotating the operator flushes out undocumented steps faster than any review. ## Cadence and automation An annual drill is theatre. The practical pattern is layered: cheap automated verification on every backup (checksum/manifest validation), an automated restore of at least one cluster per day or per week into a scratch environment with automated data assertions, and a full human-driven game day - including failover, DNS, credentials, and application startup - once or twice a year. Anything you can automate should run unattended and page when it fails, because a drill that requires a person to remember it eventually stops happening. ## What to answer when asked The answer an interviewer wants is short: job success measures the job, not the backup; only a completed restore with data-level assertions and a stopwatch validates a backup; run it regularly, automatically, from the off-site copy, with someone other than the author driving. Then note the second-order point - the drill also produces the only honest measurement of how long recovery takes, which is what the business actually buys.
- How often would you run restore drills, and does the cadence depend on the system?Cheap checks (checksum and manifest verification) run on every backup; an automated restore-and-assert runs at least weekly, and daily for tier-1 systems. A full human game day - failover, credentials, DNS, application startup - happens once or twice a year and after any material change to the backup tooling, storage layout, or schema footprint. Lower-tier systems can drop to quarterly automated restores, but never to zero.
- What evidence should each drill produce, and where does it go?A dated record with: which backup was used, where it was fetched from, wall-clock time for fetch/restore/replay/validation, the data-freshness gap measured against the failure moment, the assertions run and their results, and every manual or undocumented step. Store it where auditors and on-call engineers both look, and trend the timings - a restore that has crept from 20 minutes to 3 hours is a recovery-objective breach nobody declared.
A fire extinguisher with a current inspection sticker. The sticker says someone walked past it, not that it will spray anything when the kitchen is on fire.
saying these in an interview costs you the question
- Treating a successful backup job or a monitoring green tick as proof the backup is restorable.
- Validating by checking that the file exists and is roughly the expected size.
- Doing the drill by restoring the local copy on the database host, which never exercises the off-site path or the credentials a real disaster needs.
- Restoring but never starting the engine or querying the data, so a structurally broken or stale dataset passes.
- Not timing the drill, so the organisation never learns that its recovery-time promise is fiction.