skip to content

Describe the layers of verification you can apply to a database backup, from cheapest to most convincing - for example checksum and manifest verification (as pgBackRest's verify command does) versus a full restore with data assertions. What does each layer catch, and what does each miss?

level: middleimportance: should knowfreq 45%

answer

  1. ladder: exists, checksums, restores, asserts
  2. manifest verify = bytes, not semantics
  3. clean verify on a stale replica still passes
  4. only restore yields a restore-time number
  5. assert freshness + object census

basics

~20 s

Four layers: metadata checks (file present, size, expected members), checksum and manifest verification of the stored bytes, a restore that starts the engine and replays logs, and application-level assertions on the restored data. Cheap layers catch corruption; only a restore proves recoverability.

solid answer

~60 s

I think of it as a ladder, each rung more expensive and more convincing. 1. **Inventory checks** - the expected backup exists, has plausible size, and contains the expected set of files. Catches a job that never ran or wrote nothing; misses everything about content. 2. **Checksum / manifest verification** - re-read every stored block and compare against the checksums recorded at backup time (this is what a tool such as pgBackRest's verify does). Catches bit rot, truncation, and a missing file in a repository; misses logical problems - a stale source, a missing schema, an unusable log chain. 3. **Restore and start** - materialise the backup on a scratch instance, replay the log chain, reach a consistent open state. This is the first layer that proves the artifact is a database. It also catches key-management failures and log-chain gaps. 4. **Application assertions** - row counts, freshness of the newest row, object counts versus production, referential spot checks, a few real queries. Cheap layers run on every backup; the expensive ones run on a schedule. Neither substitutes for the other.

code

text · 11 lines
text
verify: full/20260812-010000F
  checked 4812 files / 1.9 TB
  manifest checksums: OK
  wal chain 0000000100000A2B..0000000100000A9F: complete
  result: PASS

not covered by this pass:
  - whether the source instance was replicating at backup time
  - whether every application schema was included
  - whether the encryption key is still retrievable
  - how long a restore of this backup takes

go deeper

for a junior

Know the two ends of the ladder: a checksum check tells you the file is intact, and only an actual restore tells you the database can come back.

for a middle

Lay out all four layers with what each catches and misses, and be able to name a failure that passes verification and fails restore, such as backing up a stalled replica.

for a senior

Design the cadence: cheap layers on every backup, restore-and-assert on a rotation that guarantees every cluster is proven inside its tier window, with timings fed back into the recovery-time objective.

for a principal

Discuss it as a coverage and cost problem across a fleet - sampling strategy, egress and compute budget for verification, and what evidence you present when someone asks whether the estate is recoverable.

## Why layers, not one check Backup verification has a cost gradient. Reading a manifest is milliseconds; re-reading a multi-terabyte repository is hours of I/O and object-store egress; a full restore is hours of wall clock plus a scratch host. Because you cannot afford the top of the ladder on every backup, you build a pyramid: the cheap layers run continuously and catch the common, dumb failures early, while the expensive layers run on a cadence and catch the failures the cheap ones structurally cannot see. ## Layer 1 - existence and inventory Does the expected artifact exist for the expected time window, in the expected location, with a plausible size and the expected set of member files? Alerting on "newest full backup is older than N hours" and "backup size deviates more than X percent from the trailing median" catches a shocking share of real incidents: schedulers that stopped firing, credentials that expired, a repository pointed at a bucket that was renamed, and a dump that produced a 4 KB file containing only an error message. It cannot say anything about the content of those bytes. ## Layer 2 - checksum and manifest verification Good physical-backup tools record a per-file (often per-block) checksum in a manifest at backup time. A verify operation re-reads the repository and recomputes those checksums, confirming that every file the manifest references is present and byte-identical to what was written, and that the chain of dependent backups and archived log segments is complete. This catches: silent media corruption, truncated uploads, an accidentally deleted archived log segment that breaks the chain, and a partial or interrupted retention expiry that removed a file some later backup still depends on. It is the right thing to run on every backup because it is automatable, needs no scratch host, and yields a hard pass/fail. It misses everything logical. A perfectly checksummed backup of a replica that stopped replicating three months ago verifies clean. So does a backup that omitted an entire schema, or one whose encryption key has since been rotated away. ## Layer 3 - restore and open This layer materialises the backup on an isolated instance, applies whatever log or incremental chain is needed, and brings the engine to a consistent, open state. It is the first layer that proves the artifact is a *database* rather than a well-checksummed pile of files. What only this layer catches: an incompatible engine version or missing extension, a broken or gapped log chain that verification's file-level view thought was fine, encryption keys that are no longer retrievable, permissions and tablespace paths that do not exist on the target, tooling changes that silently altered the restore command, and - critically - **how long the restore takes**. Restore duration is the dominant, unmeasurable-any-other-way component of a recovery-time objective. ## Layer 4 - application-level assertions An engine that opens can still hold the wrong data. This layer asks whether the restored database is *usable by the application*: - Freshness: the maximum timestamp or identifier in an append-only table, compared with the moment the backup claims to represent. A large gap means you were backing up a stale source. - Completeness: counts of tables, indexes, sequences, views, routines, and constraints against a production census; row counts for the largest tables within an expected band. - Integrity: spot checks that foreign-key relationships resolve and that a few business invariants hold. - Behaviour: run two or three queries the application really issues, and ideally boot the application in a test mode against the restored instance. This is what catches selective-dump mistakes, missing large objects or sequences, and "the restore succeeded but the app cannot start" - which is the failure mode that turns a 30-minute recovery into a 6-hour one. ## Putting it together A defensible programme looks like: inventory and freshness alerts on every backup; checksum/manifest verification on every backup or at least daily across the repository; automated restore-and-assert of at least one cluster per day, rotating across the fleet so every cluster is proven within its tier's window; and a full human game day periodically. Record the outcome and the timings of every restore, because those timings are the only honest input to the recovery-time objective. The answer an interviewer is listening for is the distinction: **verification proves the bytes; restoration proves the database; assertions prove the data.** Candidates who name only the middle layer typically have never had to explain a clean-verify, unusable-restore incident.

  • Give a concrete failure that passes checksum verification but fails a real restore.
    Backups taken from a replica whose replication stopped weeks ago: every file is intact and every checksum matches, so verification is green, but the restored database is weeks stale. Another is an encryption key that has been rotated or destroyed - the ciphertext verifies perfectly against its manifest, yet nothing can decrypt it. Both are logical failures invisible to byte-level checks.
  • How would you assert on restored data without hardcoding brittle expectations?
    Use relative and structural assertions rather than fixed values: the newest row's timestamp must be within the expected lag of the backup time, row counts must be within a band of the trailing production values, and the object census (tables, indexes, sequences, routines, constraints) must match a snapshot exported from production at backup time. Store that census alongside the backup so the check compares like with like.

saying these in an interview costs you the question

  • Believing checksum or manifest verification proves the backup is restorable.
  • Skipping the checksum layer because 'we do a restore drill every quarter' - three months of undetected corruption is a real loss window.
  • Verifying only the latest full backup and never the incremental or log chain it depends on.
  • Calling a restore successful when the engine starts, without any check that the data is current or complete.
  • Running verification against a locally cached copy rather than the actual stored repository.

context