skip to content

Your database backups are encrypted at rest. What can go wrong at restore time because of that encryption, and how do you manage backup encryption keys so a restore is still possible during a real disaster?

level: seniorimportance: should knowfreq 32%

answer

  1. encryption turns confidentiality risk into availability risk
  2. key must not share fate with the data
  3. envelope encryption: rotate the wrap, not the data
  4. never delete a key version a retained backup needs
  5. break-glass decrypt role + offline escrow

basics

~20 s

An encrypted backup is only as recoverable as its key. Keys stored on the lost host, rotated away, region-pinned, or reachable only through the failed environment make backups unrecoverable. Keep keys in a separate, replicated, escrowed key store, retain old key versions for the backup's lifetime, and prove it by restoring.

solid answer

~1 min

Encryption converts a data-confidentiality risk into an **availability** risk: the backup is now unrecoverable without a key, and the key becomes a single point of failure. The realistic failure modes: the key or passphrase lived on the database host that died; the key material sits in a key service in the same region that just went down; a rotation policy destroyed the old key version while backups encrypted under it are still retained; the key policy grants decrypt only to a role that only exists in the failed environment; or a cross-account copy was never re-encrypted with a key the recovery account can use. What I do: keep backup keys in a managed key service that is multi-region or explicitly replicated to the recovery region; **never destroy a key version while any retained backup depends on it** - rotation adds a new version for new backups and keeps old ones readable; grant decrypt to a break-glass identity in the recovery account, not only to production roles; keep an offline escrow copy of the wrapping key or passphrase in a physical safe or dedicated secret store; and validate all of it by doing the restore in the recovery account with production credentials disabled.

go deeper

for a junior

Say the essential thing: an encrypted backup is useless without its key, so the key must be stored somewhere separate from the database it protects.

for a middle

Name the concrete failure modes - key on the lost host, region-pinned key, rotation deleting old versions - and describe envelope encryption and multi-region key availability.

for a senior

Design the whole lifecycle: key retention tied to backup retention, break-glass decrypt identity separate from production, delete rights withheld from production, audit alerts on key changes, and a DR drill with production credentials disabled.

for a principal

Weigh confidentiality against availability explicitly - who holds custody, threshold escrow, the ransomware trade-off between reachable decrypt and unreachable delete, and how key custody is evidenced to auditors and regulators.

## Encryption trades confidentiality risk for availability risk Backups belong off-site, often in third-party storage, and they contain the entire database - so encrypting them at rest is not optional in most environments. The cost is that a backup now has a hard dependency on a second artefact, the key. Anything that makes the key unavailable makes the backup a pile of noise. Cryptographic erasure is a *feature* when you deliberately destroy a key to render data unrecoverable; it is a catastrophe when it happens by accident. ## The failure modes that actually happen 1. **Key co-located with the thing it protects.** A passphrase in the backup tool's config file on the database host, or a key file on the same volume. The disaster that destroys the database destroys the key. This is the most common and most embarrassing failure. 2. **Key service scoped to the failed region or account.** A region-pinned key in a managed key service cannot decrypt from the DR region during exactly the event you built DR for. Similarly, a key policy that only names production roles is unusable from a recovery account. 3. **Rotation that destroys old versions.** Rotation is good practice; *deleting* old key versions is only safe if nothing retained still depends on them. A seven-year archival backup encrypted under a key version scheduled for deletion in 90 days is a time bomb. 4. **Client-side versus server-side confusion.** With server-side storage encryption, the storage platform decrypts transparently for any authorised reader - convenient but it means storage credentials are enough to read the data. With client-side or tool-managed encryption, the tool holds the key and the storage provider never sees plaintext - stronger, but now *you* own key availability entirely. Teams often assume the other model is in force. 5. **Copies that cross a trust boundary unchanged.** Replicating an encrypted object to another account or provider without re-wrapping under a key that the destination can use produces a copy that exists but cannot be opened where it is needed. 6. **No one knows the procedure.** The key exists, is reachable, and nobody on call knows which key, which version, or which command supplies it - so recovery time balloons while people search. ## Managing keys so restores work **Envelope encryption** is the standard structure: each backup is encrypted with a per-backup data key; that data key is itself encrypted (wrapped) by a long-lived master key held in a key management service and stored alongside the backup. Rotating the master key then only re-wraps data keys, not terabytes of ciphertext, and old backups stay readable as long as the master key version that wrapped them still exists. Operational rules: - **Separate the key's fate from the data's fate.** Different host, different region, different account, ideally different provider. The question to ask is: *what single event removes both the backup and its key?* - **Replicate or make the key multi-region**, and confirm the DR region can decrypt without any dependency on the primary region's control plane. - **Retain key versions at least as long as the longest backup retention.** Wire this into the retention policy explicitly; the key lifecycle and the backup lifecycle must be reviewed together. - **Grant decrypt to a break-glass identity** that exists independently of production - a role in the recovery account, guarded by MFA and audited, so a compromised or deleted production identity does not block recovery. - **Escrow**. For the highest tiers, keep an offline copy of the wrapping key or passphrase - split with a threshold scheme (m of n custodians) and stored physically - so total loss of the online key infrastructure is survivable. - **Balance against the ransomware threat.** The decrypt path must be reachable in a disaster, but the *delete* path must not be reachable by production credentials. These are different permissions; grant them to different identities. - **Audit and alert** on key deletion, disablement, and policy changes. A scheduled key deletion should page a human the way a dropped table would. ## Prove it, don't assume it The only credible verification is a restore executed **in the recovery environment, with production credentials unavailable**, using only the break-glass path. That drill catches the region-pinned key, the missing key-policy grant, and the undocumented passphrase in one shot. Record which key version each backup used in the drill log, and re-run the drill after every key rotation and every change to key policy or DR topology - these are precisely the changes that silently break decryptability while every backup job stays green.

  • How do you rotate backup encryption keys without making older backups unrecoverable?
    Use envelope encryption: each backup gets a unique data key wrapped by a master key version, and the wrapped data key is stored with the backup. Rotation creates a new master key version used for new backups while all previous versions stay enabled, so old backups still unwrap. The rule is that a key version may only be disabled or destroyed after every backup wrapped under it has passed its retention date.
  • When would you deliberately destroy a backup encryption key?
    For cryptographic erasure - when a legal or contractual obligation requires data to become unrecoverable and physically deleting every copy is impractical, for example ciphertext already replicated into immutable, write-once storage whose retention has not expired. Destroying the key renders those copies permanently unreadable. It must be a deliberate, approved, audited action with a documented scope, because it is irreversible and can take unrelated retained backups with it if the key is shared.

A safe-deposit box: the bank's vault protects the contents, but if the only key was in the house that burned down, the box may as well be filled with concrete.

saying these in an interview costs you the question

  • Storing the backup passphrase or key file on the database host it protects.
  • Assuming a rotation policy that destroys old key versions is safe while long-retention backups still depend on them.
  • Granting decrypt only to production roles, so recovery from a separate DR account is impossible.
  • Confusing storage-provider server-side encryption with tool-managed client-side encryption when reasoning about who can read the backup.
  • Never rehearsing a restore with production credentials disabled, so the key-availability gap is discovered during the real incident.

context