How would you design a retention and rotation policy for database backups - which tiers to keep, how long to keep each, and how retention interacts with the window over which you can still recover the database to an arbitrary past moment?
answer
- tier by detection latency, not disk price
- GFS: daily / weekly / monthly
- fine window = oldest base + unbroken log
- expire by chain, never by file mtime
- keep N restorable sets; verify after expiry
basics
~20 sTier it: dense recent backups plus retained log for a fine-grained recovery window (say 7-14 days), then weekly and monthly copies kept for months or years for compliance and slow-burn corruption. Expire the base backup and its dependent log together, and never expire a backup another one depends on.
solid answer
~60 sI design retention from the **recovery scenarios**, not from disk price. - **Operational window** (days 0-14): frequent fulls or incrementals plus the complete archived log, so any moment in that window is recoverable. This covers the overwhelmingly common case - a bad deploy or bad data change noticed within hours. - **Medium tier** (weeks 2-12): weekly fulls only, no fine-grained log. Covers corruption discovered late, an unnoticed data-quality regression, or a legal question. - **Long tier** (months to years): monthly or quarterly copies, off-site and immutable, driven by regulation and contracts, not by engineering taste. The hard constraints are dependency-safe expiry and copy independence. In an incremental or differential scheme every backup depends on its ancestors, so expiry must remove whole chains: deleting an old full silently invalidates every incremental built on it. Likewise, log segments must be retained back to the oldest base backup you still intend to use, or that backup is unrestorable. I also make retention **automated and verified** - an expiry bug is only discovered during a restore, so verification must run after expiry, not before.
code
text · 7 linesoperational : daily incremental + weekly full, keep 14 days, WAL retained fully
medium : weekly full, keep 12 weeks, no WAL
long : monthly full, keep 7 years, off-site + object-lock
recoverable to ANY moment: now - 14 days .. now
recoverable only to a backup instant: 14 days .. 7 years
minimum retained restorable full-chains: 3go deeper
Know the shape - keep recent backups densely and older ones sparsely - and that keeping a backup file is not the same as being able to restore to any moment.
Explain the tiers with concrete cadences and retentions, and articulate that the fine-grained recovery window is bounded by the retained log, not by how many fulls you keep.
Lead with dependency-safe expiry, retention counted in restorable sets, verification after expiry, and the interaction between storage class, retrieval latency and recovery time.
Frame retention as a negotiated risk and cost position - regulatory obligations, storage and egress spend, per-tier recovery expectations, and how the policy is evidenced and reviewed.
## Start from scenarios, not from storage cost Retention answers one question: *how far back might we need to go, and at what granularity?* Different failures have wildly different detection latencies: - A failed migration or a bad `DELETE` is noticed in minutes to hours. It needs **fine granularity**, not depth - the ability to land on the second before the damage. - Slow data corruption or a subtly wrong batch job may go unnoticed for weeks. It needs **depth** at coarse granularity - a clean weekly copy from before the regression started. - Legal discovery, audit, or a contractual retention clause needs **depth measured in years** at very coarse granularity, and often immutability. A single flat policy ("keep 30 days") serves none of them well: it is too shallow for compliance and too expensive for the fine granularity nobody needs at day 29. ## A workable tiering A common shape, often called grandfather-father-son: | Tier | Cadence | Retention | Log retained? | Purpose | |---|---|---|---|---| | Operational | daily full or incremental | 7-14 days | yes, complete | arbitrary-moment recovery from recent mistakes | | Medium | weekly full | 4-12 weeks | no | late-discovered corruption, quarter-end questions | | Long | monthly or quarterly full | 1-7 years | no | regulatory, contractual, forensic | The operational tier is the expensive one because it retains the whole log stream; the long tier is cheap per copy but must live off-site and, for the highest assurance, under write-once retention. ## How retention constrains the fine-grained recovery window The ability to recover to an arbitrary past moment is bounded by **the oldest usable base backup plus an unbroken run of archived log from that backup forward**. Two consequences follow, and both are the classic interview trap: 1. **Expiring the log truncates the window from the far end.** If you keep 14 days of log but 90 days of fulls, the 30-day-old full can only be restored to the exact moment it was taken, not to any point after it - the log needed to roll forward is gone. That is a legitimate design (a coarse restore point), but it must be a deliberate choice, not a surprise. 2. **Expiring a base backup orphans its log.** Log segments older than the oldest retained base backup are unusable and should be expired with it; otherwise you pay to store bytes that can never be replayed. ## Dependency-safe expiry With incremental or differential backups, an archive is not a standalone object - it is a node in a chain. Deleting a full backup silently invalidates every incremental and differential built on it. Rules: - Expire by **chain**, never by file age. A good backup tool does this natively; a hand-rolled `find -mtime +N -delete` cron does not, and it is a recurring source of "we had 60 days of backups and none of them restored". - Retention counts should be expressed in **complete restorable sets**, e.g. "always keep at least 3 restorable full-chains", so a failing backup job cannot cause expiry to erode you down to zero good copies. - Never let expiry run **before** the newest backup has been verified. Otherwise a corrupted new backup plus an eager expiry removes the last good one. ## Rotation, immutability, and cost Rotation is the mechanical side: which copies move to colder or remote storage and when. A typical flow keeps the operational tier on fast same-region storage, copies weeklies and monthlies to a remote account with object-lock retention matching the retention period, and lets storage lifecycle rules transition older objects to archival classes. Two cautions: archival storage classes have **retrieval latency and cost** that must be reflected in your recovery-time expectations for those tiers, and object-lock retention must be set to at least the policy period or lifecycle rules will happily delete what compliance requires. ## Operating the policy - Encode retention in the backup tool's configuration, not in ad-hoc scripts, so expiry understands dependencies. - Alert on both directions: too few restorable sets (recovery risk) and unexpected growth (expiry silently failing). - Re-verify after every expiry cycle; expiry bugs are invisible until a restore. - Review the policy whenever data volume, regulation, or the recovery-point objective changes - retention is where the objective becomes a bill. The answer an interviewer wants: tiers keyed to detection latency, a fine-grained window bounded by log retention, expiry that respects backup chains, retention counted in restorable sets, and verification after expiry.
- What breaks if a cron job deletes backup files older than 30 days on an incremental backup repository?Incrementals depend on their ancestors, so deleting a 31-day-old full silently invalidates every newer incremental built on it, and deleting archived log segments truncates the roll-forward window. The repository still looks healthy - the files that remain are intact - and the damage only surfaces during a restore. Expiry must be performed by the backup tool, which understands chain dependencies, and expressed as a number of complete restorable sets.
- How does moving old backups to an archival storage class affect your recovery planning?Archival classes trade a much lower storage price for retrieval latency measured in minutes to hours and a per-retrieval fee. Any tier stored that way effectively has a much longer recovery time, so it must never hold the copy you rely on for an operational incident. Document a separate, longer recovery-time expectation for those tiers and keep the recent operational window on fast storage.
saying these in an interview costs you the question
- Applying one flat retention period to every backup regardless of the failure it protects against.
- Assuming a 90-day backup retention means you can recover to any moment in the last 90 days, when log retention is only 7 days.
- Deleting backups by file age with a generic cleanup script, breaking incremental chains.
- Retaining archived log older than the oldest base backup, which can never be replayed and only costs money.
- Running expiry before verifying the newest backup, so a bad new backup plus expiry leaves no good copy.