Explain the difference between full, incremental, and differential backups of a relational database, and how each choice affects the backup window, storage cost, and restore time.
answer
- Full = everything, stands alone
- Incremental = since last backup (any type)
- Differential = since last full
- Restore: 1 vs 2 vs 1+N artifacts
- Long incremental chain = fragile
basics
~20 sA full backup copies everything. An incremental copies only what changed since the previous backup of any type. A differential copies everything changed since the last full. Incrementals are smallest to take but slowest to restore, because the whole chain must be replayed.
solid answer
~50 s**Full** is a complete self-contained copy: simplest to restore, largest, slowest to take. **Incremental** copies only the blocks or rows changed since the *previous backup*, whatever type it was. Nightly runs are tiny, so the window and storage are minimal, but restore means last full + every incremental in order. One missing or corrupt link makes everything after it useless. **Differential** copies everything changed since the *last full*, ignoring intermediate differentials. Each one grows through the week, but restore needs only two artifacts: the full plus the newest differential. The tradeoff is backup cost versus restore cost and chain fragility. A common shape is a weekly full plus daily incrementals or differentials on a grandfather-father-son retention rotation. Choose incrementals when the nightly window or network bandwidth binds; choose differentials when restore speed and fewer moving parts matter more.
go deeper
Be able to define all three crisply and say which needs the most artifacts to restore. Getting the differential definition right is the actual screen.
Add the restore arithmetic (1 vs 2 vs 1+N) and the page-level meaning of 'changed', and name a concrete weekly-full schedule.
Talk about chain fragility, synthetic fulls, chain-aware retention expiry, and the fact that the choice is driven by which resource — window, bandwidth, or restore time — is scarce.
Frame it as moving cost between the steady state and the incident, tie it to the tolerated restore duration and storage budget, and note that the policy is only real once a restore has been timed.
## The three types All three answer one question: of everything the database holds, how much do we copy tonight? **Full backup** — the entire dataset. It stands alone: restoring needs only this one artifact. It is the largest and slowest to produce and re-reads data that has not changed in years. **Incremental backup** — only what changed since the *previous backup of any type*. Monday's incremental is relative to Sunday's full; Tuesday's is relative to Monday's incremental. Each one is small. **Differential backup** — everything changed since the *last full*, ignoring other differentials. Monday's is small, Friday's contains Monday-through-Friday changes. ## What "changed" means At the physical level, incrementals are computed per **page/block**: the engine or backup tool tracks which fixed-size pages were modified (a changed-block bitmap, a page LSN comparison, or a scan of the transaction log). Touching one row rewrites the whole page, so a workload of tiny scattered updates produces an incremental much larger than the bytes actually changed. At the logical level, "changed" usually means rows above a watermark column, which is why logical incrementals struggle with deletes and in-place updates. ## Restore arithmetic This is the part interviewers push on. - Restore from full: apply 1 artifact. - Restore from differential: apply 2 artifacts (full + latest differential), regardless of the day. - Restore from incrementals: apply 1 + N artifacts, where N grows to the length of the chain. Six days after the full, that is seven sequential operations, often each requiring the previous one's output. So incrementals move cost from the nightly window to the incident. Differentials do the opposite: storage grows through the week (Friday's differential may approach the size of a full on a write-heavy system) but restore stays two steps. ## Chain fragility A differential chain has one dependency: the full. An incremental chain has N. If a single incremental is corrupt, unreadable, or was silently skipped, every backup after it is worthless — you can only restore back to the last good link. This is why long incremental chains are periodically "rolled up" into a new synthetic full (merging the chain server-side into one artifact) or simply cut with a fresh weekly full. ## Where the constraints come from Pick based on which resource is actually scarce: - **Backup window / IO headroom**: incrementals, because you read and write the least. - **Off-site bandwidth or object-storage egress**: incrementals, same reason. - **Restore time and operational simplicity under pressure**: differentials, or more frequent fulls. - **Retention cost**: incrementals store fewer duplicate bytes, but many storage backends dedupe or do incremental-forever with synthetic fulls, blurring the difference. ## Retention The types combine with a retention policy, classically grandfather-father-son: keep the last N days of dailies, the last N weekly fulls, and monthly fulls for long retention. Retention is a separate axis from backup type — a differential is useless once its parent full has aged out, so expiry must be chain-aware. ## Practical default For a mid-sized OLTP database: weekly full, daily incremental if the write volume is modest relative to the dataset, daily differential if it is not and you value restore speed. Then measure both the produced sizes and an actual restore, because the arithmetic above is only a model until you run it.
- You take a full on Sunday and incrementals every night. On Thursday the Tuesday incremental turns out to be corrupt. What can you restore?Only up to Monday: the Sunday full plus the Monday incremental. Wednesday's and Thursday's incrementals describe changes relative to a state you can no longer reconstruct, so they cannot be applied. This is the practical argument for shorter chains, synthetic fulls, or differentials — and for verifying each artifact when it is written rather than discovering the gap during an incident.
- Why can an incremental backup be much larger than the number of bytes the application actually changed?Physical incrementals track modified pages, not modified rows. A single-column update to one row marks the whole 8KB or 16KB page dirty, so a workload of small scattered writes across many pages produces an incremental close to the size of the working set. Index pages are dirtied too, and maintenance operations such as a rebuild or a bulk vacuum can rewrite pages whose logical content did not change at all.
Incremental backups are a diary entry each day: to reconstruct the year you read every entry in order, and a missing page breaks the story. A differential is a rewritten summary of everything since January: longer each time, but you only ever read two documents.
saying these in an interview costs you the question
- Saying a differential is 'since the last differential' — that is an incremental
- Assuming incrementals always restore faster because the files are smaller
- Treating the choice as pure storage economics and ignoring chain fragility
- Believing an incremental chain can be restored out of order or partially
- Expiring a full while its dependent differentials or incrementals are still retained