Explain the 3-2-1 backup rule and how you would apply it to a production relational database. What failure modes do off-site and immutable copies protect against that a second local copy does not?
answer
- 3 copies, 2 media, 1 off-site
- 3-2-1-1-0 adds immutable + zero errors
- replica != backup (DROP TABLE replicates)
- one stolen credential must not reach all copies
- object lock / separate account / air gap
basics
~20 sKeep three copies of the data, on two different media or storage systems, with at least one off-site. For a database that means the live cluster, a local or same-region backup repository, and a copy in another region or account - ideally write-once so it cannot be deleted or encrypted by an attacker.
solid answer
~60 s**3-2-1** means: three copies of the data (the production copy plus two backups), on two distinct storage technologies or platforms, with one copy off-site. The point is that each copy should fail for *independent* reasons. For a relational database: copy one is the live cluster (replicas do not count as backups - they replicate a `DROP TABLE` instantly), copy two is a backup repository in the same region for fast restores, copy three is a cross-region or cross-account copy for regional loss. A second local copy only protects against media failure. An off-site copy additionally covers site loss, region outage, and blast-radius events like a bad automation run that wipes a whole account. Immutability - object-lock or write-once retention on the remote copy, held in a separate credential domain - is what covers ransomware and a compromised or careless administrator, because the same credentials that can delete production must not be able to delete the backup. Modern phrasings add 3-2-1-1-0: one immutable or offline copy, zero errors in verification.
go deeper
State the rule - three copies, two kinds of storage, one off-site - and be clear that a replica is not a backup because it copies mistakes too.
Map it to a concrete database setup and explain what each leg covers, including why immutability is what defends against ransomware and accidental deletion.
Reason about independence explicitly: name the shared dependency that would take out multiple copies, design the credential separation, and tier retention so the off-site copy stays affordable.
Treat it as a risk budget across the estate - which tiers get an immutable off-site leg, what recovery from the remote copy costs in time and money, and how the control is evidenced for auditors and insurers.
## The rule 3-2-1 is a heuristic for **independence of failure**: at least **3** copies of the data, on at least **2** different media or storage platforms, with at least **1** copy off-site. It comes from photography and general IT, but maps cleanly onto databases. An extended form, 3-2-1-1-0, adds one **immutable or offline** copy and **zero** errors in verification. The rule is not about paranoia arithmetic. It exists because copies that share a dependency fail together. Two backups in the same bucket, under the same credentials, in the same region, written by the same tool, are close to one copy for risk purposes. ## Mapping it onto a relational database - **Copy 1 - the live cluster.** The primary and its data. Note that streaming replicas and multi-AZ standbys are **not** backups: they faithfully reproduce logical damage. A dropped table, a bad migration, or a `DELETE` without a `WHERE` clause reaches the standby in milliseconds. - **Copy 2 - a local or same-region backup repository.** Physical base backups plus the archived write-ahead log, on a different storage system from the database volumes. This is the copy you actually restore from most of the time, because it is fast to fetch. - **Copy 3 - an off-site copy.** Another region, and ideally another cloud account or provider. This is the copy that exists for events that take the whole site or account with them. The "two media" leg used to mean disk plus tape. Today it usually means block storage plus object storage, or two providers - the intent is that a defect or outage in one storage technology does not take both copies. ## What the off-site copy buys you A second local copy protects against exactly one class of event: media or host failure. It does nothing about: - **Site or region loss** - power, network, flood, fire, or a provider regional outage. - **Blast-radius automation** - a Terraform apply or a cleanup script that deletes a whole account's resources, including both local copies. - **Account compromise** - credentials that reach production also reach the local repository. ## What immutability buys you Ransomware and malicious insiders explicitly hunt backups first, because destroying recovery leverage is what makes the extortion work. If the credentials that write backups can also delete them, an attacker with those credentials has both. Countermeasures: - **Object lock / write-once-read-many retention** on the remote copy, so objects cannot be deleted or overwritten before their retention date, not even by the account root. - **A separate credential and trust domain** - a different account, subscription, or project, with a one-way push or, better, a pull performed by the backup account so production holds no delete rights on it. - **Versioning plus MFA-protected deletion** so accidental or scripted deletes are recoverable. - **Air-gapped or offline copies** for the highest tiers, where nothing online can reach the media at all. The practical test question is: *which single credential, if stolen, destroys every copy?* If there is one, 3-2-1 is decorative. ## Cost and pragmatism Off-site copies cost storage plus cross-region transfer, and restoring from them is slower. The usual compromise is tiering: keep a short window of dense, frequent backups locally for fast recovery, and push a thinner set (say weekly fulls plus log, or daily fulls with a shorter retention) to the remote immutable copy for disaster and ransomware scenarios. Compress and encrypt before transfer. Finally, the rule is worthless without the "0 errors" clause. A remote immutable copy that has never been restored from is an untested backup that also happens to be expensive. Every leg of 3-2-1 needs its own periodic proof-of-restore, and the off-site leg is the one most often skipped and most often needed.
- Why does a synchronous standby or a multi-AZ replica not count as one of the three copies?Replication propagates logical changes faithfully and fast, so it reproduces the damage you most often need to recover from - a dropped table, a bad migration, an unbounded delete, or corruption written by the application. Replicas protect availability against host and zone failure; they do not give you a point in the past to return to. Only a backup with a retained history does that.
- How do you keep the off-site copy from being deleted by the same credentials that manage production?Put it in a separate account or subscription with its own identity domain, and give the production role write-only or no rights over it - either the backup account pulls, or production pushes into a bucket where it lacks delete and overwrite permissions. Layer object-lock retention on top so objects cannot be removed before their retention expiry even by an administrator, and require MFA or a break-glass process for any lock change.
Spare house keys: one on your keyring, one in a kitchen drawer, one with a friend across town. The drawer key covers losing your keyring; only the friend's key covers the house burning down.
saying these in an interview costs you the question
- Counting a streaming replica or multi-AZ standby as a backup copy.
- Two 'copies' that live in the same bucket, region, and credential domain and therefore fail together.
- Assuming cloud object storage durability guarantees make off-site copies unnecessary - durability does not protect against deletion, ransomware, or account loss.
- Giving the production backup role delete permissions on the immutable copy, which defeats the immutability.
- Never restoring from the off-site copy, so its retrieval path and credentials are untested when they matter most.