Why is a lakehouse table format's time travel not a substitute for backups?
answer
- it lives in the same bucket as the live table
- same credentials, same blast radius
- references, not copies
- retention is a very short recovery objective
- versioning is not the same thing as a backup
basics
~20 sSnapshots reference files in the same bucket, catalog and account as the live table, so anything that destroys or loses those objects destroys the history with them. Time travel protects against bad writes inside the retention window, not against losing the storage.
solid answer
~50 sTime travel and rollback are excellent at one failure class: a **logical** error — a bad merge, a wrong filter, a duplicated load — discovered inside the retention window while everything else is intact. They are useless against the failures backups exist for. Old snapshots point at objects in the same bucket, under the same credentials, registered in the same catalog, so a deleted bucket, a purging `DROP TABLE`, a compromised key that deletes objects, a region outage or a lost catalog entry takes the history along with the current data. Retention bounds it further: a problem found after the horizon has no version left to return to. And rollback itself is normally recorded as a *new* commit rather than an erasure, so it is a data-correction tool, not a disaster-recovery one. Real protection is separate: cross-account or cross-region replication of the storage, object versioning with delete protection, backups of the catalog, and periodic independent copies of critical tables — with restores actually rehearsed.
go deeper
Know the one-line distinction: time travel is versioning inside the table's own storage, a backup is a separate copy elsewhere.
Be able to name what shares fate with the history — the same bucket, credentials and catalog — and to state retention as the limit on how far back recovery reaches.
Show the layered design: rollback for logical errors, object versioning and delete protection, cross-account replication, catalog backup, and rehearsed restores mapped to specific failure classes.
Own the recovery objectives themselves: which tables justify which tier of protection, what the retention window commits you to as an RPO, and how the compliance-deletion requirement cuts against long history.
## What time travel actually protects against One thing, very well: a **logical** error on an otherwise healthy table. Someone ran a merge with an inverted predicate; a backfill double-loaded a month; a schema-mangling job wrote garbage. The bad state is a commit, the good state is the commit before it, and rollback re-points the table at the good file set in seconds. No restore, no downtime, no data movement. That is a genuinely superior answer to the classic warehouse restore, and it is why teams reach for it first. ## The failures it does not survive Every one of these takes the history down with the table: - **Storage loss.** The bucket or container is deleted, corrupted, or lost with a region. Old snapshots hold references, not copies, so their files are gone too. - **Purging drops.** A `DROP TABLE` that also removes data, or a lifecycle rule that deletes objects under the table prefix, leaves nothing for any snapshot to reference. - **Credential compromise or a hostile actor.** Anything that can delete the current data can delete the historical data — same bucket, same permissions, same blast radius. Ransomware does not respect snapshots stored beside the data. - **Catalog loss.** If the pointer to current metadata lives in a catalog and the catalog entry is lost, the objects may survive while nothing knows what the table is. Backing up storage without backing up the catalog is a half backup. - **Expiry.** Maintenance ran; the versions you wanted are gone. Self-inflicted, common, and permanent. - **Late detection.** A corruption discovered two months later, on a table with a two-week retention, has no version to return to at all. ## Retention is an RPO, and a short one Stating it as a recovery objective clarifies the argument: retention is the maximum age of a state you can return to, and it is usually days. Backup policies for important data are usually measured in months, with tiers, and with copies that live somewhere a compromise of the primary account cannot reach. Those are different requirements, and one does not imply the other. ## Rollback semantics worth knowing Rolling back typically appends a **new** commit that re-points the table at an earlier file set, rather than erasing what happened. That is a feature — the incident stays auditable and the rollback itself is reversible — but it means rollback is not a way to make data disappear. If the requirement is "this data must not exist any more," you need the files rewritten and the referencing snapshots expired, not a rollback. ## What an actual protection design looks like Layer it, and be explicit about which layer covers which failure: 1. **Time travel and rollback** — bad writes, within retention, table intact. Fast, cheap, first line. 2. **Object versioning plus delete protection** on the storage — accidental and malicious object deletion. 3. **Cross-account or cross-region replication** of the table's storage — account compromise, region loss. A different account matters more than a different region for the compromise case. 4. **Catalog backup or a redeployable catalog** — so a restored bucket can be turned back into tables. 5. **Periodic independent copies of critical tables** — the answer to "we found it three months later." 6. **Rehearsed restores.** An untested restore path is a hypothesis. ## Answering the interview version The crisp framing is: time travel is *versioning*, not *backup*, because it protects the contents of the table from you, but does not protect the table from the platform. Say what it does cover, name two failure classes it cannot (loss of the storage, and detection after retention), and describe the second copy that does cover them. Candidates who claim rollback makes backups unnecessary are describing a platform with a single point of failure and have usually never restored anything.
- What exactly does rolling back to an earlier version do to the versions in between?Typically it appends a new commit that makes the earlier file set current again, leaving the intervening versions in the table's history until they expire. The incident stays auditable and the rollback is itself reversible. It also means rollback does not erase data: if something must genuinely cease to exist, the files have to be rewritten and the snapshots referencing them expired.
- If you replicate the table's bucket to another region, is that enough for recovery?Not on its own. You also need the catalog entry that says which metadata object is current, or the replicated objects are just files nobody can interpret as a table. Replication also copies deletions unless it is configured otherwise, so a hostile or accidental purge can propagate. Pair it with object versioning, delete protection, and a rehearsed restore.
- Where does time travel genuinely beat a traditional backup?Recovery speed and precision for logical errors. Returning to the version before a bad merge is a metadata operation measured in seconds with no data movement, and you can inspect and diff the candidate version before committing to it. A backup restore of the same table is a bulk copy with real downtime. Use both, for different failure classes.
Undo in a text editor: superb for the paragraph you just ruined, worthless when the laptop is stolen — and it never survives closing the file.
saying these in an interview costs you the question
- Claims rollback makes backups unnecessary
- Assumes snapshots are stored outside the table's own bucket
- Forgets the catalog also needs backing up
- Thinks rollback erases the bad versions from history
- Treats a days-long retention window as a recovery objective for months-old errors