In Apache Iceberg, which files does the expire_snapshots procedure actually delete?
answer
- old commits still hold on to files
- age of a file is not the test
- reachable from a kept snapshot or not
- two phases: drop snapshots, then delete files
- retain_last and history.expire.max-snapshot-age-ms
basics
~20 sIt removes snapshot entries older than the retention window from table metadata, then deletes the data files, delete files, manifests and manifest lists that no remaining snapshot references. Anything still reachable from a kept snapshot survives, whatever its age.
solid answer
~40 sEvery Iceberg commit writes a new snapshot, and old snapshots keep pointing at the files they saw, so nothing is reclaimed until snapshots are expired. `expire_snapshots` works in two phases: it drops the selected snapshots from table metadata, then computes the files reachable from the snapshots that remain and physically deletes the data files, delete files, manifest files and manifest lists that only the removed snapshots referenced. **Reachability decides, not file age** — a five-year-old data file the current snapshot still lists is untouched. Arguments are `older_than` and `retain_last`, defaulting to the table properties `history.expire.max-snapshot-age-ms` (5 days) and `history.expire.min-snapshots-to-keep` (1). Snapshots held by a branch or tag are protected by that ref's own retention. It does not remove old `metadata.json` files or uncommitted orphan files — those are separate settings and a separate procedure.
code
sql · 12 lines-- keep 5 recent snapshots, expire anything older than the cutoff
CALL prod.system.expire_snapshots(
table => 'db.events',
older_than => TIMESTAMP '2026-08-14 00:00:00',
retain_last => 5
);
-- make retention declarative instead of per-call
ALTER TABLE prod.db.events SET TBLPROPERTIES (
'history.expire.max-snapshot-age-ms' = '604800000',
'history.expire.min-snapshots-to-keep' = '5'
);go deeper
Recall that Iceberg keeps every past version of a table as a snapshot, and that a cleanup procedure is what eventually frees the old files. Know the name expire_snapshots and that it trades history for storage.
Explain the two phases — drop snapshot entries, then delete only files no surviving snapshot references — and name the older_than and retain_last arguments plus the history.expire.* table properties that supply their defaults.
Show judgment about the retention window: long enough for the slowest reader and any audit or rollback requirement, short enough to control object-storage cost, with tags pinning snapshots that must outlive it. Be ready to explain why storage only drops after expiry follows compaction.
Own retention as policy rather than a job argument: retention tiers set through table properties, tagging conventions for compliance points, and the cost-versus-recoverability tradeoff you are signing the platform up for across every table.
## Why snapshots accumulate Every write to an Iceberg table — append, overwrite, `DELETE`, `MERGE`, or a compaction job — is a commit that produces a new `metadata.json` containing a new snapshot. A snapshot points at one manifest list (`snap-<id>-<seq>-<uuid>.avro`), which points at manifest files, which list the data files and delete files that make up the table at that instant. Nothing is ever mutated in place. That immutability is what gives Iceberg snapshot isolation, time travel and rollback — and it is also why a table that never expires anything keeps every data file it has ever written, plus a metadata file whose snapshot log grows without bound. ## What the procedure does `expire_snapshots` runs in two phases. 1. **Metadata phase.** It commits a new `metadata.json` in which the selected snapshots no longer appear in `snapshots` or in the snapshot log. 2. **Cleanup phase.** It computes the set of files reachable from the snapshots that remain, compares it with the files reachable from the removed snapshots, and physically deletes the difference: data files, delete files (position and equality), manifest files, and manifest lists. The decisive rule is **reachability, not age**. A data file written years ago that the current snapshot still lists is never deleted by expiry. Conversely a data file written five minutes ago can be deleted if the only snapshot that referenced it was just expired — which is exactly what happens to the pre-compaction files after `rewrite_data_files`. ## Arguments and defaults ```sql CALL prod.system.expire_snapshots( table => 'db.events', older_than => TIMESTAMP '2026-08-14 00:00:00', retain_last => 5 ) ``` `older_than` defaults to now minus `history.expire.max-snapshot-age-ms`, whose table default is 5 days. `retain_last` defaults to `history.expire.min-snapshots-to-keep`, whose default is 1, and acts as a **floor**: that many recent ancestors of the current snapshot are kept even if they are older than the cutoff. The procedure also accepts `snapshot_ids` to expire specific snapshots, `max_concurrent_deletes` to parallelise the file deletions, and `stream_results` so the driver does not collect the whole delete list. The Java equivalent is `table.expireSnapshots().expireOlderThan(ts).retainLast(n).commit()`. Setting the two `history.expire.*` properties on the table makes retention declarative, so a generic maintenance job needs no per-table arguments. Branches and tags protect what they reference. A snapshot held by a tag is retained under that ref's retention (`max-ref-age-ms`, and per-ref minimum snapshots and age), so tagging a month-end snapshot is the supported way to keep one specific point in time far beyond the general window. ## What it does not clean up - **Old `metadata.json` files.** They are governed by `write.metadata.delete-after-commit.enabled` (default `false`) and `write.metadata.previous-versions-max` (default 100). A high-frequency streaming table with the default off accumulates a metadata file per commit. - **Orphan files** — files under the table location that no metadata ever referenced, typically left by a failed or killed writer. Expiry only walks metadata, so it cannot see them; `remove_orphan_files` lists the storage location and diffs it against metadata. ## What expiry costs you Once a snapshot is expired, everything that depended on it is gone: time travel `AS OF` that snapshot id or a timestamp inside the expired range fails, `rollback_to_snapshot` to it fails, and an incremental read whose start snapshot has been expired fails. A reader that resolved an older snapshot before the job ran and is still scanning can hit missing files, which is why the retention window should comfortably exceed your longest-running query and your slowest downstream consumer's lag. Choose the window from real requirements — audit and recovery needs — rather than copying the 5-day default blindly. ## Operating it Inspect state through metadata tables: `SELECT * FROM db.events.snapshots` for what exists and when, `db.events.refs` for branches and tags, `db.events.metadata_log_entries` for metadata file history. The common surprise is "we compacted and storage went **up**" — compaction writes new files while the pre-compaction files stay reachable from the previous snapshots, so space is only reclaimed at expiry. That fixes the order of operations: rewrite first, then expire.
- We ran compaction and storage grew instead of shrinking. Why?Compaction writes new, larger data files and commits a new snapshot, but the previous snapshots still reference the small files they replaced, so both sets are on disk. Iceberg cannot delete the originals while any retained snapshot points at them. Space is reclaimed only when `expire_snapshots` removes those snapshots. That is why maintenance runs rewrite first, expiry second.
- How do you keep one specific snapshot far longer than the retention window?Tag it. `ALTER TABLE ... CREATE TAG` pins a snapshot id under a named ref, and expiry will not remove a snapshot that a branch or tag still references, subject to that ref's own retention settings such as `max-ref-age-ms`. This is the supported way to hold a month-end or pre-migration state for audit while everything else expires on the normal schedule.
- Does expire_snapshots remove old metadata.json files?No. Metadata file retention is separate: `write.metadata.delete-after-commit.enabled` is `false` by default, so previous `metadata.json` versions are kept indefinitely, and `write.metadata.previous-versions-max` (default 100) caps how many are retained once you turn deletion on. On a frequently committing table this is a real source of metadata bloat that snapshot expiry will never touch.
saying these in an interview costs you the question
- Says it deletes every data file older than the cutoff
- Confuses it with removing uncommitted orphan files
- Believes time travel still works after a snapshot expires
- Assumes compaction alone reclaims storage without expiry
- Thinks it also prunes old metadata.json versions