After compaction commits, why do the old data files still occupy storage, and how are they removed?
answer
- the commit changes the list, not the storage
- a running query is still reading the old objects
- older retained states still point at them
- some files were never committed by anyone
- the cutoff must outlast your longest job
basics
~20 sCompaction only removes files from the table's current file list; the objects stay because older table states still reference them and in-flight readers are still reading them. A separate retention-aware cleanup deletes files that no retained state references, after a safety window.
solid answer
~50 sA compaction commit swaps the table's file list — old files out, new files in — but deleting the objects at that instant would break two things. Queries that already planned against the previous state are mid-scan and would hit missing-file errors, and any retained earlier state used for rollback or historical reads still points at those files. So removal is a **separate maintenance step** governed by a retention window: files that no retained state references and that are older than the window are deleted. Two distinct kinds of garbage exist — files that *were* part of the table and are now unreferenced by any retained state, and **orphan files** that were never committed at all, left by failed or aborted writes. The second kind is invisible to the table's metadata entirely and needs a listing-based sweep, which is dangerous if its cutoff is shorter than the longest in-flight write.
go deeper
Know that compaction removes files from the table's list but not from storage, and that a separate cleanup job deletes them later once nothing needs them.
Explain why immediate deletion would break in-flight readers and rollback, and distinguish files superseded by a newer state from orphan files that were never committed at all.
An interviewer expects you to derive a retention window from the longest query plus the recovery horizon, to explain why an orphan sweep needs a generous age threshold, and to sequence compaction before cleanup.
Own the maintenance contract for a platform: which jobs run where, how retention is set and defended against storage-cost pressure, and how silent cleanup failures or over-aggressive sweeps are detected before they become an incident.
## Two layers, two lifetimes A lakehouse table separates the **table layer** (what the table currently contains) from the **data-file layer** (objects in storage). A commit changes only the first. Removing a file from the current file set does not delete the object; it merely means new queries will not consider it. That separation is not an oversight — it is what makes several core features work: - **Reader isolation.** A query that started before the compaction commit planned against the old file set and holds a list of objects it intends to read. If those objects vanished at commit time, the query would fail partway with file-not-found. Because they persist, it completes correctly against a consistent older view. - **Rollback and historical reads.** Retained earlier table states reference the old files. Being able to read the table as of yesterday, or to undo a bad write, requires those files to exist. - **Commit atomicity.** The commit is a single metadata swap; making it also a bulk delete of thousands of objects would make it non-atomic and slow. ## Garbage type 1 — unreferenced files from superseded states After compaction, the small input files are still referenced by the *previous* table state. They become deletable only once that state is itself dropped from retention. The cleanup rule is: **delete an object when no retained table state references it.** This makes retention policy the real control knob. A long retention window means excellent time-travel and rollback ability, and a large storage bill — every compaction leaves its inputs behind for the whole window, and a heavily compacted table can hold several times its logical size. A short window frees storage quickly but destroys the ability to read or roll back to older states, and — the classic incident — must still be longer than the longest-running query, or in-flight readers get their files deleted underneath them. The practical failure mode reported over and over: someone shortens retention to save money, a nightly report that takes three hours is running, and the cleanup deletes files it is reading. Set the window from the maximum expected query duration plus the recovery horizon you actually want, not from the storage bill alone. ## Garbage type 2 — orphan files An **orphan** is an object in the table's storage location that **no table state has ever referenced**. It is not garbage from a superseded state; it was never committed at all. Sources: - A write job that produced data files and then failed, was killed, or lost its executor before committing. - A speculative or retried task whose duplicate output lost the race and was never included. - A compaction job that wrote its outputs and then hit a commit conflict and abandoned them. - Anything a process wrote into the table's path outside the format's own write path. Orphans are invisible to metadata-driven cleanup precisely because metadata has never heard of them. Finding them requires listing the storage location and subtracting everything any retained state references — an expensive, potentially very slow operation on a large table in an object store, and one that must be done carefully. The danger is a **cutoff that is too aggressive**. An orphan sweep must not delete files that a *currently running, not yet committed* write job just produced — those look exactly like orphans, because they are uncommitted. Hence orphan cleanup always uses an age threshold comfortably longer than the longest write job in the system. Running it with a short threshold while a large ingest is in flight deletes that job's output and it fails or, worse, commits a partial result. ## Metadata garbage Storage is not the only thing that accumulates. The table's own metadata grows with every commit — each one adds a new state and, depending on the format, additional metadata objects describing it. A table committing every minute accrues huge numbers of metadata entries, and reading the current state can require traversing or reconciling many of them. So maintenance has a third component: pruning old metadata alongside old data, and periodically consolidating metadata so that opening the table does not require replaying its whole history. A table can be perfectly compacted at the data layer and still open slowly because its metadata history was never trimmed. ## A workable maintenance policy Four jobs, distinct and differently scheduled: 1. **Compaction** — frequent, cheap, restricted to under-sized files in partitions that actually received writes. 2. **Clustering** — occasional, expensive, justified by measured pruning failure. 3. **Retention-based cleanup** — regular, deletes files no retained state references; window derived from longest query plus desired recovery horizon. 4. **Orphan sweep** — infrequent, listing-based, age threshold well beyond the longest write job. Ordering matters: compact first, then clean up, so cleanup has the newly superseded files in view. And every one of these should be observable — bytes reclaimed, files deleted, states expired — because a silently failing cleanup shows up months later as a storage bill, and an over-aggressive one shows up immediately as a broken query. ```text s3://lake/events/data/part-0001.parquet <- referenced by current state s3://lake/events/data/part-0002.parquet <- referenced only by an older state s3://lake/events/data/part-tmp-9f2c.parquet <- referenced by nothing: orphan ```
- What is the difference between an unreferenced file and an orphan file?An unreferenced file was once part of the table and became garbage when the state referencing it fell out of retention — metadata knows about it historically. An orphan was never committed to any state at all, usually left by a failed or aborted write, so metadata has no record of it and only a storage listing compared against all retained states can find it.
- How do you choose the retention window?Take the maximum realistic query duration as a hard floor — deleting files under a running reader is the classic outage — then extend it to whatever recovery horizon you actually want for rollback and historical reads. Storage cost sets the upper bound, but it should never push the window below that floor.
- Why is an orphan sweep with a short age threshold dangerous?Files written by a currently running, not-yet-committed job are indistinguishable from orphans, because neither is referenced by any state. A short threshold deletes that job's output mid-flight, so the job fails or produces an incomplete commit. The threshold must comfortably exceed the longest write job's duration.
- Can a table be well-compacted at the data layer and still slow to open?Yes. Metadata accumulates with every commit, so a frequently written table can have an enormous history to reconcile even when its data files are perfectly sized. Maintenance must prune and consolidate metadata as well as data, otherwise opening the table stays expensive.
Removing a book from the library catalogue does not shred it — it stays on the shelf until a separate weeding pass, and someone may still be reading it.
saying these in an interview costs you the question
- Assuming a compaction commit deletes the old objects immediately
- Setting the retention window shorter than the longest running query
- Running an orphan sweep while a large write job is in flight
- Treating orphan files as findable from table metadata alone
- Cleaning data files but never pruning accumulated metadata history