How does an open table format make a multi-file write atomic on object storage?
answer
- object stores give you exactly one atomic unit
- the new files exist but nothing points at them yet
- visibility hangs on a single object write
- one conditional swap of the current pointer
- before it: old table; after it: new table
basics
~20 sIt writes all new data files first, where nothing references them, then writes new table metadata describing the resulting snapshot, and finally swaps a single pointer to that metadata in one atomic operation. Readers see the old state or the new one, never a partial write.
solid answer
~50 sObject storage gives you one atomic unit: a single object write. There is no multi-file rename and no cross-object transaction, so a format cannot make ten files appear together by any storage primitive. It sidesteps the problem by making visibility depend on **one** object. The writer uploads its new data files, which are inert because no snapshot references them yet, then writes new table metadata listing the resulting file set, then performs a **single atomic swap** of the table's current-state pointer from the old metadata to the new. That swap is the commit: before it, readers see the previous snapshot; after it, they see the new one. There is no in-between. If the writer dies at any earlier point, the uploaded files are simply unreferenced garbage that an orphan-file cleanup removes later — the table was never wrong, only slightly wasteful. The swap must be atomic and give a single source of truth for "what is current," which is exactly the guarantee the catalog layer exists to provide.
code
text · 8 linest0 current -> meta_007 table = f1.parquet, f2.parquet
t1 PUT data/f3.parquet inert: no snapshot references it
t2 PUT data/f4.parquet inert
t3 PUT metadata/meta_008 candidate: lists f1, f2, f3, f4
t4 CAS current: 007 -> 008 <- commit; all four files visible at once
-- crash between t1 and t4:
-- readers still see f1, f2; f3 and f4 are orphans awaiting cleanupgo deeper
Know that the write becomes visible all at once, at the end, and that data files are written before anything points at them. Be able to say a half-finished write is invisible rather than half-applied.
Explain the three phases and name the primitive: one conditional single-object write. An interviewer expects you to say why object storage forces this design and what orphan files are.
Show that you have operated it: orphan cleanup thresholds versus long writes, why listing is never the source of truth, and what breaks when two authorities can both claim to hold the current pointer.
Own the boundary between storage and catalog — which component is the single authority for the current pointer, what happens to multi-engine writes if that is ambiguous, and how that constrains platform-wide engine choice.
## Why this is hard on object storage A POSIX filesystem gives you a cheap atomic multi-file trick: write into a staging directory and rename it. Object stores do not have that. Their guarantees are per object — a PUT of one object either happens or does not — and there is no directory rename, no multi-object transaction, and no way to make several uploads become visible at the same instant. Historically listings were also only eventually consistent, so even *enumerating* what you just wrote was not reliable. A table format therefore never defines the table as "the files under this prefix." It defines the table as "the file set named by the current table metadata," and it reduces the whole atomicity problem to changing **one** thing. ## The commit protocol Every open table format follows the same three-phase shape, whatever it names the pieces: 1. **Write data, invisibly.** The writer uploads new data files under the table's storage location. Nothing points at them, so no reader can see them. Upload order, retries and partial failures are all harmless at this stage. 2. **Write new table metadata.** The writer produces a new metadata object describing the snapshot that results from its operation: the full resulting file set (or a delta plus a base the reader can resolve), the schema, the partitioning, and statistics. This object is also still not the table — it is a candidate. 3. **Swap the pointer, atomically.** One operation makes the candidate current. Concretely this is either a **compare-and-swap in a catalog** ("set the table's current metadata location to M8, but only if it is still M7") or an **atomic create-if-absent of the next numbered commit object** ("create commit 8; fail if it already exists"). Either way the primitive is a single-object conditional write. The instant step 3 succeeds, every subsequent reader resolves the table through the new pointer and sees all of the new files at once. A reader that resolved the pointer a millisecond earlier reads the old snapshot end to end. Both are consistent table states. ```text t0 current -> meta_007 (table = f1.parquet, f2.parquet) t1 writer PUTs f3.parquet (inert: no snapshot references it) t2 writer PUTs meta_008 (candidate: lists f1, f2, f3) t3 ATOMIC: current 007 -> 008 (the entire write becomes visible here) ``` ## Why crashes are safe A writer that dies before step 3 leaves uploaded objects that no snapshot references. Readers are unaffected because readers never list the directory — they read the file list out of metadata. The debris costs storage until an orphan-file cleanup removes objects that are unreferenced *and* older than a safety threshold. That threshold must exceed the longest possible in-flight write, or cleanup will delete files a live writer is about to commit. ## Why the pointer needs help Some object stores now offer conditional writes, and some formats can commit directly against storage. Where that primitive is missing or where multiple engines write the same table, the atomic swap is delegated to a catalog that can do a compare-and-swap and is the single authority on the current pointer. Two catalogs disagreeing about which metadata is current is not a slow table, it is a corrupt one — which is why "one source of truth for the current state, swapped atomically" is the requirement a catalog exists to meet. ## What this buys beyond atomicity - **Isolation.** Readers resolve the pointer once and then read an immutable file set, so a commit landing mid-query cannot change what that query sees. - **Durability of history.** The old metadata object still exists and still references live files, which is exactly what time travel and rollback read. - **Concurrency control.** Because the swap is conditional on the pointer's current value, two writers racing cannot both win; the loser learns it lost and retries. - **No listing dependence.** Scan planning reads a written file list instead of enumerating storage, which is both faster on large tables and immune to stale listings. ## The failure mode to name in an interview The classic broken alternative is "write files, then overwrite a manifest of the directory in place." That has a window where the manifest is half written or where a second writer clobbers the first, and it is exactly what the conditional single-object swap removes. If you can state that the commit is atomic because *one object write is atomic and everything else is invisible until it happens*, you have answered the question.
- What happens to the files a writer uploaded if it crashes before committing?They stay in storage, referenced by nothing, and are invisible to every reader because readers resolve the file list from metadata rather than by listing the directory. An orphan-file cleanup deletes unreferenced objects older than a safety threshold. That threshold must be larger than the longest in-flight write, or cleanup can delete a live writer's staged files.
- Why don't readers just list the storage prefix to find the table's files?Because the directory contains staged files from in-flight writes, orphans from failed ones, and files that older snapshots still need. Only metadata knows which files belong to which committed state. Listing is also slow on very large tables and historically not consistent enough to trust immediately after a write.
- Where does the atomic swap actually happen?Either in a catalog that supports a compare-and-swap on the table's current metadata pointer, or directly in storage via an atomic create-if-absent of the next commit object where the store offers that primitive. The requirement is the same in both cases: exactly one authority, and a conditional write so a racing writer cannot silently overwrite.
- Does the writer hold a lock while it writes its data files?No, and that is the point. Data files are written with no coordination at all, because they are invisible until commit. Contention exists only for the final swap, which is a single fast conditional operation, so long writes do not block readers or other writers while they run.
Like publishing a new edition of a catalogue: the printed pages can pile up in the back room for hours, but the shop is selling the old edition until the single moment someone changes which edition is on the counter.
saying these in an interview costs you the question
- Says the format renames a staging directory at the end
- Claims a distributed lock is held across all file uploads
- Thinks a partially uploaded write is visible until cleanup
- Says readers list the directory to find the table's files
- Believes the commit rewrites existing data files in place