When an Apache Iceberg append adds one file, which metadata files are written?
answer
- the root moves, the leaves do not
- new files at the top, references below
- one branch rewritten, the rest re-pointed
- the parent's manifests are carried by path
basics
~20 sAn append writes the data file, a new manifest listing it, a new manifest list that references that manifest plus the parent snapshot's still-valid manifests, and a new JSON metadata file carrying the new snapshot. Existing manifests and data files are reused untouched.
solid answer
~50 sOnly the top of the metadata tree is rewritten. The writer produces the data file, then a **new manifest** (`<uuid>-m0.avro`) whose single entry has `status = ADDED` and carries the file's partition tuple and column bounds. It then writes a **new manifest list** (`snap-<new-snapshot-id>-<attempt>-<uuid>.avro`) that references the new manifest *plus every manifest the parent snapshot still needs* — those older manifests are reused by reference, not rewritten. Finally it writes a **new metadata JSON** containing the new snapshot appended to `snapshots`, with `current-snapshot-id`, `snapshot-log` and `metadata-log` updated. The commit itself is the catalog swapping its pointer to that new metadata file. Nothing older is mutated, which is exactly why old snapshots remain readable for time travel. One nuance: by default Iceberg may merge small manifests during commit (`commit.manifest-merge.enabled`), so a busy table can rewrite a few manifests rather than none.
code
text · 11 lines# before the append
metadata/00006-a1b2c3d4.metadata.json <- catalog points here
metadata/snap-8712449321005630891-1-77aa.avro
metadata/77aa-m0.avro
# after appending one data file
metadata/00007-3f9c1a2b.metadata.json <- catalog now points here
metadata/snap-3055729675574597004-1-4f2c.avro (references 77aa-m0 AND 4f2c-m0)
metadata/4f2c-m0.avro (new: the one added file)
metadata/77aa-m0.avro (unchanged, reused by reference)
data/ts_day=2026-05-01/00000-7-b41c9e2f.parquetgo deeper
Remember the direction: a commit writes new files at the top of the metadata tree and leaves existing data files and manifests alone.
Name the three metadata files a commit produces — manifest, manifest list, metadata JSON — and explain that the manifest list re-references the parent's manifests by path.
Connect the shape to operations: commit cost tracks change size, small-manifest accumulation is a real planning tax, and fast appends trade planning speed for cheaper commits.
Own the write-path policy across a platform — commit cadence, whether streaming writers use fast appends, and how manifest consolidation is scheduled so planning latency stays bounded.
## What an append actually produces Iceberg metadata is a tree of immutable files, and a commit is a **copy-on-write of the path from the root to the change**. For an append of one data file, the writer produces, in order: 1. **The data file** — e.g. `data/ts_day=2026-05-01/00000-7-b41c9e2f.parquet`, written and closed before any metadata is touched. 2. **A manifest** — `b41c9e2f-m0.avro`, one `manifest_entry` with `status = 1` (ADDED), `snapshot_id` set to the new snapshot, and a `data_file` struct carrying `file_path`, `file_format`, the `partition` tuple, `record_count`, `file_size_in_bytes` and the column stats (`value_counts`, `null_value_counts`, `lower_bounds`, `upper_bounds`). 3. **A manifest list** — `snap-<new-snapshot-id>-1-<uuid>.avro`. This is the file people misjudge: it lists the *new* manifest **and re-references the parent snapshot's manifests by path**. Those older manifests are not read, not copied, not rewritten — only their paths and summary rows are carried forward. 4. **A metadata file** — a fresh `*.metadata.json` identical to the previous one except that `snapshots` gains an entry (with `snapshot-id`, `parent-snapshot-id`, `sequence-number`, `timestamp-ms`, `schema-id`, `manifest-list`, and a `summary` map recording `operation: append` plus added/total counts), `current-snapshot-id` moves, and `snapshot-log` and `metadata-log` gain rows. 5. **The commit** — the catalog is asked to move the table's current-metadata pointer from the base metadata file to the new one, atomically. Until that swap succeeds the new files exist but are unreachable; if it fails, they are orphans and no reader ever saw them. ## What is *not* written No existing data file is touched. No existing manifest is modified — Avro manifests are write-once. The previous metadata JSON, manifest list and manifests all remain on storage and remain valid, which is precisely what makes the parent snapshot still readable. Time travel is not a replay mechanism; it is the fact that the older root is still intact and still points at intact children. This structure explains the cost model: **a commit's metadata write is proportional to what changed, not to table size.** Appending one file to a table with a million files writes one small manifest and one manifest list whose row count equals the number of manifests — not the number of files. ## The manifest-merge nuance If every append wrote a brand-new manifest and nothing ever consolidated them, a table committing every minute would accumulate thousands of tiny manifests and planning would slow down. Iceberg counters this at commit time: `commit.manifest-merge.enabled` (on by default for merge appends) lets a commit combine small manifests into larger ones, governed by `commit.manifest.target-size-bytes` and `commit.manifest.min-count-to-merge`. So the honest answer is "a new manifest, a new manifest list and a new metadata file — and possibly some small manifests merged". Streaming writers that want the cheapest possible commit use a **fast append**, which skips merging entirely and accepts more manifests, leaving consolidation to a later maintenance job. ## Sequence numbers From format version 2, each snapshot receives a monotonically increasing `sequence-number`, and manifest entries inherit it. This is not decoration: it orders commits and determines which delete files apply to which data files. An appended data file with a high sequence number is unaffected by delete files committed earlier, which is what makes concurrent append-plus-delete workloads correct. ## Other operation types - **Overwrite / delete / MERGE** follow the same shape, but the new manifest contains entries with `status = 2` (DELETED) for removed files, and in v2+ may add a **delete manifest** (`content = 1` in the manifest list) referencing position or equality delete files. - **Schema or partition-spec change** may write only a new metadata JSON — a new schema in `schemas` or a new spec in `partition-specs`, with no snapshot at all, because no data changed. - **Rewrites (compaction)** produce a snapshot that adds new files and marks old ones deleted; the old files stay on storage until snapshot expiry releases them. ## The takeaway to say out loud "A commit rewrites the root and one branch of the metadata tree, and re-references everything else." That single sentence carries the atomicity story (one pointer swap), the time-travel story (old roots intact), the performance story (write cost tracks change size), and the storage story (old files linger until a retention job removes them).
- Why does the new manifest list re-reference the parent snapshot's manifests instead of copying their contents?Because manifests are immutable and already describe files that have not changed. Referencing them by path keeps commit cost proportional to the change rather than to table size, and lets multiple snapshots share the same manifest file — which is also why expiring one snapshot cannot blindly delete a manifest another still references.
- What happens to the files a commit wrote if the catalog pointer swap fails?They stay on storage but are unreachable — no metadata file references them, so no reader can see them. The writer retries against the refreshed base and writes a fresh set. The abandoned files become orphans that a later orphan-file cleanup can remove.
- Does every Iceberg commit create a snapshot?No. Changes that touch only table state — a schema addition, a new partition spec, a property change — write a new metadata JSON without a new snapshot, because no data files changed. Only operations that alter the set of data or delete files produce a snapshot entry.
saying these in an interview costs you the question
- Says the whole metadata tree is rewritten on every commit
- Thinks the new manifest list copies old manifests' contents
- Believes older data files are deleted at commit time
- Claims a commit rewrites the previous metadata.json in place
- Assumes every commit produces exactly one new manifest, always