skip to content

What do Iceberg format versions 1, 2 and 3 each change about table metadata?

level: seniorimportance: should knowfreq 45%

answer

  1. the number lives in the table, not the engine
  2. files only, then rows, then vectors
  3. ordering commits is what makes deletes correct
  4. upgrades go one way only

basics

~20 s

Version 1 tracks data files only. Version 2 adds row-level delete files, marks manifests as data or delete content, and adds sequence numbers that order commits. Version 3 adds deletion vectors stored in Puffin blobs, row lineage, default column values and new data types.

solid answer

~50 s

The `format-version` field in an Iceberg table's metadata file selects the spec revision. **v1** describes analytic tables made of data files only: manifests list data files, and changing rows means rewriting whole files. **v2** adds row-level deletes: manifests can hold delete-file entries, the manifest list marks each manifest's `content` as data or deletes, and every snapshot gets a monotonically increasing `sequence-number` that manifest entries inherit — that ordering is what decides which delete files apply to which data files. **v3**, the newest revision, replaces position delete files with **deletion vectors** stored as Puffin blobs (one per data file), and adds row lineage, default column values for added columns, new types including nanosecond timestamps and variant, and multi-argument partition transforms. Upgrading is a property change and is one-way; downgrading is not supported. Readers and writers must support the version, so upgrade only once your whole engine fleet does.

code

sql · 2 lines
sql
-- one-way upgrade; verify every reader supports the version first
ALTER TABLE db.events SET TBLPROPERTIES ('format-version' = '2');

go deeper

for a junior

Know that the version number lives in the table's metadata file and that version 2 is where row-level deletes became possible.

for a middle

Explain what v2 added structurally — delete-file entries, manifest content markers, sequence numbers — and why ordering commits is required for deletes to be correct.

for a senior

Own the migration: the upgrade is one-way, gated on every reader supporting the version, and touches only metadata — so sequence readers first, then tables.

for a principal

Set the fleet policy: which spec version the platform standardizes on, how engine upgrades are sequenced ahead of table upgrades, and when v3 features justify the compatibility cost.

## Why the version exists An Iceberg table's metadata file opens with `"format-version": N`. That number is a contract: it tells every reader and writer which structures may appear in the metadata and how to interpret them. It is stored in the table itself, changed with a property update, and cannot be lowered — a table upgraded to a newer version is unreadable by engines that only implement an older one, which makes the upgrade a fleet-wide decision rather than a per-job one. ## Version 1 — data files only v1 defines the original structure: metadata file → manifest list → manifests → data files, with schemas, partition specs, sort orders and snapshot history. Manifest entries describe **data files only**. There is no way to mark individual rows as removed, so any delete or update is a **copy-on-write** operation: read the affected files, write new ones without the removed rows, and commit a snapshot that adds the new files and marks the old ones deleted. Correct, simple, and expensive when a change touches one row in a large file. ## Version 2 — row-level deletes and sequence numbers v2 is the version most production tables run and the one interviews focus on. It adds: - **Delete files.** A manifest entry's `data_file.content` distinguishes 0 (data), 1 (position deletes) and 2 (equality deletes). Position deletes name a file path and row positions; equality deletes name column values that identify removed rows. This is what makes *merge-on-read* possible: a delete can be recorded without rewriting the data file, and the reader applies it at scan time. - **Delete manifests.** The manifest list's `content` field marks a manifest as holding data entries or delete entries, so the planner can fetch deletes separately. - **Sequence numbers.** Each snapshot gets a `sequence-number`; manifest entries inherit it, and data files also carry a `file_sequence_number`. A delete file applies only to data files at or below its sequence number. Without this, a delete committed at time T would wrongly apply to rows appended at T+1, and concurrent append-plus-delete workloads could not be correct. v2 also tightened field requirements and made partition-value inheritance from the manifest possible, keeping manifests smaller. ## Version 3 — deletion vectors, lineage and richer types v3 is the current revision and refines v2's merge-on-read rather than replacing the model: - **Deletion vectors.** Instead of accumulating many small position-delete files per data file, v3 stores a single binary vector of deleted row positions per data file, held in a **Puffin** blob. One vector per data file bounds the read-side merge cost and removes the small-delete-file pileup that v2 tables accumulate. - **Row lineage.** Rows carry identity and update tracking through metadata fields, enabling change-tracking use cases that previously required rewriting or external bookkeeping. - **Default column values.** A newly added column can declare a default, so existing files need not be rewritten for the column to read as a value rather than null. - **New types**, including nanosecond-precision timestamps, a variant type for semi-structured data, geometry and geography types, and an unknown type. - **Multi-argument partition transforms**, widening what a partition spec can express. ## Upgrading in practice The upgrade is a table property change: `ALTER TABLE db.events SET TBLPROPERTIES ('format-version' = '2')`. It rewrites the metadata file with the new version and nothing else — existing data files, manifests and snapshots are untouched, because both versions describe the same data-file layer. What changes is what *future* writes may produce. The risk is entirely on the read side. If any engine in the estate — an older Trino, an older Spark runtime, a downstream tool with a pinned Iceberg library — does not implement the version, it will refuse the table or, worse, must be prevented from reading it at all. So the sequence is: upgrade the readers, verify, then upgrade the tables. The same care applies to v3, which is newer and less universally implemented than v2. ## What to say when asked Name the axis rather than reciting a changelog: **v1 = files only; v2 = row-level deletes plus the sequence numbers that make them correct; v3 = deletion vectors, lineage, defaults and new types.** Then add the operational point that the version lives in the table, is one-way, and gates which engines can read it. That combination — the mechanism plus the migration risk — is what a senior answer sounds like. ## Common mistakes Saying v2 "added ACID" (v1 already had atomic snapshot commits). Confusing Iceberg's format version with the write-mode setting that chooses copy-on-write versus merge-on-read behaviour for an operation — the format version says what *may* be written, the write mode says what a given operation *does*. And assuming an upgrade rewrites data: it does not.

  • Why did row-level deletes require sequence numbers rather than timestamps?
    A delete must apply to data written before it and not to data written after, and that ordering has to be exact and machine-independent. Sequence numbers increase monotonically per snapshot and are inherited by manifest entries, so the comparison is unambiguous. Wall-clock timestamps across writers would be skewed and could apply a delete to rows that arrived later.
  • Does upgrading a table's format version rewrite any data?
    No. It rewrites the metadata file with the new `format-version` and leaves data files, manifests and snapshot history untouched, because both versions describe the same data-file layer. Only subsequent writes may use the new structures. The real cost is compatibility: every reader must support the new version first.
  • How do deletion vectors differ from the position delete files they replace?
    A position delete file is one more small file per delete operation, so a hot table accumulates many per data file and the reader merges them all at scan time. A deletion vector is a single binary vector of deleted positions per data file, stored in a Puffin blob and replaced on update, which bounds both file count and merge cost.

saying these in an interview costs you the question

  • Says format version 2 is what introduced ACID commits
  • Confuses the format version with an engine or library version
  • Claims upgrading the format version rewrites the data files
  • Thinks the format version can be downgraded if needed
  • Equates the format version with a copy-on-write write-mode setting

context