In Apache Iceberg, how does a reader reach the data files a query must scan?
answer
- one pointer at the top, files at the bottom
- four hops, and no directory listing
- the snapshot names an Avro file, not data files
- manifest list prunes manifests, manifests prune files
basics
~20 sThe catalog names the current metadata file. That JSON gives the current snapshot, whose manifest-list Avro file names the manifests. Each manifest lists data files with partition values and column bounds, so the reader prunes manifests then files and scans only survivors.
solid answer
~50 sIt is a four-hop chain, and no directory listing happens anywhere in it. The **catalog** holds one pointer: the path of the current `metadata.json`. That JSON contains the schemas, partition specs, the `snapshots` array and `current-snapshot-id`; the reader picks the snapshot (or an older one for time travel) and takes its `manifest-list` path — an Avro file named `snap-<snapshot-id>-<attempt>-<uuid>.avro`. Each row of that manifest list describes one **manifest**, including per-partition-field summaries, so manifests whose partition ranges cannot match the predicate are skipped without being opened. The surviving manifests are read, and each of their rows describes one data file: its path, partition tuple, `record_count`, and `lower_bounds`/`upper_bounds` per column. Files that cannot match are pruned, and the rest become scan tasks. In format version 2 and later the same chain also yields delete files, tracked in manifests whose `content` marks them as deletes.
code
json · 27 lines{
"format-version": 2,
"table-uuid": "9f1e6a80-3d1b-4a52-9f0c-1c2b3d4e5f60",
"location": "s3://lake/db/events",
"current-schema-id": 1,
"default-spec-id": 0,
"current-snapshot-id": 3055729675574597004,
"snapshots": [
{
"snapshot-id": 3055729675574597004,
"parent-snapshot-id": 8712449321005630891,
"sequence-number": 42,
"timestamp-ms": 1746100000000,
"schema-id": 1,
"manifest-list": "s3://lake/db/events/metadata/snap-3055729675574597004-1-4f2c9a1e.avro",
"summary": {
"operation": "append",
"added-data-files": "3",
"added-records": "120000",
"total-data-files": "1841"
}
}
],
"refs": {
"main": { "snapshot-id": 3055729675574597004, "type": "branch" }
}
}go deeper
Memorize the order: catalog, metadata JSON, manifest list, manifests, data files. Being able to recite the chain in the right direction is most of the credit at this level.
Explain what each hop contributes — snapshot selection at the metadata file, manifest pruning from partition summaries, file pruning from per-column bounds — and why no directory listing occurs.
Use the chain diagnostically: when planning is slow or a query reads too many files, walk the snapshots, manifests and files metadata tables in that order and point at the hop that failed to prune.
Own the consequences at scale — planning cost tracks manifest count, so ingest cadence, manifest sizing and snapshot retention are platform decisions, not per-job ones.
## The chain, top to bottom Iceberg's read path is a fixed sequence of pointer follows. Learn it as a chain and most other Iceberg questions answer themselves. **Hop 0 — the catalog.** The catalog (Hive Metastore, AWS Glue, a REST catalog, JDBC, Nessie, or a Hadoop directory with `version-hint.text`) stores exactly one durable fact about the table: the path of the **current metadata file**. Nothing else about the table lives there. This single pointer is what makes commits atomic — swapping it is the commit. **Hop 1 — the metadata file (JSON).** It describes the entire table: `format-version`, `table-uuid`, `location`, `schemas` with `current-schema-id`, `partition-specs` with `default-spec-id`, `sort-orders`, `properties`, `refs` (branches and tags), the `snapshot-log` and `metadata-log` histories, and the `snapshots` array with `current-snapshot-id`. Each snapshot entry carries `snapshot-id`, `parent-snapshot-id`, `timestamp-ms`, `schema-id`, a `summary` map (operation, added/total counts), and — the important field — `manifest-list`, an absolute path. The reader chooses a snapshot here. Normally `current-snapshot-id`; for `FOR SYSTEM_VERSION AS OF <id>` or `FOR SYSTEM_TIME AS OF <ts>` it picks a historical entry from the same array. **Time travel is just picking a different row of `snapshots`** — nothing else in the chain changes. **Hop 2 — the manifest list (Avro, `snap-*.avro`).** One per snapshot. Each row is a `manifest_file` record: `manifest_path`, `manifest_length`, `partition_spec_id`, `added_snapshot_id`, counts of added/existing/deleted files, and `partitions` — an array of field summaries (`contains_null`, `contains_nan`, `lower_bound`, `upper_bound`) *per partition field*. From format version 2 the record also has `content` (0 = data manifest, 1 = delete manifest) and `sequence_number` / `min_sequence_number`. Those partition summaries are the first pruning step: a manifest whose `ts_day` range is 2026-01-01…2026-01-31 is skipped entirely for a query filtering on March, without opening the file. **Hop 3 — manifests (Avro).** Each row is a `manifest_entry`: `status` (0 EXISTING, 1 ADDED, 2 DELETED), `snapshot_id`, sequence numbers, and a nested `data_file` struct with `content` (data / position deletes / equality deletes), `file_path`, `file_format`, the `partition` tuple, `record_count`, `file_size_in_bytes`, `column_sizes`, `value_counts`, `null_value_counts`, `nan_value_counts`, `lower_bounds`, `upper_bounds` and `split_offsets`. This is the second pruning step, and the precise one: partition tuples eliminate whole partitions, and per-column min/max bounds eliminate individual files even inside a matching partition. **Hop 4 — data files.** What survives becomes split-level scan tasks handed to the engine, paired with any delete files that apply. ## Why this beats listing A Hive-style reader lists directories, then opens file footers to prune. On object storage, listing is slow, eventually consistent in some stores, and never atomic — a concurrent writer can make a listing self-contradictory. Iceberg replaces listing with reading a bounded set of metadata files whose paths are all recorded absolutely. Planning cost scales with the number of *manifests*, not with the number of files or objects in a prefix, and the whole plan is derived from **one immutable snapshot**, so it is consistent by construction. ## Where deletes enter (v2 and later) Format version 2 introduced row-level delete files. They are tracked by the same chain: delete manifests are marked `content = 1` in the manifest list, and their entries describe delete files rather than data files. A snapshot's **sequence number** decides applicability — a delete file only applies to data files with a lower or equal sequence number — which is how Iceberg avoids re-deleting rows in files written after the delete. Format version 3 replaces position delete files with **deletion vectors** stored in Puffin blobs, but the chain that finds them is unchanged. ## Practical inspection Every hop is queryable through metadata tables: `db.tbl.snapshots`, `db.tbl.manifests`, `db.tbl.files`, `db.tbl.entries`, `db.tbl.history`, `db.tbl.refs`. When someone asks "why did this query read 4,000 files?", the answer is found by walking those tables in the same order as the reader. ## Common mistakes Saying the catalog stores the file list (it stores one pointer). Saying the manifest list contains data-file paths (it contains *manifest* paths). Believing planning lists directories (it never does). And believing time travel replays a log — it does not; it reads a different snapshot's manifest list, which is why an expired snapshot's time travel fails immediately rather than degrading.
- Where exactly does time travel enter this chain?At the metadata file. The reader picks a different entry from the `snapshots` array — by `snapshot-id` or by the nearest `timestamp-ms` — and follows that entry's `manifest-list` instead of the current one. Every hop below is identical, which is why time travel costs nothing extra and why it fails outright once a snapshot has been expired.
- Why does planning cost scale with manifest count rather than file count?The reader opens the manifest list, then only the manifests whose partition summaries can match the predicate. Files inside skipped manifests are never examined. That is why a table with many tiny manifests plans slowly even when the data volume is modest, and why manifest rewriting is a real tuning lever.
- In a format-version-2 table, how do delete files reach the reader through this chain?They travel the same path. The manifest list marks each manifest with `content` — 0 for data, 1 for deletes — and delete manifests' entries describe delete files. A delete file applies only to data files whose sequence number is at or below its own, so newer data written after the delete is unaffected.
- What does the reader do when the current snapshot's schema differs from the table's current schema?Each snapshot entry records `schema-id`, so the reader resolves the data using the schema that snapshot was written with and projects it onto the requested schema by field ID. That is how a time-travel read of an older snapshot returns sensible columns even after the table schema has changed.
saying these in an interview costs you the question
- Claims the catalog stores the table's list of data files
- Says the manifest list contains data-file paths directly
- Thinks planning lists the data directory to find files
- Describes time travel as replaying a change log
- Cannot distinguish the manifest list from a manifest