skip to content

What does an Apache Iceberg table store under its metadata/ and data/ directories?

level: juniorimportance: should knowfreq 62%

answer

  1. two halves: one you query, one that decides
  2. the directory listing is not the table
  3. JSON on top, Avro in the middle
  4. snap-*.avro sits above the manifests

basics

~20 s

An Iceberg table's data/ directory holds the immutable data files (Parquet by default, optionally ORC or Avro). Its metadata/ directory holds the JSON metadata files, the Avro manifest lists named snap-*.avro, and the Avro manifest files that describe those data files.

solid answer

~40 s

An Iceberg table is a location on a filesystem or object store with two halves. `data/` holds the immutable data files — Parquet by default, ORC or Avro if configured — usually written into partition-value subdirectories such as `ts_day=2026-05-01/`. `metadata/` holds the table layer: one JSON metadata file per commit (`v7.metadata.json` under a Hadoop catalog, or `00007-<uuid>.metadata.json` under Hive, Glue, JDBC or REST catalogs), one Avro **manifest list** per snapshot named `snap-<snapshot-id>-<attempt>-<uuid>.avro`, and the Avro **manifest** files (`<uuid>-m0.avro`) that enumerate data files with their partition values and column statistics. The crucial part: the directory listing does not define the table. Membership is decided purely by the manifests reachable from the current metadata file, so a stray Parquet file in `data/` is simply invisible to readers.

code

text · 12 lines
text
s3://lake/db/events/
  metadata/
    v1.metadata.json
    v2.metadata.json
    v3.metadata.json
    version-hint.text
    snap-3055729675574597004-1-4f2c9a1e.avro
    4f2c9a1e-m0.avro
  data/
    ts_day=2026-05-01/
      00000-3-9b7d1c2a.parquet
      00001-4-9b7d1c2a.parquet

go deeper

for a junior

Be ready to name the two directories and say what each holds: data files in data/, and JSON metadata plus Avro manifest lists and manifests in metadata/.

for a middle

Explain the file types by shape — .metadata.json, snap-*.avro, *-m0.avro — and why the manifests, not the directory listing, define table membership.

for a senior

Show the operational consequences: orphan files waste storage but are invisible; hand-deleting a data file breaks queries; layout can be relocated with write.data.path or object-storage mode without touching readers.

for a principal

Own the storage-layout policy — where metadata and data live, whether object-storage hashing is on, and how that interacts with bucket lifecycle rules that could silently delete files a manifest still references.

## The two halves of an Iceberg table Apache Iceberg is a **table format**: a specification for a set of files that, taken together, describe a table sitting on top of ordinary data files. Physically an Iceberg table is just a location — `s3://lake/db/events/`, `hdfs://.../events/`, or a local directory — containing two subdirectories by default: `data/` and `metadata/`. `data/` is the **data-file layer**. It contains the actual rows, written in a columnar or row format: Parquet by default, with ORC and Avro also supported. These files are **immutable** — Iceberg never edits a data file in place. Changing rows means writing new files (and, from format version 2 onward, small delete files that mark rows as removed). `metadata/` is the **table layer**. It contains three kinds of file, and knowing which is which is the whole point of this question. ## What lives in metadata/ **1. Metadata files (JSON).** Each commit writes a brand-new JSON metadata file describing the entire table state: `format-version`, `table-uuid`, `location`, the list of `schemas` with `current-schema-id`, the list of `partition-specs` with `default-spec-id`, `sort-orders`, table `properties`, the list of `snapshots`, `current-snapshot-id`, a `snapshot-log` and `metadata-log` of past states, and `refs` (named branches and tags). Naming depends on the catalog: a Hadoop catalog writes `v1.metadata.json`, `v2.metadata.json`, … plus a `version-hint.text`; Hive, Glue, JDBC and REST catalogs write names like `00007-3f9c1a2b-….metadata.json` and keep the pointer to the current one in the catalog itself. **2. Manifest lists (Avro).** One per snapshot, named `snap-<snapshot-id>-<attempt>-<uuid>.avro`. Each row describes one manifest file: its path and length, which partition spec it was written with, how many files it added/kept/deleted, and per-partition-field summaries (`contains_null`, `lower_bound`, `upper_bound`) used to skip whole manifests at planning time. **3. Manifest files (Avro).** Named like `4f2c9a1e-….-m0.avro`. Each row describes one data file (or, in v2+, one delete file): `file_path`, `file_format`, the partition tuple, `record_count`, `file_size_in_bytes`, and column statistics such as `value_counts`, `null_value_counts`, `lower_bounds` and `upper_bounds`. ## Why the split matters In a plain Hive-style table, "what files are in this table?" is answered by **listing directories**. That is slow on object storage, is not atomic, and gives no statistics. In Iceberg it is answered by **reading metadata**: catalog pointer → current metadata file → the current snapshot's manifest list → manifests → the exact set of data-file paths. Two consequences follow directly: - **A file in `data/` that no manifest references is not part of the table.** Copying a Parquet file into the directory by hand does nothing; such a file is an *orphan* and a maintenance job can later delete it. - **Deleting a data file from `data/` does not remove rows** — it breaks the table, because manifests still reference a path that no longer exists, and queries fail with a missing-file error. It also means partition directories are a **convenience, not a mechanism**. Iceberg prunes using partition values stored in manifests, not by parsing directory names, which is why the physical layout can change without breaking queries. ## Where the layout can differ The two-directory layout is the default, not a requirement: - `write.data.path` and `write.metadata.path` can point data and metadata at different prefixes or buckets. - `write.object-storage.enabled` inserts a hash component into data-file paths so that writes spread across object-store partitions instead of hammering one key prefix — the resulting paths look nothing like `ts_day=…/`. - Delete files (v2+) live alongside data files and are tracked by delete manifests, distinguished by a `content` field rather than by directory. ## Reading a listing Given a listing, identify files by shape: `*.metadata.json` is the table state; `snap-*.avro` is a manifest list (one per snapshot); other `*-m*.avro` files are manifests; everything under `data/` is data. If you can name those four things and say which one the catalog points at, you have answered the question. ## Common confusions The manifest list is Avro, not Parquet or JSON. There is one metadata JSON **per commit**, not one per table — old ones remain until a maintenance job cleans them. And `metadata/` holds far more than the schema: schema history, partition-spec history, every retained snapshot, and the pointers that make time travel possible.

  • What happens if you copy a Parquet file into the table's data/ directory by hand?
    Nothing — queries never see it. Iceberg builds its file set from the manifests reachable from the current metadata file, not from a directory listing. The file is an orphan, it consumes storage, and a later maintenance job that removes orphan files can delete it permanently.
  • Can an Iceberg table's data and metadata live in different buckets?
    Yes. `write.data.path` and `write.metadata.path` relocate each half independently, and `write.object-storage.enabled` adds a hash prefix to data-file paths so writes spread across object-store key ranges. Because file paths are recorded absolutely in manifests, readers follow them wherever they point.
  • Why does an Iceberg table keep several .metadata.json files instead of overwriting one?
    Each commit writes a new metadata file and the catalog swaps its pointer, which makes the commit atomic — no reader ever sees a half-written state. Older metadata files remain as history, listed in `metadata-log`, and are removed later by retention settings or maintenance jobs.

data/ is the warehouse shelves; metadata/ is the ledger that says which boxes are actually part of today's inventory. A box on the shelf that the ledger never mentions is not stock — it is clutter.

saying these in an interview costs you the question

  • Says listing data/ tells you which files are in the table
  • Thinks metadata/ contains only the table schema
  • Calls the manifest list a Parquet or JSON file
  • Believes deleting a file from data/ deletes those rows
  • Assumes partition directories are how Iceberg prunes files

context