skip to content

What is the difference between a file format like Parquet and a table format like Iceberg?

level: juniorimportance: must knowfreq 85%

answer

  1. one file knows nothing about its neighbours
  2. bytes inside a file versus the file list
  3. which layer knows the table's current version
  4. one describes a file, one describes a table

basics

~20 s

A file format defines how bytes are laid out inside one immutable file. A table format is a metadata layer above a set of such files that records which files, which schema and which version make up the table right now.

solid answer

~50 s

A **file format** — Parquet, ORC, Avro — specifies how a single file encodes its rows and columns: schema, encoding, compression and its own internal statistics. It knows nothing about any other file. A **table format** sits above a set of those files and answers the question no file can answer: *which files are this table, right now*. It keeps an authoritative, versioned list of data files plus the table's current schema and partitioning, so a reader gets one consistent set of files rather than whatever storage happens to contain at that moment, and a writer can add and remove files in a single atomic commit. Parquet gives you efficient bytes; the table layer gives you a table — atomic commits, schema evolution, time travel and safe concurrent writes. The two are largely orthogonal: a table-format table is usually *made of* Parquet files.

code

text · 6 lines
text
s3://lake/warehouse/orders/
  data/00000-0-9a1f.parquet      # data-file layer: immutable bytes
  data/00001-0-3c7d.parquet
  data/00002-0-b40e.parquet
  metadata/                      # table layer: which files are the table now
    <table metadata + per-version file lists>

go deeper

for a junior

Be ready to state the split in one breath: Parquet describes the inside of one file, a table format describes which files make up the table. Knowing that a table-format table is made of Parquet files is the point.

for a middle

Explain the mechanics the table layer adds — an explicit versioned file list, atomic add/remove commits, table-level schema, and per-file statistics — and why a directory listing cannot provide any of them.

for a senior

Show you reason about the boundary in production: which failures belong to the file layer (corrupt file, bad encoding, tiny files) and which belong to the table layer (a commit conflict, a missing snapshot, files present on storage but absent from metadata).

for a principal

Own the platform consequence: the file format and table format are chosen independently, so standardizing the table layer across teams buys interoperability and governance without forcing anyone to re-encode petabytes of existing data.

## Two layers, two jobs Every lakehouse table is built from two independent layers. The bottom layer is the **data-file layer**: a set of immutable files in object storage or HDFS, normally Parquet, sometimes ORC or Avro. Each file is self-describing — it carries its own schema and its own encoded, compressed data — and once written it is never edited in place. The top layer is the **table layer**: metadata written by a table format such as Apache Iceberg, Delta Lake or Apache Hudi. It records which data files constitute the table at each version, what the table's current schema and partitioning are, and what each commit added or removed. ## What the file format owns A file format decides how one file stores data: how columns are laid out and encoded, which compression codec is used, how the file records its own schema, and how it can be split so several tasks read different chunks in parallel. It also stores statistics about its own contents so a reader can skip parts of the file. What a file cannot know is everything *about the table*: whether sibling files exist, whether it is still part of the current version of the table, whether some of its rows have since been logically deleted, or what the table looked like yesterday. A file is a box of rows; it has no opinion about the collection it belongs to. ## What the table layer adds On top of that file set, a table format supplies: - **An authoritative file list per version.** Readers resolve the file set from metadata instead of listing a directory. - **Atomic commits.** Adding new files and removing replaced ones happens as one version bump, so a reader never sees half a job's output. - **Snapshot isolation and time travel.** Because old versions still reference the files they used, a query can read a consistent past state. - **Schema and partition evolution.** The table's schema lives in metadata, so a column can be added, renamed or dropped without rewriting existing files. - **File-level statistics for planning.** Row counts and per-column bounds per file let the planner skip whole files before opening any of them. - **Concurrency control.** Two writers committing at once are serialized by the metadata layer rather than racing on a directory. ## Why "a folder of Parquet files" is not a table The pre-lakehouse arrangement was a Hive-style table: a metastore entry pointing at a directory, where the table's contents were simply whatever files that directory held at read time. That is fragile. A job writing a hundred files makes its partial output visible file by file. A rewrite deletes files a running query is still reading. Listing millions of objects is slow and expensive. Nothing enforces that all the files agree on a schema. The table layer exists precisely to replace "whatever is in the folder" with "exactly this list of files, as of this version". ## Choosing a format for each layer The layers are chosen separately. Parquet and ORC are columnar and suit analytical scans that touch a few columns of very many rows. Avro is row-oriented and suits write-heavy, whole-row, append-style workloads. Iceberg can store data files as Parquet, ORC or Avro; Delta Lake's data files are Parquet. So "we use Parquet" and "we use Iceberg" are not competing statements — the usual answer is both. ## The confusion to avoid The classic interview stumble is treating a table format as a replacement for Parquet, or calling Iceberg "a file format". It is not one, and it does not re-encode your bytes. It writes metadata that points at the Parquet files you already have. That is also why adopting a table format over an existing lake can often be done without rewriting the data at all: only the table layer is new. ## How to say it in an interview "Parquet describes one file. Iceberg or Delta describe the table those files belong to — the file list, the schema, the version history, and the atomic commit that moves the table from one file set to the next." Everything else in the lakehouse — time travel, partition evolution, safe concurrent writes, compaction — is a consequence of that split.

  • If a table format does not change how bytes are stored, what does it actually write to storage?
    Metadata files: a record of the table's current schema, partitioning and properties, and a list of the data files that belong to each version, usually with per-file row counts and column bounds. It also writes a version history so old file sets remain resolvable for time travel. The data files themselves are ordinary Parquet or ORC files, unchanged.
  • Can two different table formats point at the same Parquet files at once?
    Physically yes — nothing stops two metadata layers from referencing the same paths, and some migration tooling relies on that briefly. Operationally it is dangerous: each format believes it owns the file set, so one format's maintenance job can delete files the other still references, and neither sees the other's commits. Treat one format as the owner and cut over cleanly.
  • Does adopting a table format require rewriting existing Parquet files?
    Usually not. Because the table layer only references file paths, an existing set of Parquet files can be adopted by writing metadata over it, typically reading each file's footer to collect statistics. You inherit the existing file sizes and layout, so a rewrite is still worth doing if the files are small or badly clustered — but it is a separate, optional step.

Parquet is a book: self-contained, readable on its own. The table format is the library catalogue: it says which books are on the shelf today, which edition is current, and what was checked in or out yesterday.

saying these in an interview costs you the question

  • Calling Iceberg or Delta Lake a file format that replaces Parquet
  • Saying the table format re-encodes or compresses the data files
  • Believing the engine lists the directory to find a table's files
  • Claiming Parquet files themselves provide ACID transactions
  • Treating schema evolution as a property of the file rather than the table

context