skip to content

Table Format Concepts

You will learn the format-agnostic ideas every lakehouse table format implements: what a metadata layer buys you over a directory of files, how partitioning and its evolution work, how snapshots give ACID and time travel over object storage, why small files ruin performance, and what a catalog actually does. Interviewers use these to check whether you can compare Iceberg, Delta and Hudi on mechanism instead of on marketing.

on this pageshow

explore

questions

29

Why does a query engine need a catalog to read a lakehouse table instead of a path?

level: juniorimportance: must knowfreq 60%

answer

  1. a folder of files cannot say which files count
  2. several metadata versions exist at once
  3. something must name the current one
  4. one small pointer per table name
  5. name resolution plus an atomic pointer

basics

~20 s

A catalog maps a table name to the location of that table's current metadata, so every engine agrees on which version is live. A storage path only lists files, including leftovers from failed or in-flight writes that no commit includes.

solid answer

~40 s

An open table format keeps data as ordinary files in object storage plus a metadata layer that records which files, schema and partitioning make up the table **right now**. Because each commit writes new metadata rather than editing the old in place, several metadata versions sit in the directory at once, and nothing in a file listing says which one is current — a listing also shows files from crashed jobs and writes still in flight. The catalog is the transactional service that stores, per `namespace.table`, one value: the pointer to the current metadata. An engine resolves `analytics.orders` through the catalog, follows the pointer, and plans its scan from that metadata. That single indirection is also what makes multi-engine reads consistent and what gives you a place to hang grants, audit and discovery.

go deeper

for a junior

Be able to say in one breath that the catalog turns a table name into the current metadata location, and that a bare storage path cannot tell committed files from leftovers.

for a middle

Explain why several metadata versions coexist in storage at once and why the pointer therefore has to live somewhere transactional outside those files.

for a senior

Show that you treat the catalog as the correctness boundary for multi-engine access, and can name what breaks operationally when engines resolve a table by path instead.

for a principal

Frame the catalog as the platform's interoperability and governance contract: one source of truth for table identity, grants and audit, chosen before engines are.

## A directory of files is not a table Open table formats store table data as ordinary files — usually Parquet — in object storage, and add a **metadata layer** above them. That metadata layer records the schema, the partitioning scheme, and above all the exact set of data files that belong to the table at this instant. It is the thing that distinguishes "these 812 files" from "those 40 leftovers a crashed job wrote at 03:12". But the metadata layer inherits the same problem one level up. Object storage has no safe in-place edit, so a commit does not modify the existing metadata: it writes a **new** metadata version and leaves the previous ones alone. At any moment, several complete metadata versions exist side by side in storage. Something outside them has to record which one is current. A directory listing cannot do that job. Listings include objects from failed writes and from writes still in flight; nothing in a filename proves a commit finished; and listing a large prefix on object storage is slow and gives no ordering guarantee you can rely on for correctness. ## What the catalog actually stores A catalog for lakehouse tables holds, for each table identity, essentially one small value: **the location of that table's current metadata**. Around that it provides: - **Name resolution.** `SELECT * FROM analytics.orders` is a logical name. The catalog turns it into a physical metadata location, so queries do not hardcode bucket paths and tables can move. - **Namespaces.** A hierarchy (database/schema, sometimes three levels such as catalog.schema.table) that groups tables and gives grants somewhere to attach. - **Atomicity.** The pointer update is conditional — swap only if the current value is still the one I read — which is what makes a commit all-or-nothing for concurrent writers. - **A governance surface.** Because every engine goes through it to find the table, it is the natural place for grants, audit logs and discovery. Notice what it does **not** store: the file list, the column statistics, the snapshot history. Those live in the table's own metadata, which the catalog merely points at. Keeping the catalog tiny is deliberate — the hot, large metadata scales in object storage, and the catalog only has to be transactional about one pointer per table. ## The read path 1. The engine parses the query and asks the catalog to resolve `analytics.orders`. 2. The catalog returns the current metadata location (and, in some catalogs, configuration or temporary storage credentials). 3. The engine reads that metadata, prunes to the files it needs using the schema, partitioning and file-level statistics recorded there, and scans only those files. Every reader that resolves the table at the same moment plans from the same metadata version, which is what gives readers a stable, consistent view even while a writer is producing the next version. ## Which catalogs people actually use - **Hive Metastore** — the incumbent: a Thrift service backed by a relational database. Widely deployed, JVM-centric, originally designed for Hive-style tables. - **AWS Glue Data Catalog** — a managed, Hive-shaped catalog on AWS. - **The Iceberg REST catalog** — not a product but an HTTP protocol specification that many implementations expose. - **Apache Polaris (incubating)** and **Databricks Unity Catalog** — governance-oriented catalogs that speak open catalog protocols. - **JDBC-backed and Nessie catalogs** — simpler or version-control-flavoured options. ## Can you skip the catalog? Sometimes, with caveats. A filesystem-style catalog keeps the pointer in a file next to the data; that is only safe on a filesystem offering an atomic rename or a conditional write, which is why it is discouraged on S3-like stores. Some formats put commit ordering inside their own log and rely on an atomic put-if-absent from the storage layer (or an external coordinator) rather than a catalog swap — there, a metastore entry exists for discovery and governance, not for correctness. Either way, the requirement is identical: exactly one committed version, agreed on by everybody. ## Why interviewers ask Because "which catalog?" decides whether Spark, Trino, Flink and your warehouse can see one another's writes. Candidates who think a table is "the folder in S3" produce lakes where two engines disagree about the data, and where a half-finished write is indistinguishable from committed data.

  • If the catalog only stores a pointer, where do the table's file list and statistics live?
    In the table's own metadata in object storage, which the pointer addresses. Keeping the file list out of the catalog lets metadata scale with the table while the catalog stays a tiny transactional record — one row per table — that any number of engines can hit cheaply.
  • What happens to a reader that is mid-scan when a writer commits a new version?
    Nothing. The reader resolved the pointer once at plan time and keeps reading that metadata version's files, which are immutable and still present until a retention job removes them. The writer's new version becomes visible only to queries that resolve the table after the swap.
  • Why is listing the storage prefix a poor substitute for the catalog even on a consistent object store?
    A listing shows files, not commits. It cannot distinguish a completed write from a crashed or in-flight one, cannot express deletes recorded in metadata, gives no schema or partitioning, and gets slower as the table grows. Correctness needs a recorded current version, not an inventory.

It is the difference between a phone book and a street: the street exists whether or not anyone lists it, but only the listing tells you which door is currently the right one to knock on.

saying these in an interview costs you the question

  • Says the table is just the folder of Parquet files
  • Thinks the catalog stores every data file path
  • Believes a directory listing is enough to plan a scan
  • Confuses the catalog with a query result cache
  • Assumes any metastore entry alone makes commits atomic

context

open as a page

Why do thousands of tiny data files slow down queries on a lakehouse table?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Each file costs fixed overhead: a metadata entry to plan, an object-store request to open, and a footer to parse. With thousands of tiny files that per-file overhead dominates, so the engine spends its time on bookkeeping rather than reading rows.

open as a page

What is the difference between a file format like Parquet and a table format like Iceberg?

level: juniorimportance: must knowfreq 85%

basics

~20 s

A file format defines how bytes are laid out inside one immutable file. A table format is a metadata layer above a set of such files that records which files, which schema and which version make up the table right now.

open as a page

What does Hive-style directory partitioning do to a table's file layout on object storage?

level: juniorimportance: must knowfreq 72%

basics

~10 s

Hive-style partitioning splits a table's files into directories named column=value, such as dt=2024-05-01. A query that filters on that column reads only the matching directories, so it touches a fraction of the table.

open as a page

How does an open lakehouse table format let you query a table as it looked yesterday?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Every write produces a new immutable snapshot: a record of the complete set of data files that form the table at that instant. Older snapshots are kept, so a time-travel query reads an older file list instead of the current one.

open as a page

How does a catalog make a commit to a lakehouse table atomic when two writers race?

level: middleimportance: must knowfreq 72%

basics

~20 s

Each writer stages new files, then asks the catalog to move the table's metadata pointer only if it still holds the value that writer read. This conditional swap succeeds for exactly one writer; the loser re-reads, revalidates its change and retries.

open as a page

What does bin-packing compaction do to a table's data files, and how is the target size chosen?

level: middleimportance: must knowfreq 72%

basics

~20 s

Bin-packing compaction groups many small files into bins that add up to a target size, rewrites each bin as one new file, and atomically swaps the table's file list. It preserves row content and order-agnostic semantics; the target is a band, commonly a few hundred megabytes.

open as a page

Why is listing a directory of Parquet files an unreliable way to define a table's contents?

level: middleimportance: must knowfreq 70%

basics

~20 s

A directory listing shows whatever is on storage at that instant, including half-written job output and files a rewrite is about to delete. A table format instead reads an explicit versioned list of files, so every reader sees one consistent set.

open as a page

Why does a table partitioned by dt scan every partition when a query filters only on event_ts?

level: middleimportance: must knowfreq 66%

basics

~20 s

Pruning only works on the partition column itself. A predicate on event_ts says nothing about dt, so the engine lists every partition and discards rows afterwards. Filter dt too, or use a format that derives dt from event_ts.

open as a page

How does an open table format make a multi-file write atomic on object storage?

level: middleimportance: must knowfreq 70%

basics

~20 s

It writes all new data files first, where nothing references them, then writes new table metadata describing the resulting snapshot, and finally swaps a single pointer to that metadata in one atomic operation. Readers see the old state or the new one, never a partial write.

open as a page

How can a table format change a table's partitioning without rewriting existing data?

level: seniorimportance: must knowfreq 58%

basics

~20 s

The table records which partition layout each data file was written under. Changing the layout is a metadata commit that applies to new writes only; existing files keep their old partition values, so nothing is rewritten and no downtime is needed.

open as a page

What does the Iceberg REST catalog spec change compared with using a Hive Metastore?

level: middleimportance: should knowfreq 48%

basics

~20 s

The REST catalog defines an HTTP protocol where the server owns commit logic, so any client — Python, Rust, Go — can be thin. A Hive Metastore is a Thrift service with a Hive-shaped model that pushes commit and storage logic into each engine's JVM client library.

open as a page

What is write amplification in a lakehouse table, and what raises it?

level: middleimportance: should knowfreq 50%

basics

~20 s

Write amplification is the ratio of bytes physically written to bytes logically changed. Immutable files force whole-file rewrites, so changing one row in a 512 MB file writes 512 MB. Larger files, frequent small updates, and repeated compaction of the same data all raise it.

open as a page

Why does a table format's metadata layer keep its own per-file min/max column statistics?

level: middleimportance: should knowfreq 55%

basics

~20 s

So the planner can eliminate whole data files from a query by reading metadata alone. Without those bounds in the table layer, the engine would have to open every file's footer just to learn it contains nothing relevant.

open as a page

Why partition a lakehouse table by a hash bucket of user_id rather than by user_id itself?

level: middleimportance: should knowfreq 45%

basics

~10 s

Hash bucketing maps a high-cardinality key onto a fixed number of partitions. Equality lookups on user_id still skip almost everything, while the table keeps a bounded number of partitions holding properly sized files.

open as a page

How does an open lakehouse table format handle two writers committing at the same time?

level: middleimportance: should knowfreq 62%

basics

~20 s

Both writers work optimistically against the snapshot they read, then race to swap the current pointer. The swap is conditional, so exactly one wins; the loser re-reads the new snapshot, checks whether the winner's changes conflict with its own, and retries or fails.

open as a page

A lakehouse table's files are registered in both a Hive Metastore and AWS Glue — what breaks?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Each catalog keeps its own pointer to a current metadata version, so their conditional swaps never see each other. Commits made through one become invisible or get overwritten from the other's view, and a cleanup job run through one deletes files the other still references.

open as a page

After compaction commits, why do the old data files still occupy storage, and how are they removed?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Compaction only removes files from the table's current file list; the objects stay because older table states still reference them and in-flight readers are still reading them. A separate retention-aware cleanup deletes files that no retained state references, after a safety window.

open as a page

When is sorting or clustering during a rewrite worth its cost compared with plain bin-packing?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Bin-packing fixes file count; sorting fixes pruning. Pay for a sort-based rewrite only when queries filter on a column whose values are scattered across every file, so per-file statistics overlap and nothing can be skipped. It costs a full shuffle.

open as a page

An ETL job copied new Parquet files into a lakehouse table's data directory — why don't queries return them?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Readers resolve a table's contents from its metadata layer, never from a directory listing. Files copied straight into storage were never registered by a commit, so they are orphans: invisible to queries and eligible for deletion by the table's cleanup job.

open as a page

A lakehouse table partitioned by hour and customer_id has millions of partitions — what breaks?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Cost moves from scanning to planning: the engine must enumerate and filter millions of partition entries before reading anything, files fall far below target size, and skew grows. Coarsen the time grain and hash the customer key into buckets.

open as a page

Why does expiring old snapshots in a lakehouse table break time travel to last month?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Snapshot expiry drops old versions from the table's history and deletes the data files no surviving snapshot still references. Once that runs, any AS OF query or rollback to a point before the retention horizon has neither a snapshot to resolve nor files to read.

open as a page

How would you choose one catalog for a platform where Spark, Trino, Flink and a warehouse share tables?

level: principalimportance: should knowfreq 38%

basics

~20 s

Start from which engines must write, since writers need a maintained client for the same catalog and only one catalog may own a table. Then weigh governance, operability and migration cost; readers can be served by federation, but writers cannot be split.

open as a page

How would you migrate a large Hive-style Parquet lake to a table format without rewriting the data?

level: principalimportance: should knowfreq 35%

basics

~20 s

Write a table layer over the existing files instead of copying them: the conversion reads each file's footer for schema and statistics and commits a file list referencing the current paths. You inherit the old file sizes and layout, so plan compaction as a separate follow-up.

open as a page

How would you set snapshot retention for lakehouse tables across audit, rollback and cost?

level: principalimportance: should knowfreq 38%

basics

~20 s

Derive retention from how long a bad write takes to detect, then check it against storage cost, the longest running query, and any deletion deadline. Tier it per table rather than setting one platform-wide number, and pin the few versions that must stay reproducible.

open as a page

Why do lakehouse catalogs vend short-lived storage credentials to query engines?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

So that access to a table's files follows the grants held in the catalog. The catalog authorizes the operation and returns temporary credentials scoped to that table's storage path, instead of every engine holding broad bucket credentials that bypass table-level permissions.

open as a page

Why is a lakehouse table format's time travel not a substitute for backups?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Snapshots reference files in the same bucket, catalog and account as the live table, so anything that destroys or loses those objects destroys the history with them. Time travel protects against bad writes inside the retention window, not against losing the storage.

open as a page

After changing a table's partition spec, when do you rewrite historical data under the new layout?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Rewrite history only when old data is queried often enough that the mixed layout measurably hurts, and the rewrite fits your compute budget. Otherwise let the new layout govern new writes and let retention retire the old data.

open as a page