Why does a query engine need a catalog to read a lakehouse table instead of a path?
answer
- a folder of files cannot say which files count
- several metadata versions exist at once
- something must name the current one
- one small pointer per table name
- name resolution plus an atomic pointer
basics
~20 sA catalog maps a table name to the location of that table's current metadata, so every engine agrees on which version is live. A storage path only lists files, including leftovers from failed or in-flight writes that no commit includes.
solid answer
~40 sAn open table format keeps data as ordinary files in object storage plus a metadata layer that records which files, schema and partitioning make up the table **right now**. Because each commit writes new metadata rather than editing the old in place, several metadata versions sit in the directory at once, and nothing in a file listing says which one is current — a listing also shows files from crashed jobs and writes still in flight. The catalog is the transactional service that stores, per `namespace.table`, one value: the pointer to the current metadata. An engine resolves `analytics.orders` through the catalog, follows the pointer, and plans its scan from that metadata. That single indirection is also what makes multi-engine reads consistent and what gives you a place to hang grants, audit and discovery.
go deeper
Be able to say in one breath that the catalog turns a table name into the current metadata location, and that a bare storage path cannot tell committed files from leftovers.
Explain why several metadata versions coexist in storage at once and why the pointer therefore has to live somewhere transactional outside those files.
Show that you treat the catalog as the correctness boundary for multi-engine access, and can name what breaks operationally when engines resolve a table by path instead.
Frame the catalog as the platform's interoperability and governance contract: one source of truth for table identity, grants and audit, chosen before engines are.
## A directory of files is not a table Open table formats store table data as ordinary files — usually Parquet — in object storage, and add a **metadata layer** above them. That metadata layer records the schema, the partitioning scheme, and above all the exact set of data files that belong to the table at this instant. It is the thing that distinguishes "these 812 files" from "those 40 leftovers a crashed job wrote at 03:12". But the metadata layer inherits the same problem one level up. Object storage has no safe in-place edit, so a commit does not modify the existing metadata: it writes a **new** metadata version and leaves the previous ones alone. At any moment, several complete metadata versions exist side by side in storage. Something outside them has to record which one is current. A directory listing cannot do that job. Listings include objects from failed writes and from writes still in flight; nothing in a filename proves a commit finished; and listing a large prefix on object storage is slow and gives no ordering guarantee you can rely on for correctness. ## What the catalog actually stores A catalog for lakehouse tables holds, for each table identity, essentially one small value: **the location of that table's current metadata**. Around that it provides: - **Name resolution.** `SELECT * FROM analytics.orders` is a logical name. The catalog turns it into a physical metadata location, so queries do not hardcode bucket paths and tables can move. - **Namespaces.** A hierarchy (database/schema, sometimes three levels such as catalog.schema.table) that groups tables and gives grants somewhere to attach. - **Atomicity.** The pointer update is conditional — swap only if the current value is still the one I read — which is what makes a commit all-or-nothing for concurrent writers. - **A governance surface.** Because every engine goes through it to find the table, it is the natural place for grants, audit logs and discovery. Notice what it does **not** store: the file list, the column statistics, the snapshot history. Those live in the table's own metadata, which the catalog merely points at. Keeping the catalog tiny is deliberate — the hot, large metadata scales in object storage, and the catalog only has to be transactional about one pointer per table. ## The read path 1. The engine parses the query and asks the catalog to resolve `analytics.orders`. 2. The catalog returns the current metadata location (and, in some catalogs, configuration or temporary storage credentials). 3. The engine reads that metadata, prunes to the files it needs using the schema, partitioning and file-level statistics recorded there, and scans only those files. Every reader that resolves the table at the same moment plans from the same metadata version, which is what gives readers a stable, consistent view even while a writer is producing the next version. ## Which catalogs people actually use - **Hive Metastore** — the incumbent: a Thrift service backed by a relational database. Widely deployed, JVM-centric, originally designed for Hive-style tables. - **AWS Glue Data Catalog** — a managed, Hive-shaped catalog on AWS. - **The Iceberg REST catalog** — not a product but an HTTP protocol specification that many implementations expose. - **Apache Polaris (incubating)** and **Databricks Unity Catalog** — governance-oriented catalogs that speak open catalog protocols. - **JDBC-backed and Nessie catalogs** — simpler or version-control-flavoured options. ## Can you skip the catalog? Sometimes, with caveats. A filesystem-style catalog keeps the pointer in a file next to the data; that is only safe on a filesystem offering an atomic rename or a conditional write, which is why it is discouraged on S3-like stores. Some formats put commit ordering inside their own log and rely on an atomic put-if-absent from the storage layer (or an external coordinator) rather than a catalog swap — there, a metastore entry exists for discovery and governance, not for correctness. Either way, the requirement is identical: exactly one committed version, agreed on by everybody. ## Why interviewers ask Because "which catalog?" decides whether Spark, Trino, Flink and your warehouse can see one another's writes. Candidates who think a table is "the folder in S3" produce lakes where two engines disagree about the data, and where a half-finished write is indistinguishable from committed data.
- If the catalog only stores a pointer, where do the table's file list and statistics live?In the table's own metadata in object storage, which the pointer addresses. Keeping the file list out of the catalog lets metadata scale with the table while the catalog stays a tiny transactional record — one row per table — that any number of engines can hit cheaply.
- What happens to a reader that is mid-scan when a writer commits a new version?Nothing. The reader resolved the pointer once at plan time and keeps reading that metadata version's files, which are immutable and still present until a retention job removes them. The writer's new version becomes visible only to queries that resolve the table after the swap.
- Why is listing the storage prefix a poor substitute for the catalog even on a consistent object store?A listing shows files, not commits. It cannot distinguish a completed write from a crashed or in-flight one, cannot express deletes recorded in metadata, gives no schema or partitioning, and gets slower as the table grows. Correctness needs a recorded current version, not an inventory.
It is the difference between a phone book and a street: the street exists whether or not anyone lists it, but only the listing tells you which door is currently the right one to knock on.
saying these in an interview costs you the question
- Says the table is just the folder of Parquet files
- Thinks the catalog stores every data file path
- Believes a directory listing is enough to plan a scan
- Confuses the catalog with a query result cache
- Assumes any metastore entry alone makes commits atomic