What does the Iceberg REST catalog spec change compared with using a Hive Metastore?
answer
- one of these two is not a product
- who decides whether a commit is allowed
- the client carried too much logic before
- assertions plus updates, evaluated server-side
- HTTP and JSON need no JVM client
basics
~20 sThe REST catalog defines an HTTP protocol where the server owns commit logic, so any client — Python, Rust, Go — can be thin. A Hive Metastore is a Thrift service with a Hive-shaped model that pushes commit and storage logic into each engine's JVM client library.
solid answer
~50 sThe Hive Metastore is an implementation: a Thrift service over a relational database, with a table model inherited from Hive, and every engine reaching it through a JVM client that also carries the logic for building and committing table metadata. The REST catalog is a **specification** — an OpenAPI-described HTTP contract for namespaces, table loading and commits — with many implementations behind it. Three practical consequences. First, clients get thin: a commit is sent as *requirements* (assertions about the state the writer read) plus *updates*, and the **server** decides whether the swap succeeds, so upgrading commit behaviour no longer means upgrading every engine. Second, non-JVM engines can be first-class, because HTTP plus JSON needs no Thrift or Hive client. Third, the server can return per-client configuration and can vend scoped storage credentials, making the catalog a real authorization boundary rather than just a lookup.
go deeper
Know that a catalog can be reached over HTTP rather than through a Hive-specific Thrift client, and that this is what lets Python and other non-JVM tools work with these tables.
Draw the distinction between a specification and an implementation, and explain that the commit arrives as assertions plus updates that the server evaluates.
Discuss the operational payoff: one place to evolve commit behaviour, authentication and multi-tenancy on a normal HTTP service, and scoped credentials instead of broad bucket access for every engine.
Weigh migrating an estate already wired to a metastore: dual-running during transition, client version floors across engines, and whether a façade over the existing metastore buys you the protocol without the move.
## Two different kinds of thing A useful first move in the interview is to refuse the false symmetry. **Hive Metastore (HMS)** is a *product*: a long-running Thrift service backed by a relational database, storing databases, tables, partitions, columns and storage descriptors. **The Iceberg REST catalog** is a *protocol*: an OpenAPI specification for how a client asks a catalog to list namespaces, load a table, and commit a change. Polaris, Unity Catalog, Nessie, Gravitino, cloud vendors' catalogs and small in-house services can all speak it; even HMS can sit behind a REST-speaking front end. ## Where the logic lives With HMS, the pointer to a table's current metadata is stored as a table property, and the client library — running inside Spark, Trino or Flink — is what builds the new metadata, decides whether the commit is valid, and performs the conditional update. That means: - Every engine ships its own copy of catalog logic, and behaviour depends on which version each shipped. - A change to commit semantics is a fleet-wide client upgrade. - The catalog cannot enforce much, because it does not understand the operation; it sees a property update. Under the REST spec, the client sends a structured commit: a set of **requirements** — assertions such as "the table's current snapshot is still the one I read" — and a set of **updates** to apply. The server evaluates the requirements and applies the updates atomically or rejects the whole request. The commit decision moves server-side. That is the architectural change everything else follows from. ## Consequences that matter in production **Thin, polyglot clients.** Talking Thrift and the Hive object model in Python, Rust or Go is painful; talking HTTP and JSON is not. This is why the REST spec is what made non-JVM readers and writers practical, and it is the answer interviewers are usually fishing for. **Server-side evolution.** Fix a conflict-detection bug, add support for a newer table format version, change how retries behave — deploy the catalog, not fifty engine images. **Configuration handshake.** A client fetches catalog configuration at initialization, so the server can hand out defaults and overrides (warehouse location, storage settings) instead of every engine holding its own copy in a config file. **Authorization boundary.** Because the server understands the operation and identity, it can authorize at table and namespace level, and can return temporary, scoped storage credentials for the specific table being read rather than requiring every engine to hold broad bucket credentials. **A real multi-tenant surface.** HMS was designed for a Hadoop cluster's trusted interior; putting it on the open internet with proper authentication and tenancy is uncomfortable. An HTTP service with standard auth is a normal thing to operate. ## What does not change - **The table format is untouched.** Data files and metadata still live in your object storage in the same layout. A REST catalog is not a storage system and does not hold your rows. - **The atomicity requirement is identical.** Whether the swap happens as a conditional database update inside HMS or as requirement evaluation in a REST server, exactly one writer may win a given transition. - **A table still needs one owning catalog.** Speaking a standard protocol does not make it safe for two catalogs to accept writes to the same files. - **Migration is not free.** Existing tables registered in HMS or a cloud catalog must be moved deliberately, and every engine in the estate needs a client that speaks the new protocol at a compatible version. ## Why HMS persists anyway Because it is everywhere. A decade of Hive, Spark and Trino deployments resolve tables through it; non-Iceberg tables in the estate may depend on its partition model; and tooling, permissions and lineage systems are wired to it. Many platforms therefore run both: HMS for legacy Hive-style tables, a REST-speaking catalog for open table format tables, or a REST façade in front of the existing metastore so clients converge on one protocol while the storage of record stays put. ## How to answer this in an interview Lead with "specification versus implementation", then name the one structural change — the commit decision moves to the server — and let the consequences (polyglot clients, server-side evolution, credential vending, real multi-tenancy) fall out of it. Candidates who instead describe the REST catalog as "a faster metastore" or "a metastore in the cloud" have missed the point entirely.
- Why does moving commit decisions server-side help a fleet with many engine versions?Because conflict detection, retry semantics and support for newer table format versions live in one deployable service instead of in every engine image. You fix or extend the catalog once; Spark, Trino and Flink keep their existing thin clients and immediately get the new behaviour.
- Does adopting a REST catalog change how the data files are laid out in storage?No. The table format decides the file layout, and the catalog only records which metadata version is current. Migration is a matter of re-registering table identity and pointing engines at a new endpoint; the Parquet files and the metadata objects stay exactly where they are.
- Can a Hive Metastore still be involved once you adopt the REST protocol?Yes, in two shapes. A REST-speaking service can front an existing metastore so the pointer of record still lives there while clients converge on one protocol; or HMS keeps serving legacy Hive-style tables while open-format tables move to the new catalog. Running both is common during migration.
saying these in an interview costs you the question
- Calls the REST catalog a faster or cloud-hosted metastore
- Thinks the REST catalog stores the table's data
- Assumes the protocol removes the need for atomic commits
- Says any engine can use it without a compatible client
- Believes one shared protocol makes dual-catalog writes safe