skip to content

What does the Hive Metastore provide that query engines outside Hive still need?

level: middleimportance: must knowfreq 55%

answer

  1. The files alone are not a table
  2. Something must name the columns
  3. It lives in a real relational database
  4. Thrift on 9083, hive.metastore.uris
  5. Trino and Flink read the same catalog

basics

~20 s

The Hive Metastore is a relational catalog mapping table and partition names to storage paths, columns, types, file formats and statistics. Engines such as Trino, Presto and Flink read it so that directories of files behave as SQL tables.

solid answer

~50 s

HDFS stores bytes; it has no idea that `/warehouse/sales/dt=2026-01-01/` is a partition of a table with eight typed columns in ORC. The Hive Metastore (HMS) holds exactly that: databases, tables, columns and types, partition keys and each partition's location, the SerDe and input/output format, table properties and statistics. It keeps this in a real RDBMS (MySQL or PostgreSQL in production) and exposes a Thrift API, by default on port 9083, configured through `hive.metastore.uris`. That API is why it outlived Hive as a query engine. Trino, Presto, Flink and Spark all speak it, so one catalog definition makes the same files queryable from several engines with consistent schemas — the practical origin of separated storage and compute. Modern replacements (AWS Glue Data Catalog, an Iceberg REST catalog, Unity Catalog) mostly reimplement the same role, and Glue deliberately keeps the HMS API shape.

go deeper

for a junior

Know that a table in this ecosystem is files plus a catalog entry, and that the catalog holds the schema and location while the filesystem holds the bytes. Be able to say what DESCRIBE FORMATTED shows you.

for a middle

Explain the mechanics: an RDBMS behind a Thrift service, one row per partition with its own location, and how hive.metastore.uris points clients at it. Expect to be asked why the embedded Derby default fails.

for a senior

Show you have operated it: metastore database backups, partition-count blowups that make planning slow, MSCK REPAIR timing out, and keeping several engines consistent against one catalog.

for a principal

Own the catalog strategy — whether the platform standardizes on HMS, Glue or an Iceberg REST catalog, what the migration path is, and how authorization and lineage attach to whichever catalog you choose.

## The problem it solves A distributed filesystem stores files. It knows a path, a size, a modification time and a set of blocks. It does not know that the CSV under `/warehouse/sales/` has a header, that the third column is a `DECIMAL(12,2)`, that the directories named `dt=2026-01-01` are a partition column rather than a literal folder, or that the whole thing should be read as ORC with a particular SerDe. Something has to hold that interpretation. In the Hadoop ecosystem that something is the **Hive Metastore (HMS)**. ## What it actually stores The metastore is a catalog of *metadata only* — never table rows. Per object it records: - **Databases** (namespaces) and their default warehouse location. - **Tables**: name, owner, table type (`MANAGED_TABLE` or `EXTERNAL_TABLE`), the list of columns with types and comments, and free-form `TBLPROPERTIES`. - **Storage descriptors**: the base location (an HDFS or S3 URI), input format, output format, SerDe class and its properties, compression, bucketing columns and sort columns. - **Partitions**: one row per partition value combination, each with *its own* location and storage descriptor. This is why a partition can live outside the table's base directory. - **Statistics**: row counts, file sizes, and per-column stats such as NDV, min/max and null counts, used by cost-based optimizers. `DESCRIBE FORMATTED <table>` prints most of this back at you, and is the fastest way to see what the catalog actually believes. ## Architecture HMS is a Java service backed by a relational database, accessed through DataNucleus/JDO. Three deployment shapes exist: - **Embedded** — an in-process Apache Derby database. This is the default when nothing is configured and it is the classic first-day trap: Derby in embedded mode allows a single JVM to hold the database, so the second client fails to connect, and the `metastore_db` directory is created wherever you happened to launch the shell. Never use it beyond a tutorial. - **Local** — the metastore code runs inside the client JVM but talks to a shared MySQL/PostgreSQL. - **Remote** — a standalone metastore process exposing Thrift on port 9083, which every client reaches via `hive.metastore.uris=thrift://host:9083`. This is the production shape, and the only one that gives you a single authoritative catalog plus a place to enforce authorization. Since Hive 3 the metastore ships as a **standalone artifact** that can be run without the rest of Hive at all — an acknowledgement that most of its clients are no longer Hive. ## Why non-Hive engines depend on it The Thrift API became the de-facto catalog interface for the whole ecosystem. Trino/Presto, Flink and Spark can all register HMS as a catalog. The payoff is that a table is defined once and read by many engines with the same schema, the same partition layout and the same location — the concrete mechanism behind "separate storage from compute". Without a shared catalog every engine needs its own DDL and the definitions drift. This also explains why HMS is a **hard dependency** long after nobody runs Hive queries: rip it out and a hundred `SELECT`s across three engines stop resolving table names. ## Partitions, the sharp edge Because each partition is a row in a relational database, partition metadata is both the metastore's greatest value and its main failure mode: - New directories written directly to HDFS by another job are **invisible** until the catalog learns about them — hence `MSCK REPAIR TABLE t;` (scan the location and add what is missing) or explicit `ALTER TABLE t ADD PARTITION (...)`. - Over-fine partitioning (hour, then customer, then region) produces hundreds of thousands of partition rows. Planning a query then means a large metastore query, and `MSCK REPAIR` on such a table can run for a very long time or time out. - Because pruning happens against the catalog, wrong or missing statistics silently degrade plans in engines that use cost-based optimization. ## Where it is going In 2026 the direction of travel is toward catalogs designed for table formats rather than directories: the **AWS Glue Data Catalog** (deliberately HMS-API-compatible so existing clients work), **Iceberg REST catalogs**, Unity Catalog and Polaris. These add atomic metadata commits, snapshot/time-travel and file-level tracking that HMS never had. The role, though, is unchanged — and in an interview the useful point is that HMS defined the role, so knowing what it stores transfers directly to whatever replaces it. ## Interview framing If you say only "it stores Hive table metadata", you have answered a 2014 question. The 2026 answer is: it is the shared catalog that turns files into tables for several engines at once, it lives in an RDBMS behind a Thrift API, its partition table is where scale problems appear, and its successors copy its interface.

  • Your team writes new partition directories to HDFS with a non-Hive job. Why do queries not see them?
    Partition pruning resolves against the metastore, not the filesystem, so a directory nobody registered simply does not exist to the planner. Fix it with `ALTER TABLE ... ADD PARTITION` from the writing job, which is cheap and exact, or with `MSCK REPAIR TABLE`, which rescans the whole location and can be very slow on a table with many partitions. Registering partitions as you write them is the pattern that scales.
  • What breaks if the metastore's backing database is lost and you only have the HDFS files?
    The data survives, but every table definition, partition registration, schema, SerDe choice and statistic is gone, so nothing is queryable until the DDL is recreated and partitions re-registered. That is why the metastore RDBMS is a first-class backup target with its own recovery drill, and why storing DDL in version control matters more than it looks.
  • How is the Hive Metastore different from a modern Iceberg catalog?
    HMS tracks tables and partitions as directories, with no atomic multi-file commit and no snapshot history, so concurrent writers can be seen mid-write. An Iceberg catalog tracks explicit file lists per snapshot, giving atomic commits, time travel, schema evolution and file-level pruning without listing directories. Many deployments run Iceberg tables while still registering them in HMS during migration.

It is the library's card catalog. The shelves hold the books; the catalog is what tells you that a particular shelf is the history section, how the books are ordered, and where to find one.

saying these in an interview costs you the question

  • Says the metastore stores the table's actual data rows
  • Thinks only Hive can read the Hive Metastore
  • Confuses metastore metadata with HDFS NameNode block metadata
  • Runs the default embedded Derby metastore in production
  • Believes new HDFS directories become partitions automatically

context