skip to content

How would you choose one catalog for a platform where Spark, Trino, Flink and a warehouse share tables?

level: principalimportance: should knowfreq 38%

answer

  1. some engines only need to read
  2. one of these constraints is absolute
  3. count the writers before comparing products
  4. specified protocol beats a product's client library
  5. the warehouse is usually the hardest constraint

basics

~20 s

Start from which engines must write, since writers need a maintained client for the same catalog and only one catalog may own a table. Then weigh governance, operability and migration cost; readers can be served by federation, but writers cannot be split.

solid answer

~50 s

Frame it as a constraint problem, not a product comparison. **Hard constraints first:** every engine that writes needs a maintained client for the chosen catalog at a compatible version, and each table has exactly one owning catalog — you cannot split writers across two. **Then governance:** do you need table-level authorization enforced at access time, credential vending, audit and lineage, or is discovery enough? **Then operability:** who runs it, what is its availability target given it sits on every query's path, how does it back up and fail over. **Then migration:** how many tables are registered somewhere already, and can you dual-run with a read-only mirror during cutover. In practice the warehouse is the hardest constraint, because it often insists on its own catalog; the workable answer is usually one owning catalog for the lake with read-only federation into the warehouse, or a catalog that speaks the warehouse's protocol natively.

go deeper

for a junior

Know the outcome rather than the method: engines that must share tables should share one catalog, and this is decided deliberately rather than per team.

for a middle

Be able to say why writers constrain the choice and readers do not, and name the common catalog options with the protocol each speaks.

for a senior

Argue the operational side: availability on the query path, credential and rate-limit behaviour under large jobs, and a table-by-table cutover that never leaves two writable registrations alive.

for a principal

Own the whole decision — writer set, governance level matched to a real requirement, protocol-level replaceability, migration plan and the warehouse boundary — and state plainly what you traded away.

## Why this is a design question, not a shortlist There is no best catalog. There is a set of constraints, and the honest answer names them in priority order and shows what you would give up. An interviewer at this level is checking whether you know which constraints are hard. ## Constraint 1: who writes This dominates everything. Reading a table through a second, read-only path is tractable; writing through two catalogs is not, because atomicity lives in a single pointer and a single conditional swap. So enumerate the writers — batch jobs, streaming jobs, the warehouse, dbt-style transformation tools, ingestion services — and require that every one of them has a maintained client for the candidate catalog at a version you can deploy. A candidate catalog that a single writer cannot speak is either eliminated or forces that writer into a read-only role. ## Constraint 2: governance requirements Ask what "secure" has to mean here: - **Discovery only** — the catalog resolves names and holds the pointer; permissions live in storage policy. Cheapest, weakest. - **Table-level authorization enforced at access time** — the catalog authorizes and vends scoped, short-lived storage credentials, and the direct bucket path is closed. This is what regulated environments usually need, and it eliminates plain metastores. - **Row and column controls, lineage, tagging, retention policy** — a governance catalog, at the cost of a heavier dependency and often a vendor. Do not buy governance you have no requirement for; do not pretend a metastore provides it. ## Constraint 3: operability and blast radius Once chosen, the catalog is on the resolution path of every query and, if it vends credentials, on the data path too. Treat it as a tier-one service: availability target at least as strong as the engines that depend on it, tested restore of its backing store, a failover story, and rate limits that survive a thousand-task Spark job resolving tables at once. Managed offerings trade control for that burden; self-hosting an open implementation trades cost for the pager. ## Constraint 4: protocol and lock-in Prefer a catalog whose protocol is specified rather than one whose interface is a specific product's client library. A specified HTTP protocol means today's clients keep working if you replace the implementation, and new engines are cheap to onboard. It also puts commit logic server-side, so evolving behaviour is a catalog deploy rather than a fleet upgrade. Ask explicitly: if we want to move off this in three years, what has to change — table metadata, or just an endpoint and some credentials? ## Constraint 5: migration cost Count the tables already registered somewhere, and the pipelines whose configuration names that catalog. The cutover per table is: stop writers, register the current metadata in the new catalog, repoint every writer and reader, delete the old writable entry. Anything that leaves two writable entries alive is worse than not migrating, because it silently loses commits. Plan for a dual-run period where the old catalog is read-only, and for a detection sweep that flags any two entries with overlapping storage locations. ## The warehouse problem Warehouses frequently insist on resolving tables through their own catalog. Options, roughly in order of preference: 1. The warehouse can act as a client of the lake's catalog, or the lake's catalog speaks a protocol the warehouse understands natively. One owner, everybody sees the same commits. 2. The warehouse mirrors the lake catalog read-only — federation or one-way sync — and never writes those tables. Acceptable; accept some staleness and be explicit about it. 3. The warehouse owns some tables and the lake catalog owns others, with a clear ownership boundary per table and no overlap. Workable but requires discipline. 4. Both write the same tables. Not an option; this is the doubly-registered failure. ## Decision shape to state out loud "One owning catalog per table, chosen by the union of writers that must speak it; a specified protocol so engines stay replaceable; governance level matched to an actual requirement rather than aspiration; run as a tier-one service; migration table by table with the old entry made read-only, then deleted." Then name what you gave up — usually some engine's native features, or a period of staleness on the mirror — because a principal-level answer that claims no tradeoff is not credible. ## Anti-patterns to call out Choosing by benchmark; choosing the catalog your favourite engine defaults to; letting each team pick their own; deferring the decision until "we see what sticks", which in practice means several catalogs claiming the same tables; and treating a crawler that auto-registers bucket contents as a catalog strategy.

  • What makes the writer set the first thing you enumerate?
    Because only one catalog can own a table's pointer, so every writer must speak the same catalog. Readers can be served by a read-only mirror or federation and can tolerate staleness; writers cannot be split without losing commits silently. The writer list therefore eliminates candidates before any feature comparison starts.
  • How do you cut over an estate of a few thousand tables without an outage?
    Table by table, in dependency order. For each: quiesce writers, register the current metadata in the new catalog, repoint writers and readers, then make the old entry read-only and later delete it. Run a sweep that flags overlapping storage locations across catalogs so no table ends up with two writable owners.
  • When is running two catalogs the right answer rather than a failure?
    When ownership is partitioned cleanly by table, not shared. A warehouse owning its native tables while the lake catalog owns open-format tables is fine, as is a read-only mirror for discovery. What is never fine is two catalogs accepting writes to the same table, however well-intentioned the sync between them.
  • What availability target does the catalog need?
    At least that of the engines depending on it, because it sits on every query's resolution path and, with credential vending, on the data path too. Budget for restore testing of its backing store, a failover plan, and rate limits that survive large jobs resolving tables from thousands of tasks simultaneously.

saying these in an interview costs you the question

  • Picks by feature matrix without listing which engines write
  • Plans to let each team choose its own catalog
  • Treats two writable catalogs as acceptable if they sync
  • Ignores the catalog's availability on the query hot path
  • Assumes migration is just repointing a configuration property

context