skip to content

What do technical, business and operational metadata each describe in a data catalog, and who typically supplies each kind?

level: juniorimportance: must knowfreq 55%

answer

  1. structure, meaning, behaviour
  2. schemas come from systems
  3. meaning comes from people
  4. runs and usage come from logs

basics

~20 s

Technical metadata describes structure (schemas, types, location); business metadata describes meaning (descriptions, glossary terms, owners, classification); operational metadata describes behaviour (freshness, run history, usage, quality results). Systems supply the first and third, people the second.

solid answer

~40 s

A catalog combines three kinds of metadata. **Technical metadata** is the structure: databases, tables, columns and types, partitions, storage location, row counts — harvested automatically from the systems. **Business metadata** is the meaning: a plain description of the dataset, glossary terms linked to columns, the owner and steward, and sensitivity classification — mostly written by people who know the data, helped by suggestions. **Operational metadata** is the behaviour over time: when the data was last refreshed, run history, who queries it and how often, data-quality results and lineage — collected from pipelines, query logs and checks. A catalog that only has the first kind is a schema browser; the other two are what let a reader decide whether to **trust and use** a dataset.

go deeper

for a junior

Be ready to name the three kinds of metadata with an example of each and say where each comes from.

for a middle

Explain how each kind is kept current and why business metadata needs owners and review.

for a senior

Show how you would prioritise documentation using usage data and connect the catalog to pipelines and query logs.

for a principal

Decide which metadata the organisation makes mandatory, who is accountable for it, and how its completeness is measured.

## What a data catalog is for A **data catalog** is an inventory of an organisation's datasets with the information needed to **find** them, **understand** them and decide whether to **trust** them. Its content is **metadata** — data about data — and it is conventionally split into three kinds. ## The three kinds | Kind | Describes | Examples | Usually supplied by | |---|---|---|---| | **Technical** | structure | schemas, tables, columns, data types, partitions, storage location, row counts, table size | harvested automatically from databases, warehouses and file stores | | **Business** | meaning and accountability | dataset description, column definitions, glossary terms, owner and steward, sensitivity classification, certification status | people who know the data, often seeded by automated suggestions | | **Operational** | behaviour over time | last refresh time, job run history, query counts and top users, data-quality check results, lineage | pipelines, schedulers, query logs and quality tools | ## Why all three matter to a reader Imagine an analyst searching for "active customers": - **Technical** metadata shows three candidate tables and their columns. - **Business** metadata says which one is the certified source, what "active" means there, and whom to ask. - **Operational** metadata shows that one candidate has not refreshed for two months and another is queried daily by finance — strong hints about which to use. Without business and operational metadata the analyst is back to asking colleagues; the catalog has only replaced a database's own schema browser. ## How each kind stays current 1. **Technical** metadata changes with every schema change, so it must be harvested automatically and frequently. 2. **Business** metadata decays when people leave or definitions change; it needs named owners and review. 3. **Operational** metadata is a by-product of running things; it is kept current by connecting the catalog to schedulers, query logs and checks rather than by typing. ## Common mistakes - Treating a filled-in schema as a finished catalog. - Asking engineers to type descriptions for thousands of columns at launch instead of prioritising the most-used datasets. - Recording owners without a process to update them when people change roles. ## Why interviewers ask it It checks whether a candidate sees the catalog as more than a list of tables. The good answer names the three kinds, gives concrete examples, and says **who or what** supplies each — which is also the answer to how the catalog stays accurate.

  • Which kind of metadata goes stale fastest, and why?
    Business metadata. Technical and operational metadata are regenerated by systems as things change and run, but descriptions, owners and glossary links depend on people and silently rot when teams reorganise or definitions shift.
  • Where would you start filling business metadata in a warehouse of thousands of tables?
    With the most-queried and most business-critical datasets, found from operational metadata such as query counts. Describing the top few percent covers most real usage, while the rest can carry automated suggestions until someone needs them.

saying these in an interview costs you the question

  • Describing a catalog as just a list of table schemas
  • Expecting engineers to hand-document every column before launch
  • Treating ownership as a one-time field rather than something maintained
  • Confusing operational metadata with the data itself