skip to content

In the OpenLineage standard, what does a run event carry and what are facets used for?

level: middleimportance: nice to knowfreq 28%

answer

  1. three nouns hold the model
  2. one execution has several events
  3. the core stays small on purpose
  4. optional typed blocks bolt on metadata
  5. row counts ride along with the edges

basics

~20 s

An OpenLineage run event ties three entities together — a job, one run of it, and the input and output datasets — plus an event type such as START, COMPLETE or FAIL and a timestamp. Facets are optional typed blocks that attach extra metadata to any of them.

solid answer

~50 s

OpenLineage models lineage as events posted during execution rather than as a graph someone writes. Each event names the **job** (the recurring unit of work, in a namespace), the **run** (one execution, identified by a UUID), and the **inputs** and **outputs** — datasets, also namespaced. It carries an `eventType` (`START`, `RUNNING`, `COMPLETE`, `FAIL`, `ABORT`, `OTHER`) and an `eventTime`, so a consumer stitches multiple events for the same `runId` into a full picture. **Facets** keep the core small and the standard extensible. They are optional, typed, versioned JSON blocks attached to a run, a job or a dataset — for example a `schema` facet on a dataset, `columnLineage` for field-level edges, `outputStatistics` for rows and bytes written, `dataQualityMetrics` on an input, `sql` on a job, `parent` to link a task run to its enclosing pipeline run. A metadata server such as Marquez receives the events and materialises the graph.

code

json · 16 lines
json
{
  "eventType": "COMPLETE",
  "eventTime": "2026-08-21T02:14:53.101Z",
  "producer": "https://example.com/integrations/orchestrator",
  "run": {
    "runId": "3f1b9d2e-6c44-4d1a-9c31-0a7e5b2f8a10",
    "facets": { "nominalTime": { "nominalStartTime": "2026-08-20T00:00:00Z" } }
  },
  "job": { "namespace": "warehouse-prod", "name": "build_orders_daily" },
  "inputs": [ { "namespace": "warehouse-prod", "name": "raw.orders" } ],
  "outputs": [ {
    "namespace": "warehouse-prod",
    "name": "mart.orders_daily",
    "facets": { "outputStatistics": { "rowCount": 9133, "size": 2201984 } }
  } ]
}

go deeper

for a junior

Recognise the three entities — job, run, dataset — and that lineage here is emitted as events during execution rather than drawn by hand.

for a middle

Explain why multiple events share one run identifier, and what facets are for: optional versioned metadata blocks that keep the core small and extensible.

for a senior

Discuss what the model enables operationally — run-scoped history, statistics travelling with lineage, parent links from task run to pipeline run — and where naming collisions break the graph.

for a principal

Own the interchange decision: adopting a shared event format versus per-tool integrations, and the naming conventions every producer must honour for the graph to connect.

## Why a standard exists at all Before an interchange format, every tool grew its own lineage representation, and stitching a pipeline that spans an ingestion tool, an orchestrator, a processing engine and a warehouse meant writing N×M adapters. **OpenLineage** defines one event format that producers emit and consumers ingest, so integrations are written once per tool rather than once per pair. It is also **push, at runtime**: the thing doing the work reports what it did. That means the recorded lineage reflects the execution that actually happened, including runs that failed and runs that read nothing. ## The core model: job, run, dataset Three entities carry almost everything. - **Job** — the recurring unit of work: a scheduled transformation, a task in a pipeline, a query. Identified by a **namespace** plus a **name**, so `warehouse-prod` and `warehouse-dev` do not collide. - **Run** — one execution of a job, identified by a **UUID** (`runId`). Runs are what carry timing, state and per-execution metrics. - **Dataset** — an input or output: a table, a file path, a topic. Also namespace plus name, where the namespace conventionally identifies the storage system or connection. An event ties them together: *this run of this job read these datasets and wrote those.* ## The run event A `RunEvent` carries an `eventTime`, an `eventType`, the `run` and `job` objects, `inputs` and `outputs` arrays, and a `producer` identifying the integration that emitted it. The event types are `START`, `RUNNING`, `COMPLETE`, `FAIL`, `ABORT` and `OTHER`. The important consequence is that **lineage is assembled from multiple events**: a `START` at the beginning, optionally `RUNNING` updates, and a terminal `COMPLETE` or `FAIL`. Both carry the same `runId`, and a consumer merges them. This is why you can know a run began, know what it intended to read, and still record that it failed — a state a static graph cannot express. ```json { "eventType": "COMPLETE", "eventTime": "2026-08-21T02:14:53.101Z", "run": { "runId": "3f1b9d2e-6c44-4d1a-9c31-0a7e5b2f8a10" }, "job": { "namespace": "warehouse-prod", "name": "build_orders_daily" }, "inputs": [ { "namespace": "warehouse-prod", "name": "raw.orders" } ], "outputs": [ { "namespace": "warehouse-prod", "name": "mart.orders_daily" } ] } ``` ## Facets: the extension mechanism If every useful piece of metadata were a core field, the specification would be enormous and every addition would be a breaking change. **Facets** solve that: optional, self-describing, independently versioned JSON objects attached to a run, a job, a dataset, or specifically to an input or an output dataset. Each declares the schema it conforms to, so a consumer that does not recognise a facet can ignore it without failing. Commonly used ones include: - **Dataset facets** — `schema` (field names and types), `dataSource` (where it lives), `columnLineage` (field-level edges, including the input fields behind each output field), `documentation`, `ownership`, `version`. - **Output dataset facets** — `outputStatistics`, carrying row count and size written by this run. This is the metric that turns lineage into observability: not just *that* a table was written, but how much. - **Input dataset facets** — `dataQualityMetrics`, carrying row counts and per-column statistics of what was read. - **Run facets** — `nominalTime` (the logical window the run covers, distinct from wall-clock start), `parent` (linking a task run to the run of its enclosing pipeline), `errorMessage` on failures. - **Job facets** — `sql` (the query text), `sourceCodeLocation`, `documentation`, `ownership`. Organisations can also define custom facets for internal metadata without forking the specification. ## Producers and consumers A **producer** is anything that emits events: an integration inside an orchestrator that reports each task run, a plugin in a processing engine, a wrapper around a transformation tool, or your own code posting to the HTTP endpoint. A **consumer** stores and serves them — **Marquez** is the reference implementation, exposing the accumulated graph and run history through an API and a UI, and several commercial catalogs ingest the same format. ## Why the model matters in an interview The interesting points are not the field names but the design choices they encode. Events rather than a static graph means lineage reflects reality including failures. Run-scoped identity means you can ask when a dataset was last produced and by which execution. Facets mean lineage, schema, statistics and quality metrics travel in **one** message, so the metadata store can answer "what fed this, how many rows, and was it fresh" without joining three systems. And a shared format means a pipeline crossing four tools can still produce one connected graph — provided all four agree on how datasets are named, which in practice is the hardest part of any real deployment.

  • Why does one run produce more than one event?
    Events are emitted as execution progresses — a START when the run begins, optional RUNNING updates, and a terminal COMPLETE, FAIL or ABORT — all sharing the same runId. The consumer merges them, so it can record that a run started and what it intended to read even if it never finished, which a single end-of-run message would lose.
  • What does the facet mechanism buy over adding fields to the core specification?
    Extensibility without breakage. Facets are optional, self-describing and versioned independently, so new metadata can appear without changing the core schema, consumers ignore facets they do not recognise, and organisations can attach custom facets for internal needs without forking the standard.
  • Which facet turns a lineage record into an observability signal?
    Statistics carried alongside the outputs — row count and size written by that run — plus quality metrics on the inputs. With those, the metadata store answers not only what fed a dataset but how much data each run produced, which is the history a volume or freshness alert compares against.
  • What is the hardest part of getting a connected graph from several tools emitting this format?
    Dataset naming. Two producers must refer to the same physical table with the same namespace and name, or the graph silently splits into disconnected halves that look complete to each tool. Agreeing naming conventions across ingestion, orchestration, processing and the warehouse is where most deployments spend their effort.

saying these in an interview costs you the question

  • Thinks a run is the same thing as a job
  • Assumes one event per run carries everything
  • Believes facets are mandatory fields of the specification
  • Confuses the event format with the metadata server that stores it
  • Expects a connected graph without agreeing dataset naming across producers

context