skip to content

What is data lineage in a data pipeline, and what questions does it let a team answer?

level: juniorimportance: should knowfreq 58%

answer

  1. a map of what fed what
  2. two directions: cause and blast radius
  3. nodes are datasets, edges are jobs
  4. recorded per run, not hand-drawn
  5. impact analysis before a schema change

basics

~20 s

Data lineage is the recorded graph of which datasets each pipeline job read and which it wrote. It answers where a table's numbers came from, what breaks if an upstream column changes, and which reports a failed job affected.

solid answer

~50 s

Lineage is a graph whose nodes are datasets (tables, files, topics) and whose edges are the jobs that read some and produce others. A useful lineage system records it **per run**, not just as static documentation, so you can also say *when* a dataset was last produced and by which run. It answers two directions of question. Walking **upstream** from a suspicious number gives root cause: which job wrote this table, and what did that job read. Walking **downstream** from a table or column gives impact analysis: if this source drops a field, or today's load failed, which models, marts, dashboards and machine-learning features are affected and who owns them. Secondary payoffs are trust (a consumer can see the provenance of a figure), compliance (where personal data travelled), and safe deletion — nothing downstream reads it, so it can go.

go deeper

for a junior

Be able to define lineage as the recorded graph of datasets and the jobs that derive them, and give one concrete use: tracing a wrong number back to its source table.

for a middle

Explain how lineage is captured mechanically — parsed from code or emitted by each run — and why per-run records add information that a static graph cannot provide.

for a senior

Show you use it operationally: walking downstream during an incident to size blast radius and identify owners, and upstream to isolate which run introduced a bad value.

for a principal

Own the argument that lineage coverage is a platform capability with a cost, and be able to say which datasets deserve full coverage and what you accept not knowing.

## What lineage is **Data lineage** is a recorded description of how data moved and was transformed to become what it is now. Formally it is a directed graph: **datasets** are nodes (a warehouse table, a file prefix in object storage, a stream topic, a dashboard) and **jobs** are edges (a transformation, a load, a copy) that consume one or more input datasets and produce one or more output datasets. An edge says nothing about *when* on its own, which is why serious lineage is recorded per **run**: this execution of this job, starting at this time, read these dataset versions and wrote that one. Job-level lineage is the map; run-level lineage is the map plus a history of journeys on it. ## Two directions of question Lineage is valuable because production questions are almost always directional. **Upstream (root cause).** A finance user says the revenue number is wrong. Without lineage you ask around and grep repositories. With lineage you walk back from the mart to the model that built it, to the staging table it read, to the raw landing dataset, and you look at each producing run: which one ran late, which one produced a fifth of its usual rows, which one changed code yesterday. **Downstream (impact analysis).** An upstream team announces a column is being renamed, or a nightly load has failed and will not be rerun until noon. The question is what breaks and who to tell. Lineage turns that from guesswork into a graph traversal: every dataset reachable downstream, plus the owners attached to each, plus the dashboards at the leaves. ## Design-time versus run-time lineage Hand-maintained lineage — a diagram in a wiki, a spreadsheet of table dependencies — decays within weeks because nothing forces it to match reality. Two mechanical alternatives dominate: - **Static/parsed lineage**: read the SQL or code and infer inputs and outputs from it. Cheap, works before anything runs, but misses anything dynamic and does not know whether the job actually ran. - **Runtime-emitted lineage**: the job (or the orchestrator running it) emits an event at start and end saying what it read and wrote. It reflects what genuinely happened, including a run that read nothing because a partition was empty. Most platforms end up with both, and store the result in a metadata service or catalog that exposes search and graph traversal. ## What lineage buys beyond debugging - **Trust.** A consumer looking at a figure can see its provenance and its last refresh, instead of asking an analyst on chat. - **Change management.** Deprecating a table safely means proving nothing reads it. Lineage plus query history is the proof. - **Regulatory work.** "Where does customer email travel?" is a lineage question; so is showing an auditor how a regulatory report was derived. - **Cost and cleanup.** Datasets with no downstream consumers and no query traffic are candidates for deletion — often a large share of a mature warehouse. - **Prioritising incidents.** A failure that reaches the executive dashboard and the billing feed is a different severity from one that reaches a sandbox table, and lineage is what tells them apart. ## What lineage is not It is not a catalog, though catalogs display it. A catalog holds descriptions, owners, tags and schemas; lineage is the connective tissue between those entries. It is also not the same as a workflow's **task dependency graph**. A workflow says task B runs after task A; lineage says the dataset B wrote was derived from the dataset A wrote. They frequently disagree: two tasks may be ordered for convenience without any data flowing between them, and data may flow between two pipelines that have no ordering relationship at all — which is exactly the case where lineage tells you something the workflow definition cannot. Finally, lineage records structure, not semantics. It tells you `orders_daily` was built from `orders_raw`; it does not tell you whether the aggregation logic is correct. Pair it with data-quality signals — row counts, freshness, distribution checks — to get from "where did this come from" to "is today's version sane". ## A minimal practical shape A usable first version records, for every run: job identity, start and end time, terminal state, the input datasets, the output datasets, and the row count written. That alone supports root-cause walks, impact analysis and staleness alerts, and it is small enough that a team can adopt it without a large platform project.

  • How is lineage different from a workflow's task dependency graph?
    A dependency graph encodes ordering — B runs after A — which may exist for convenience with no data flowing between them. Lineage encodes derivation: the dataset B wrote was built from the dataset A wrote. Data often flows between pipelines that have no ordering relationship at all, and those edges are invisible in the workflow definition but visible in lineage.
  • Why does hand-maintained lineage documentation stop being useful so quickly?
    Nothing forces it to match reality. A wiki diagram is correct on the day it is drawn and drifts with every merged change, and because it is never wrong loudly, nobody notices until an impact analysis based on it misses a consumer. Lineage has to be derived mechanically from code or emitted by runs to stay trustworthy.
  • What can lineage tell you during an incident that a failure alert cannot?
    The alert says a job failed. Lineage says what that failure reaches: which downstream datasets are now stale, which dashboards and feeds sit at the leaves, and who owns them. That converts a red task into a severity judgement and a list of people to notify.

It is the supply chain record for a number: which ingredients went into the dish, which kitchen cooked it, and which meals you have to recall if one ingredient turns out to be bad.

saying these in an interview costs you the question

  • Describes lineage as a diagram someone maintains by hand
  • Confuses lineage with a data catalog's schema and descriptions
  • Treats task ordering in a workflow as lineage
  • Only walks upstream; never mentions downstream impact analysis
  • Claims lineage tells you whether the transformation logic is correct

context