An auditor asks which exact source data produced a regulatory report published last quarter; why is a static lineage graph not enough to answer?
answer
- graph today versus graph then
- which run, which input version
- code changes and backfills since
- snapshots or retained versions of inputs
- evidence kept with the report
basics
~20 sA static graph shows today's dependencies, not what ran then. You need run-level records naming which job version ran, which input dataset versions it read, and when, plus retained copies or snapshots of those inputs.
solid answer
~40 sA **static lineage graph** describes the dependencies as they are *now*. Since last quarter, jobs may have been rewritten, sources swapped, tables backfilled or corrected in place — so today's graph can point at inputs the report never read. To answer the auditor you need **run-level provenance**: for the report's publishing run, which job and code version ran, which **input dataset versions** or snapshots it read, and the timestamps. That record must be **captured at publication time** and retained with the report, together with a way to retrieve those exact input versions — a snapshot, a retained table version, or an archived extract. Without that you can only describe how the report would be built today, which is not what the auditor asked.
go deeper
Know that a lineage graph shows current dependencies, and that a past report may have been built from different logic or data.
Explain run-level provenance, which facts it records, and why it must be captured at publication time.
Design the evidence a regulated report keeps, including input versions and how long they stay retrievable.
Decide with legal and finance owners which reports need reproducible inputs, and how that retention is reconciled with deletion duties.
## Two kinds of lineage record - **Design-time (static) lineage**: the dependency graph between datasets and jobs as currently defined. It answers "what feeds what" today. - **Run-level lineage (provenance)**: records of individual executions — this run of this job, at this time, with this code version, read these versions of these inputs and wrote this version of that output. It answers "what fed *this particular* output". An audit question about a report already published is a **run-level** question. ## Why today's graph misleads | What happened since the report | Effect on a static graph | |---|---| | The transformation was rewritten | the graph shows the new logic, not the logic that ran | | A source was replaced by another feed | the graph points at a source the report never read | | An input table was backfilled or corrected | the table now holds different rows than it did then | | A dependency was added or removed | the graph's edges no longer match that run | In each case, describing the report from the current graph produces a confident but wrong account. ## What an adequate answer needs 1. **The publishing run identified**: the run that produced the report, with its time and the code version of each job on the path. 2. **Input versions recorded**: for each input dataset, which version the run read — a snapshot identifier, a table version, a partition and load timestamp, or a file manifest. 3. **Those inputs still retrievable**: the version must still exist. If snapshots were expired or the table was rewritten in place, the record points at data that is gone. 4. **The evidence stored with the report**: the run record, input versions and code versions kept for as long as the report must be defensible. ## Retention tension Keeping exact input versions costs storage, and it collides with **deletion duties**: an erasure request may require removing a person's rows from those very inputs. Teams usually keep the **record** of what was read for the report's retention period, and decide per dataset whether to retain the data itself or a reproducible aggregate. That decision belongs with whoever owns the report's regulatory obligations. ## Why interviewers ask it It tests whether the candidate knows the difference between a dependency diagram and **provenance evidence**. Regulated reporting, finance close and audit all ask this question, and the strong answer says the evidence has to be captured **when the report is published**, not reconstructed later.
- How would you make this answerable for every future quarterly report?Make the publishing job record its run — code version, input dataset versions and output version, and store that record with the report. Retain the referenced input versions for the report's retention period, or archive the exact extract the report was built from.
- The input table was corrected in place after publication. What can you still say?If you kept the run record but not the old rows, you can say which table and load the report read and what was corrected since, but you cannot reproduce the figure. That gap is why in-place corrections to reported data need their own record.
saying these in an interview costs you the question
- Answering an audit question from today's dependency graph
- Assuming input tables are unchanged since the report ran
- Planning to reconstruct provenance after the auditor asks
- Ignoring that expired snapshots make recorded versions unretrievable