skip to content

How can column-level lineage show where a sensitive field such as a customer email ended up downstream, and where does that tracing break down?

level: seniorimportance: should knowfreq 38%

answer

  1. start from the classified source column
  2. follow derived columns downstream
  3. transformations that mask versus copy
  4. aggregates, free text, exports lose the trail
  5. propagate the tag, then verify

basics

~20 s

Start at the classified source column and follow its column-level lineage downstream to every derived column, noting which steps copy it and which mask or aggregate it. Tracing breaks at free-text fields, opaque code, exports and other unobserved copies.

solid answer

~50 s

With **column-level lineage**, each output column is linked to the input columns it was computed from. Starting from the column classified as sensitive — the customer email — I follow those links **downstream** and list every column derived from it. For each hop I record whether the step **copies** the value, **transforms but keeps it identifying** (lower-casing, concatenation), or **masks or removes** it (hashing with a secret key, dropping, aggregating away); some lineage records even carry a masking flag per transformation. That list lets me **propagate the classification tag** to derived columns and check that each has the right protection. It breaks down where the value hides inside **free-text or JSON fields**, passes through **code the lineage cannot parse**, or leaves through **exports and copies** the graph does not see — so I verify with content scans on the suspected destinations.

go deeper

for a junior

Know that column-level lineage can follow a sensitive column into the tables derived from it.

for a middle

Explain how to judge each hop as copy, masked or removed, and why derived columns inherit the classification.

for a senior

Show where tracing breaks, such as free text, opaque code and exports, and how scans and logs close those gaps.

for a principal

Decide which pipelines must emit column-level lineage and how tag propagation feeds access policy, erasure and breach scoping across the platform.

## Why trace a sensitive field Privacy and security teams repeatedly ask one question: **where did this personal data go?** It matters when classifying new tables, answering an erasure or access request, scoping a breach, or checking that analytics copies are protected. **Column-level lineage** — a mapping from each output column to the input columns used to compute it — gives the first answer without scanning every table. ## The tracing method 1. **Start from the classified source column**, for example `customers.email`, tagged as personal data. 2. **Walk downstream through column-level edges** to every column derived from it, across models, marts and extracts. 3. **Classify each hop by what the transformation does to the value:** | Transformation | Is the output still identifying? | |---|---| | Direct copy or rename | yes — same value | | Normalisation (trim, lower-case) or concatenation | yes | | Hash without a secret key | usually yes — common values can be recomputed and matched | | Keyed hash or tokenisation with a separately held secret | pseudonymous — still personal data if the secret can link it back | | Aggregation to counts over large groups | usually no, but small groups can still single people out | | Dropped | no | 4. **Propagate the classification tag** to every derived column that remains identifying, so access and masking policies keyed on the tag cover it automatically. 5. **Report the downstream set** to whoever asked — the privacy team, the erasure workflow, or the incident lead. Some lineage standards let a producer record this per transformation: the OpenLineage column-lineage facet describes each input field's transformation as direct or indirect and carries a **masking** flag, so a tracing tool can tell a copy from a masked derivation without reading the code. ## Where the trail breaks - **Free text and semi-structured payloads.** An email typed into a support note or nested in a JSON blob is not a column-level edge. - **Opaque code.** Dynamic SQL, user-defined functions or notebook code the lineage cannot parse leave a gap. - **Indirect influence.** A filter or join on the email changes which rows appear without copying the value; column lineage may record it as an indirect dependency or not at all. - **Unobserved copies.** Exports, spreadsheets and syncs to other tools leave the graph entirely. - **Aggregates over small groups.** A count grouped by a rare attribute can re-identify people even though no email column survives. ## Closing the gaps - Run **content scans** for the value pattern on suspected destinations, especially free-text columns. - Use **access logs** to find readers that took copies out of the platform. - Require new pipelines to emit column-level lineage, so the graph's coverage grows. ## Why interviewers ask it It connects lineage to privacy work. A senior answer shows the method (source column, downstream walk, per-hop verdict, tag propagation) and is honest that lineage is **necessary but not sufficient**: the known breaks need scanning and logs behind them.

  • Why is an unkeyed hash of an email treated as still identifying?
    Anyone with a list of candidate emails can hash them and match the results, because the same input always gives the same output and the function is public. A hash with a secret key held apart from the data resists that, but whoever holds the key can still link records, so the result is pseudonymous rather than anonymous.
  • How would this tracing feed an erasure request?
    The downstream set from the source column tells the erasure workflow which tables and extracts may hold the person's data. The workflow still has to locate the person's rows in each, and the known breaks — free text, exports — need their own search.

saying these in an interview costs you the question

  • Assuming any hashed column is anonymous
  • Treating the lineage walk as a complete inventory of copies
  • Ignoring free-text and JSON fields when tracing personal data
  • Tagging only the source column and not derived columns