skip to content

Data Governance, Lineage & Catalogs

Catalogs, lineage, classification of personal data, access policy, retention and data ownership. Interviewers probe it to see whether you can run a data platform, not only ship pipelines.

on this pageshow

explore

questions

page 1 of 2

How does role-based access control differ from attribute-based access control for governing who can read warehouse data?

level: juniorimportance: must knowfreq 55%

answer

  1. who you are versus what is true
  2. roles bundle permissions
  3. attributes on user and data
  4. role explosion at scale
  5. policies written once, evaluated per request

basics

~20 s

Role-based access grants permissions to roles and users to roles. Attribute-based access evaluates rules over attributes of the user and the data, such as department and a column's sensitivity tag, so one rule covers many tables.

solid answer

~40 s

With **role-based access control (RBAC)** permissions are attached to **roles** — `finance_analyst` may read these schemas — and users get roles. It is simple to reason about, but in a large warehouse roles multiply: finance-EU-read, finance-EU-read-with-PII, and so on, the problem known as **role explosion**. **Attribute-based access control (ABAC)** evaluates **rules over attributes** of the user (department, region, clearance), the data (classification tag, owning domain, region) and sometimes the context (purpose, time). One rule — *users may read rows whose region matches their own; columns tagged restricted are masked unless the user holds the restricted clearance* — then covers every table carrying those attributes. Most data platforms combine them: roles for coarse access to schemas, attributes and tags for row and column rules.

go deeper

for a junior

Be ready to define RBAC and ABAC with a warehouse example of each.

for a middle

Explain role explosion and how attribute or tag-based rules cover new tables without new grants.

for a senior

Show how you combine roles for coarse access with attribute rules for rows and columns, and how you govern the attributes.

for a principal

Decide the access model for the whole data estate, including who owns user and data attributes and how effective access is audited.

## The question access control answers For every query the platform must decide: **may this user read this data?** — possibly down to individual rows and columns. Two models dominate how that decision is expressed. ## Role-based access control In **RBAC**, permissions are granted to **roles**, and users are made members of roles. - `marketing_analyst` → read on the `marketing` schema. - `finance_analyst` → read on `finance`, but not on `payroll`. It is easy to explain and to review ("who has this role?"). The trouble starts when access depends on more than one dimension. Region, sensitivity and department multiply: `finance_eu_read`, `finance_eu_read_pii`, `finance_us_read`… This **role explosion** makes roles hard to review, and every new table needs grants to the right roles. ## Attribute-based access control In **ABAC**, a policy is a rule over **attributes**: - **User attributes**: department, region, employment type, clearance. - **Data attributes**: classification tag (public, confidential, restricted), owning domain, the region a row belongs to. - **Context attributes**: stated purpose, time, network location. Example rules: 1. A user may read rows whose `region` equals the user's region attribute. 2. Columns tagged `restricted` are masked unless the user holds the `restricted` clearance. 3. Contractors may not read datasets tagged `customer-personal`. Each rule applies to **every table carrying the attribute**, including tables created tomorrow. ## Comparison | Aspect | RBAC | ABAC | |---|---|---| | Unit of policy | role → permission on objects | rule over attributes | | Scales with | number of role combinations | number of rules and quality of attributes | | New table | needs new grants | covered if its tags and attributes are set | | Row and column rules | awkward (one role per slice) | natural | | Reviewability | easy: list role members | harder: must evaluate rules to see effective access | | Main risk | role explosion, stale grants | wrong or missing attributes silently change access | ## How platforms combine them Most real deployments use **roles for coarse access** (which schemas or domains a team may query) and **attribute or tag-based policies** for fine-grained row filters and column masks. The attributes themselves — a user's department, a column's classification — must be governed: an ABAC policy is only as correct as the tags it reads. ## Why interviewers ask it It is the entry question for data access governance. A good answer defines both, names **role explosion** as the reason ABAC exists, and notes that ABAC moves the risk to **attribute quality**.

  • What is the main operational risk of attribute-based policies?
    Wrong or missing attributes. If a column is not tagged as restricted, or a user's department attribute is stale, the policy grants or denies the wrong access with no visible error, so the tagging and user-attribute sources need their own governance.
  • Why is RBAC easier to review in an access audit?
    Because effective access is visible as role membership plus grants. With attribute rules an auditor must evaluate each rule against current attributes to know who can read what, which usually needs tooling.

saying these in an interview costs you the question

  • Solving every new access need by creating another role
  • Assuming attribute-based policies are correct regardless of tag quality
  • Believing RBAC and ABAC cannot be used together
  • Granting access per user instead of through roles or policies
open as a page

What do technical, business and operational metadata each describe in a data catalog, and who typically supplies each kind?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Technical metadata describes structure (schemas, types, location); business metadata describes meaning (descriptions, glossary terms, owners, classification); operational metadata describes behaviour (freshness, run history, usage, quality results). Systems supply the first and third, people the second.

open as a page

What are the four principles of data mesh, and what problem of centralised data platforms is each one meant to solve?

level: juniorimportance: must knowfreq 50%

basics

~20 s

Domain ownership moves analytical data to the teams that know it; data as a product makes it usable by others; a self-serve platform makes that affordable for every domain; federated computational governance keeps it interoperable through shared rules the platform enforces.

open as a page

When classifying warehouse columns, how do direct identifiers, quasi-identifiers and sensitive attributes differ, and why does each need a different control?

level: juniorimportance: must knowfreq 55%

basics

~20 s

A direct identifier names a person by itself (email, national ID). Quasi-identifiers identify people only in combination (birth date, postcode, sex). Sensitive attributes are what must not leak (diagnosis, salary). Each class needs a different protection.

open as a page

Why is deleting one customer's data from a data lake much harder than deleting their row from an application database?

level: middleimportance: must knowfreq 60%

basics

~20 s

Lake data sits in immutable files copied through raw, cleaned and curated layers, kept in old snapshots, backups and extracts. Deleting a person means rewriting files everywhere they appear, expiring old versions, and stopping replays from reloading them.

open as a page

A revenue dashboard shows a figure that finance says is wrong; how do you use lineage to find where the error entered?

level: middleimportance: must knowfreq 60%

basics

~20 s

Start at the dashboard and walk the lineage upstream one hop at a time, comparing each dataset's figure with a trusted reference. The first hop where the value diverges is where the error entered; then check what changed there recently.

open as a page

To protect an email column that analysts use as a join key, when would you tokenize it, encrypt it, or replace it with a keyed hash?

level: middleimportance: must knowfreq 55%

basics

~20 s

Use a keyed hash when analysts only need a stable join key and nobody needs the email back. Tokenize when an authorised process must recover it. Encrypt when the value itself must be recoverable by key holders. None of the three makes it anonymous.

open as a page

What can a data lineage graph tell an analyst about where a dashboard number comes from, and what can it not tell them?

level: juniorimportance: should knowfreq 45%

basics

~20 s

Lineage shows which sources and transformation steps feed the number, and in what order. It does not say whether the data is correct, what the metric is supposed to mean, or whether the logic matches the business definition.

open as a page

In data governance, how do the roles of data owner, data steward and data custodian differ?

level: juniorimportance: should knowfreq 45%

basics

~20 s

The data owner is accountable for a dataset and decides its use and access; the data steward maintains its definitions, quality and metadata day to day; the data custodian operates the systems and technical controls that store and protect it.

open as a page

What does a periodic access review of a data platform check, and how do read logs keep it from becoming a rubber stamp?

level: middleimportance: should knowfreq 40%

basics

~20 s

An access review asks each data owner to confirm that every holder of access to their data still needs it. Read logs show which grants were never used, so owners revoke on evidence instead of approving everything by default.

open as a page

How does a business glossary differ from a data dictionary, and why link glossary terms to the columns that implement them?

level: middleimportance: should knowfreq 40%

basics

~20 s

A glossary defines business concepts in plain language, independent of any system; a data dictionary describes the columns of specific tables. Linking terms to columns shows where each concept lives and exposes columns that claim a term but compute it differently.

open as a page

How do push-based and pull-based metadata ingestion into a data catalog differ, and how does each one go stale?

level: middleimportance: should knowfreq 40%

basics

~20 s

Pull ingestion has the catalog crawl sources on a schedule, so it is stale between crawls and misses short-lived objects. Push ingestion has systems emit metadata when things change, so it is fresh but depends on every producer emitting correctly.

open as a page

How do you define and enforce retention periods across a warehouse and a lake so that expired data is actually deleted on time?

level: middleimportance: should knowfreq 45%

basics

~20 s

Give every dataset a retention class tied to a named date column, partition by that date, run automated jobs that drop expired partitions and then physically purge old versions, and log evidence. Legal holds and exceptions are explicit, recorded overrides.

open as a page

An auditor asks which exact source data produced a regulatory report published last quarter; why is a static lineage graph not enough to answer?

level: middleimportance: should knowfreq 35%

basics

~20 s

A static graph shows today's dependencies, not what ran then. You need run-level records naming which job version ran, which input dataset versions it read, and when, plus retained copies or snapshots of those inputs.

open as a page

What is federated computational governance in a data mesh, and which decisions stay global while others stay with each domain?

level: middleimportance: should knowfreq 35%

basics

~20 s

Domain and platform owners jointly agree a small set of global rules, and the platform enforces them automatically. Global rules cover interoperability, such as how a shared entity is identified, security and discoverability; each domain decides its own data model.

open as a page

In a data mesh, what distinguishes a data product from a table a team happens to publish?

level: middleimportance: should knowfreq 45%

basics

~20 s

A data product has an accountable owner and is built for consumers - discoverable, addressable, self-describing, trustworthy with stated objectives, interoperable through global standards and secured by global access rules. A published table has none of those guarantees.

open as a page

How does automated sensitive-data discovery find personal data in a warehouse, and why do column-name rules alone miss much of it?

level: middleimportance: should knowfreq 45%

basics

~20 s

Scanners sample column values and match them with patterns plus checksums, reference dictionaries and trained entity recognisers, then propose tags with a confidence score. Names alone miss badly named columns, free text and JSON that hide personal data.

open as a page

In a data lake, table-level access rules are enforced by the query engine; why can a user sometimes still read the protected data, and how do platforms close that path?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Table rules run in the query engine, but the files sit in object storage with separate permissions, so a direct storage reader skips filters and masks. Deny users direct storage access and give only governed engines scoped, short-lived credentials.

open as a page

When masking policies are attached to classification tags, what happens to a newly added column of customer emails that nobody has tagged, and how do you close that gap?

level: seniorimportance: should knowfreq 35%

basics

~20 s

With no tag, no tag-based policy applies, so the emails are readable by anyone who can read the table. Close the gap by scanning new columns, inheriting tags along lineage, and defaulting unclassified columns in sensitive tables to masked.

open as a page

A data catalog launched a year ago is now ignored because descriptions are stale and listed owners have left; how do you make it trusted again?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Reassign ownership, automate every fact a system can supply, focus human effort on the most-used datasets, surface trust signals such as certification and freshness, deprecate dead assets, and embed the catalog in daily tools so it gets used and corrected.

open as a page

How would you fulfil a data subject's erasure request end to end across a data platform, from intake to proof of completion?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Verify the request, resolve the person to every identifier, locate their data through classification and lineage, delete or anonymise per store while honouring retention duties, notify processors, purge old versions, suppress re-ingestion, and log evidence.

open as a page

Why does a lineage graph miss some real consumers of a dataset, and how do you find the ones it cannot see?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Lineage sees only tools that emit or expose it, so ad hoc queries, notebooks, spreadsheet exports, dynamic SQL and external copies go missing. Warehouse query and access logs, plus owners, reveal readers the graph cannot see.

open as a page

How can column-level lineage show where a sensitive field such as a customer email ended up downstream, and where does that tracing break down?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Start at the classified source column and follow its column-level lineage downstream to every derived column, noting which steps copy it and which mask or aggregate it. Tracing breaks at free-text fields, opaque code, exports and other unobserved copies.

open as a page

In a data mesh, how do source-aligned and consumer-aligned data products differ, and who should own a dataset that combines several domains?

level: seniorimportance: should knowfreq 30%

basics

~20 s

Source-aligned products publish a domain's business facts close to how they happen; consumer-aligned products reshape and combine them for a group of use cases. A cross-domain dataset should be owned by a team whose purpose is serving that use, not by one source domain.

open as a page

What does k-anonymity guarantee for a released dataset, how is it achieved, and which attacks does it still leave open?

level: seniorimportance: should knowfreq 35%

basics

~20 s

k-anonymity guarantees each record shares its quasi-identifier values with at least k-1 others, achieved by generalising and suppressing values. It still leaks when a group shares one sensitive value or an attacker has background knowledge.

open as a page

Who should approve access requests to a domain's sensitive dataset, and how do you keep approvals fast without making them meaningless?

level: principalimportance: should knowfreq 30%

basics

~20 s

The data owner should be accountable for access to sensitive data, but not click every request. Tier data by sensitivity, auto-approve low tiers by policy, require owner review with a stated purpose for high tiers, make grants expire, and track lead time.

open as a page

Engineers want to keep raw event data with identifiers forever so pipelines can always be rebuilt; how do you decide what the platform actually keeps?

level: principalimportance: should knowfreq 30%

basics

~20 s

Weigh the real value of full rebuilds against the risk, erasure cost and legal limits of keeping identifiable raw data. Usually keep identifiable raw data for a bounded rebuild window, then keep pseudonymised or aggregated forms, with legal minimums as explicit exceptions.

open as a page

A company of a few hundred people asks whether to adopt data mesh; how do you decide whether a central data team is still the better model?

level: principalimportance: should knowfreq 35%

basics

~20 s

Adopt it only if the central team is a real bottleneck, the business has several distinct domains, domains can staff data ownership, and a self-serve platform exists or can be funded. Otherwise improve the central model, adding domain ownership step by step.

open as a page

When does crypto-shredding, encrypting each person's data with their own key and destroying the key on erasure, beat physically deleting rows in a lakehouse, and what does it cost?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Crypto-shredding encrypts each person's identifying fields with a per-person key; erasure destroys the key, making every copy unreadable, including backups and old snapshots. It wins where rewriting copies is impractical, but costs key management at scale and query-time decryption.

open as a page

showing 1–30 of 32