When masking policies are attached to classification tags, what happens to a newly added column of customer emails that nobody has tagged, and how do you close that gap?
answer
- no tag, no policy
- fails open by default
- scan on schema change
- deny or mask until classified
- inherit tags from lineage
basics
~20 sWith no tag, no tag-based policy applies, so the emails are readable by anyone who can read the table. Close the gap by scanning new columns, inheriting tags along lineage, and defaulting unclassified columns in sensitive tables to masked.
solid answer
~50 sTag-based policies are **fail-open for untagged data**: the mask attaches to the classification tag, so a new `email` column with no tag is served **in clear text** to everyone who can read the table — the rule never sees it. I would close it at three points. **Detect**: run classification scans on every schema change and new table, and inherit tags from upstream columns through **column-level lineage**, since a copied email is still an email. **Default**: in schemas holding customer data, treat any **unclassified column as masked or restricted** until a steward reviews it — fail closed instead of open. **Prevent**: add a publishing check that blocks new columns in sensitive datasets unless they declare a classification. Then monitor the count and age of unclassified columns, because that number is the size of the gap.
code
pseudocode · 11 lines# evaluated for every column a query returns
function visible_value(column, value, reader):
tags = catalog.tags(column)
if tags is empty:
if schema_of(column).holds_personal_data:
return MASKED # fail closed until a steward classifies it
return value # non-personal schema: fail open, scans follow up
for tag in tags:
if tag in SENSITIVE_TAGS and not reader.has_clearance(tag):
return MASKED # any sensitive tag the reader lacks masks it
return valuego deeper
Know that a tag-based rule only applies to columns that carry the tag.
Explain why an untagged sensitive column is exposed and how scans on schema change detect it.
Design detection, a fail-closed default for sensitive schemas, publishing checks and a metric for unclassified columns.
Decide per domain whether unclassified data fails open or closed, and own the trade-off between protection and breaking consumers.
## How tag-based policies work Instead of writing a masking rule per table, a platform attaches the rule to a **classification tag**: *mask any column tagged `personal-email` unless the reader has clearance*. When a column is tagged, the policy applies wherever that column appears. This scales well — but only for columns that **carry the tag**. ## The gap A pipeline adds `contact_email` to an existing table. Nobody tags it. The masking policy is keyed on the tag, so it does not apply; the table-level grant lets analysts read the table; the emails are returned **in clear text**. No error is raised, and no alert fires. The policy is **fail-open** with respect to classification. ## Closing it | Layer | Control | What it catches | |---|---|---| | **Detect** | classification scans triggered by schema changes and new tables | new columns whose values look personal | | **Detect** | tag inheritance through column-level lineage | copies and renames of already-tagged columns | | **Default** | unclassified columns in sensitive schemas are masked or hidden until reviewed | anything detection missed | | **Prevent** | publishing checks require a declared classification for new columns in sensitive datasets | the gap at the moment it is created | | **Monitor** | count and age of unclassified columns, per domain | whether the gap is shrinking | ## Fail open or fail closed? A **fail-closed default** — mask what is not yet classified — protects data but can break dashboards when a harmless new column appears masked. Teams usually accept that cost for schemas known to hold customer or employee data, and keep fail-open defaults for clearly non-personal domains such as infrastructure metrics. The decision should be explicit and recorded per domain. ## Tag inheritance Much sensitive data arrives by **copying**: a staging model renames `email` to `customer_email`, a mart concatenates it into a contact string. Column-level lineage links these to the tagged source, so a platform can propose or apply the same tag automatically. Inheritance needs care: an aggregate or a keyed hash of the column may deserve a different tag, so inherited tags should be reviewable. ## Measuring the exposure Useful numbers for the governance dashboard: - unclassified columns in sensitive schemas, and their median age; - columns where a scan disagrees with the current tag; - reads of unclassified columns by users without clearance. ## Why interviewers ask it Tag-based governance is widely promoted, and its weak point is exactly this. A senior answer states the **fail-open behaviour** plainly and closes it with detection, a deliberate default, prevention at publication and a metric.
- Why not simply mask every new column by default everywhere?It breaks harmless changes and trains teams to request unmasking reflexively, which erodes the control. Defaulting to masked is worth its cost in schemas that hold personal data; elsewhere, detection and review are a better balance.
- Should inherited tags be applied automatically or proposed for review?Direct copies and renames can be applied automatically because the value is unchanged. Derived columns, such as a keyed hash or an aggregate, should be proposed for review, since their sensitivity may differ from the source.
saying these in an interview costs you the question
- Assuming tag-based masking protects columns nobody has tagged
- Relying on a one-time classification scan at onboarding
- Masking every new column everywhere without considering domain
- Applying inherited tags to aggregates without review