skip to content

How does automated sensitive-data discovery find personal data in a warehouse, and why do column-name rules alone miss much of it?

level: middleimportance: should knowfreq 45%

answer

  1. names lie, values tell
  2. sample the values
  3. patterns plus checksums
  4. dictionaries and entity models
  5. confidence, review, re-scan

basics

~20 s

Scanners sample column values and match them with patterns plus checksums, reference dictionaries and trained entity recognisers, then propose tags with a confidence score. Names alone miss badly named columns, free text and JSON that hide personal data.

solid answer

~40 s

A discovery scan combines **metadata** and **content**. Column names (`email`, `ssn`) give cheap hints, but names are often misleading — `field_7`, `notes`, `payload`. So the scanner **samples values** and tests them: **patterns** with validation such as a checksum for card numbers or national IDs, **dictionaries** of first names or cities, and **trained entity recognisers** for names and addresses inside free text. Each column gets a proposed class and a **confidence** from the share of sampled values that matched. Results are **reviewed** by the data owner rather than applied blindly, because false positives (order numbers that look like phone numbers) and false negatives (rare formats, sparse columns) are both common. Scans rerun on **new and changed columns**, and free-text and nested fields get deeper scanning because that is where personal data hides.

go deeper

for a junior

Know that scanners look at sample values as well as column names, and why names can mislead.

for a middle

Explain the signals a scanner combines, how confidence is computed, and the common false positives and negatives.

for a senior

Design the review loop, re-scan triggers and defaults for unreviewed columns, and secure the scanner itself.

for a principal

Decide how much classification to automate across the estate and how to measure and fund the residual manual review.

## The problem Classification only helps if it covers the columns that actually hold personal data. In a large warehouse nobody knows them all: pipelines copy data under new names, analysts create tables, and support tools write free text that customers fill with anything. **Automated discovery** proposes classifications at scale so humans review a list instead of hunting. ## Signals a scanner combines | Signal | How it works | Strength | Weakness | |---|---|---|---| | Column name and description | keyword and synonym matching (`email`, `dob`, `tax_id`) | cheap, no data read | misses badly named columns; misfires on `email_opt_in` | | Value patterns | regular expressions for emails, phone numbers, IDs | precise for structured formats | ID-like numbers match everything | | Validation checks | checksums (for example the Luhn check on card numbers), valid ranges | removes most false positives | only for formats that have one | | Dictionaries | lists of given names, surnames, cities | catches names in short fields | ambiguous words ("Rose", "Paris") | | Entity recognisers | trained models that find names, addresses in text | works on free text | slower; probabilistic | | Statistics | cardinality, uniqueness, value length | spots ID columns | cannot say *what* the ID is | ## How a scan runs 1. **Sample** values per column — enough rows to be representative, drawn across partitions, not just the first rows. 2. **Score** each column: the share of non-null sampled values matching each class, adjusted by name hints. 3. **Propose** a class and confidence; above a threshold the proposal is auto-applied as *suggested*, below it goes to review. 4. **Review** with the dataset owner, who confirms, corrects or rejects. 5. **Re-scan** incrementally: new tables, new columns and schema changes trigger a scan; free-text columns are re-sampled periodically because their content drifts. ## Why name rules alone fail - **Opaque names**: `attr_12`, `payload`, `notes` hold emails and phone numbers. - **Free text**: support comments and chat logs contain names, addresses and account numbers typed by people. - **Nested data**: JSON or semi-structured columns hide fields a name rule never sees. - **Misleading names**: `email_sent_at` is a timestamp; `phone_verified` is a flag. ## Operating it well - Treat scanner output as **proposals**, measured by precision and recall against reviewed samples. - Default **untagged columns to restricted** in sensitive schemas until reviewed. - Keep the scanner's own access **tightly controlled**: it reads raw values, and its findings list where the sensitive data is. - Scan **copies and exports** too, not only the primary warehouse. ## Why interviewers ask it Any platform with more than a few hundred tables needs this. The good answer names **content sampling** as the core, admits the error modes, and puts a **human review and re-scan loop** around the tool.

  • How do you keep the discovery scanner itself from becoming a privacy risk?
    Run it under a dedicated, audited identity with read access only where scanning requires it, keep sampled values in memory rather than storing them, and protect its findings, because a list of where every sensitive column lives is itself sensitive.
  • How would you measure whether discovery is good enough?
    Build a reviewed sample of columns with known classes and compute precision and recall per class. Track how many columns stay unreviewed and how long new columns wait before classification.

saying these in an interview costs you the question

  • Trusting column names as the only signal
  • Auto-applying every scanner proposal without owner review
  • Scanning once at onboarding and never again
  • Ignoring free-text and nested JSON columns