A data-loss policy can match on a regex pattern, a document fingerprint or a sensitivity label - what does each detect?
answer
- three modes, three different failure shapes
- one guesses, one remembers, one is told
- shape of the data versus a known document
- metadata is a claim, not evidence
- print to PDF and the label may not follow
basics
~20 sA pattern rule matches content shaped like sensitive data, so it also fires on lookalikes. A fingerprint matches content that resembles a specific indexed document or record set. A label matches metadata someone attached at creation, not the content at all.
solid answer
~50 sThe three modes differ in what they know before they look. **Pattern matching** is a regular expression plus supporting evidence - a checksum, nearby keywords, a minimum number of hits - so it generalises to data the engine has never seen but fires on anything shaped alike, such as a 16-digit part number read as a card number. **Fingerprinting** indexes a known source: the document is cut into overlapping segments and hashed, or a real customer table is hashed cell by cell, and a candidate is scored on how much of it matches. Precision is high, but it only knows what was indexed. **Labels** are metadata written into the file when it was classified; the check is trivial and works even when the body is unreadable, but it is a claim about the file rather than evidence from it, and a print-to-PDF can leave the sensitive bytes with no label on them.
go deeper
Be ready to name the three modes and give one example of each failing: a part number read as a card number, an export made after the last index run, and a labelled file printed to PDF.
An interviewer expects the mechanics: checksums, proximity keywords and minimum match counts behind a pattern rule; segment hashing and a similarity threshold behind a fingerprint; and why a label survives editing but not re-rendering.
Show how the mode changes your next investigative move, and be precise that a hit proves a match at an enforcement point - not that data left, and not that anyone intended harm.
Own the coverage argument: which data classes you can realistically fingerprint, what labelling costs the business to apply and keep accurate, and where you accept that pattern matching is the only mode you can afford.
## What a content-inspection control is actually doing A data-loss control sits at an enforcement point - an agent on the endpoint, the outbound mail path, or a scanner attached to a SaaS tenant over its API - and asks one question of everything crossing it: *does this match something I was told to care about?* There are three families of answer. They cost different amounts, they generalise differently, and they fail in three completely different ways, which is why the mode is the first thing an analyst should look at on an incident. ### 1. Pattern matching The rule is a regular expression, almost never on its own. A usable card-number rule is a digit pattern *plus* a checksum test (a card number has a check digit, so most random 16-digit strings can be rejected outright), *plus* proximity keywords such as expiry or CVV within a few words, *plus* a minimum number of distinct matches in one file, all rolled into a confidence level. - **Strength:** it needs no prior knowledge of the document. A spreadsheet created five minutes ago is inspected as well as one that existed when the policy was written. - **Weakness:** everything shaped like the target matches. Order identifiers, internal part numbers, test files full of synthetic card numbers, a developer fixture of fake national ID numbers. These are *false positives* in the strict sense: the content is not what the rule claimed. - **What a hit proves:** bytes at this enforcement point are *shaped like* the data class. Nothing more. ### 2. Document and record fingerprinting Here the engine is given the real thing in advance. For unstructured documents, the source is cut into overlapping segments and each is hashed; the index stores hashes, not the document, so the index itself is not a new copy of the secret. A candidate file is processed the same way and scored on how many segments match - a similarity ratio and a matched-segment count. For structured data the same idea is applied to a record set: hash every cell of an exported customer table, then require, say, three fields from the same row to appear together before calling it a match. - **Strength:** precision. The reference is your actual data, so a hit is rarely a lookalike. - **Weakness:** it only knows what was indexed, and only as of when it was indexed. An export produced after the last index run is invisible. Retyped, summarised or heavily reformatted content falls below the similarity threshold. And shared boilerplate is a trap in the other direction: if the indexed source carried a standard header block or template, every document built from that template inherits some similarity. - **What a hit proves:** this content resembles a *specific indexed source* above a threshold. ### 3. Sensitivity labels A label is metadata attached at creation or classification - by a human choosing it, or automatically. The check is the cheapest of the three and it is the only one that still works when the engine cannot read the body at all, for example because the labelling stack encrypted the file as part of applying the label, or because the format is one the extractor does not parse. - **Strength:** it survives the content changing. Edit the document heavily and the label rides along; no pattern or fingerprint has that property. - **Weakness:** it is an assertion *about* the file, not evidence *from* it. A file nobody labelled is invisible to label rules no matter what is in it, and an over-labelled estate turns the mode into noise. Metadata also does not survive every transformation: printing to PDF, pasting into a fresh document, exporting a table to CSV, or photographing the screen can all produce bytes that are just as sensitive with no label attached. Some stacks propagate labels through some of those paths; never assume it without testing the specific path. - **What a hit proves:** someone, or some classifier, once said this file was sensitive. ## Why the mode changes the analyst's job Real policies combine modes, and the combination is what sets the false-positive rate: label OR fingerprint to catch known material, pattern with a high confidence threshold to catch the unknown. When an incident lands in the queue, the mode tells you which question to ask next. A pattern hit invites *is this really that kind of data?* A fingerprint hit invites *which indexed source, and how much of it?* A label hit invites *who applied the label, and is it still accurate?* And none of the three answers the two questions people most often assume they answer. A hit does not mean data left - inline modes fire *instead of* the transfer, and out-of-band scanners fire *after* it. A hit also says nothing about intent: if the content genuinely is customer data and the person moving it was authorised to move it, the finding is correct and the behaviour is fine. That is a **benign true positive**, and it is a different verdict from a false positive, where the rule was simply wrong about what the bytes were.
- Why does a card-number pattern still fire on an internal 16-digit part number, and what reduces that?Because a bare digit pattern only tests shape. Adding a check-digit test rejects most random 16-digit strings, requiring keywords such as expiry or CVV within a short window demands corroborating context, and setting a minimum match count stops a single incidental number from tripping the policy. Together they trade a little recall for a large drop in lookalike hits.
- A user prints a labelled spreadsheet to PDF and uploads the PDF. Which of the three modes still has a chance?The label mode probably loses, because the rendered PDF is a new file and the label metadata is not guaranteed to be carried onto it. Content-based modes can still fire if the agent extracts text from the PDF - a pattern rule on the visible numbers, or a fingerprint if enough segments survive the render. If the PDF is image-only and there is no OCR, all three miss.
- What is the difference between a false positive and a benign true positive on a data-loss hit?A false positive means the rule was wrong about the content - the 16-digit string was a part number, not a card. A benign true positive means the rule was right and the behaviour was fine: it really was the customer list, and the person moving it was authorised to move it. They are fixed differently. The first is a rule-quality problem; the second is a question about who is allowed to do what.
Pattern matching is a bouncer checking whether an ID looks like an ID; fingerprinting is checking it against a list of real ones on file; a label is trusting the sticker somebody put on the folder.
saying these in an interview costs you the question
- Says a data-loss hit means the data left the company
- Treats a sensitivity label as proof of what the file contains
- Assumes regex matching is exact and cannot produce lookalike hits
- Thinks fingerprinting stores the source document itself rather than hashes
- Calls every hit on authorised activity a false positive