skip to content

Why does replacing every value in a column with one placeholder pass privacy review yet break tests?

level: seniorimportance: should knowfreq 36%

answer

  1. unarguably private, and quietly useless
  2. a column has shape, not just values
  3. distinct count, spread, empty ratio, selectivity
  4. ties make sorted paged cases flaky
  5. assert the shape before the suite runs

basics

~20 s

A constant column is unarguably private and destroys the column's shape: distinct values, spread, lengths, how many rows a filter returns. Cases that sort, group, deduplicate or search that field then pass or fail for reasons unrelated to the code.

solid answer

~50 s

A privacy reviewer asks one question - can a real person be read out of this? - and a column holding the same placeholder in every row answers it perfectly. That is why the change survives review, and why nobody mentions it to the team whose cases read the column. What is lost is the column's **shape**: how many distinct values it holds, how they spread, how long they are, how many are empty, and how many rows a filter returns. A constant collapses all of it. A uniqueness constraint fails at load; sorting ties every row so paged cases turn flaky; a search matches everything or nothing; a grouped assertion collapses into one bucket. The repair is a **shape contract** per column - minimum distinct count, empty ratio, length range, value range - asserted against the masked dataset before any suite runs, so the dataset fails instead of the tests.

code

yaml · 8 lines
yaml
shape_contract:
  customer.city:      { min_distinct: 200,  empty_ratio_max: 0.05, length: [2, 40] }
  customer.balance:   { min_distinct: 5000, empty_ratio_max: 0.01, range: [0, 250000] }
  customer.reference: { unique: true,       empty_ratio_max: 0.00, length: [12, 12] }

on_breach: fail_the_masking_job
recheck: every_refresh
publish_measured_values_with: the_dataset

go deeper

for a junior

Know that a column can be perfectly safe and still useless: if every row holds the same replacement, anything that sorts, searches or counts distinct values over that column stops meaning anything. Safety and usefulness are two separate checks on the same dataset.

for a middle

Explain what a column's shape is - distinct count, spread, empty ratio, lengths, how many rows a filter returns - and name a concrete case that each of those properties supports.

for a senior

Trace an intermittent failure back to a masking decision: ties in a sort making paged results non-deterministic, a search matching every row, a load breaking a uniqueness constraint. Then say how the masking job itself should have caught it.

for a principal

Own the position that a dataset carries a stated contract on both privacy and shape, checked on every refresh, and settle who signs it off when the two requirements pull against each other.

## Why the change survives review A privacy review of a test dataset asks one question: can a real person be read out of this? A column holding the same placeholder in every row answers it as completely as any transformation ever will. There is no residual signal, no rare value, nothing to link against. The reviewer signs it off, and quite reasonably. Nobody in that conversation is asking the other question, which is whether the dataset still supports the work it exists for. That question belongs to the people whose cases read the column, and they usually meet the change as a set of unexplained failures some days later. The gap is organisational rather than technical: the transformation was chosen against one requirement and graded against one requirement. ## What shape means Shape is the set of properties of a column that code notices even though no single row is interesting: - **Cardinality** - how many distinct values the column holds. - **Distribution** - how those values spread across the range, including the long thin end. - **Empty ratio** - what proportion of rows carry no value at all. - **Length spread** - the shortest and longest values, and where the bulk sits. - **Selectivity** - how many rows a typical filter on the column returns. - **Ordering** - whether a sort on the column yields a stable, meaningful sequence. A constant collapses every one of them at once: cardinality falls to one, distribution to a spike, length spread to a point, selectivity to all-or-nothing, and ordering to arbitrary. ## How it breaks things | What the case does | What a constant column does to it | |---|---| | Loads the dataset | fails immediately if the column is declared unique | | Sorts and pages | ties every row, so page boundaries move between runs and the case turns flaky | | Searches or filters | matches every row or none, so the assertion is trivially true or trivially false | | Groups and counts | collapses into one bucket, and the per-group assertion loses its meaning | | Follows a reference | points every child row at one parent, so a fan-out defect cannot appear | | Reads a rendered screen | shows a column of identical text, hiding truncation and layout defects | | Exercises the empty path | never runs it - or always runs it, if the placeholder is itself empty | The flaky paged case deserves a note, because it costs the most time. Sorting on a column where every value is equal leaves the order of tied rows unspecified, so the same query can return the same rows in a different sequence on different runs. The case fails intermittently, gets labelled unstable, and gets a retry wrapped around it - three steps away from the actual cause, which is a masking decision taken in another team's ticket. ## A shape contract, asserted before the suite runs The repair is to make shape an explicit, checked property of the masked dataset rather than an accident of the transformation: 1. **State the contract per column.** For each transformed column, record what the dataset promises: a minimum distinct count, an empty ratio within a tolerance of the source, a length range, a value range, and whether ordering has to be meaningful. 2. **Assert it in the masking job.** Compute the same statistics over the source and over the output, compare both against the contract, and fail the job on a breach. The dataset is what is wrong, so the dataset is what should fail. 3. **Publish the numbers alongside the dataset.** Anyone diagnosing a strange failure can then see in one place that the column they are looking at holds four distinct values where the source holds forty thousand. 4. **Re-check on every refresh.** Shape drifts as the source changes and as rules are added; a contract checked once is a comment. None of this argues for keeping more realism than the tests need. A column no case reads should be emptied, and emptying it is exactly the right call - the difference is that it is then a decision taken with both requirements in view and written down, rather than a privacy win that silently withdrew a capability. The rule of thumb is short: a transformation may remove meaning from a column, but it should never remove the column's shape by accident.

  • Which single statistic would you check first to catch a collapsed column?
    The distinct-value count against the source, expressed as a ratio. It is one number per column, it costs a single pass over the data, and it catches a fixed placeholder, an over-wide band and a truncated substitution in the same check. Empty ratio is the natural second, because emptying a column is the other transformation that silently changes which paths the cases take.
  • A paged case turns intermittent after a masking change. How do you connect the two?
    Look at what the case sorts on. If that column now holds very few distinct values, the order of tied rows is unspecified and page boundaries move between runs, so the failure is non-deterministic ordering rather than a defect in paging. Confirm it cheaply by adding a tiebreaker to the sort, then fix the dataset's cardinality instead of leaving the retry in place.

saying these in an interview costs you the question

  • Judges a masked column only on whether anything can leak from it
  • Believes a column of identical values is harmless because nothing is lost
  • Adds a retry to a paged case instead of asking why ordering became unstable
  • Checks masked data by looking at it rather than by measuring it
  • Assumes a suite that still passes proves the dataset kept its shape