skip to content

questions

4

What do substitution, shuffling, banding and blanking each preserve when masking a test dataset?

level: middleimportance: must knowfreq 62%

answer

  1. four ways to replace a field's values
  2. what each one leaves for the tests
  3. one keeps the distribution, rows lie
  4. ranges keep order, kill boundary cases
  5. an emptied field tests the missing path

basics

~20 s

Substitution swaps in a plausible replacement and keeps type and shape; shuffling keeps the column's real distribution but detaches values from their rows; banding keeps magnitude and order while losing precision; blanking keeps nothing and forces the empty-value path.

solid answer

~50 s

Masking a dataset is a **per-field decision**, and the four common transformations fail in different directions. - **Substitution** replaces the value with a plausible one drawn from a lookup set. Type, length and format survive, so parsers and screens still work; agreement with neighbouring fields does not. - **Shuffling** redistributes the column's own values across rows. The distribution is exactly right, which suits a case that reads a total or a spread, but every value now sits on the wrong record. - **Banding** replaces the value with the range it fell in. Ordering and rough magnitude survive; per-row precision and boundary cases do not. - **Blanking** removes the value, so nothing survives and every read of the field exercises the missing-value path. Choose per field from what the cases read, and record the choice beside the field so the next reader can see the trade you made.

code

pseudocode · 16 lines
pseudocode
masking_map:
  full_name   -> substitute(from: name_list, keep: length_class)
  postal_code -> substitute(from: valid_codes(region), keep: format)
  birth_date  -> band(width: 10 years)
  salary      -> shuffle(scope: column)
  notes       -> blank()

pass_one(rows):                    # row-local rules, one row at a time
  for row in rows:
    for field, rule in masking_map where rule.scope == row:
      row[field] = rule(row[field])

pass_two(rows):                    # column-global rules need every row first
  for field, rule in masking_map where rule.scope == column:
    values = collect(rows, field)
    assign(rows, field, permute(values))

go deeper

for a junior

Be ready to say what masking a test dataset is for and to name the common per-field choices: swap the value, shuffle the column, replace it with a range, or empty it. The point to carry is that the choice is made per field, not once for the whole file.

for a middle

Explain the mechanics of each transformation and what each one leaves behind - type and format, exact distribution, ordering, or nothing at all. An interviewer expects you to pair every choice with a field it suits and a case it would ruin.

for a senior

Show that you drive the choice from the cases that read the field, and that you can spot the failure where a masked dataset still passes for the wrong reason. Be ready to say how the choice is recorded, reviewed and re-applied on the next refresh.

for a principal

Own the trade between how much realism the test estate keeps and how much exposure it carries, including where the transformation runs and which choices that rules out. Be able to defend a standard set of per-field rules teams apply without re-arguing them each time.

## What masking means here Masking, in the test-data sense, is replacing the values in a dataset with substitutes so the dataset can live somewhere less protected than the system it came from. It is applied field by field, and the interesting part is not privacy - a blanked column is perfectly private - but what each transformation leaves behind for the tests that will read it. Four transformations cover most of the work. ### Substitution from a lookup set The real value is swapped for a plausible one drawn from a prepared set: names from a name list, street lines from a street list, an identifier regenerated to the same pattern. Type, length, character classes and formatting survive, so screens render, parsers accept and validators pass. What does not survive is any agreement the field had with its neighbours: a substituted city no longer matches the postal code beside it unless the substitution is built to move the pair together. ### Shuffling within a column The column keeps its own values; they are redistributed across rows. This is the only one of the four that preserves the distribution exactly - same minimum, maximum, spread, distinct count and long thin end, because the multiset is unchanged. It is excellent when a case reads an aggregate over the column and useless when a case reads the row: every salary now belongs to a different employee, so any rule comparing two fields of one record is nonsense. Shuffling also removes nobody - every real value is still in the file, which some reviewers refuse regardless of where the rows sit. ### Banding The value is replaced by the range it fell into: an exact age becomes 30-39, an exact balance becomes 10,000-24,999. Ordering and rough magnitude survive, and code that groups or filters by range still behaves. Precision does not survive, which kills boundary cases: if the rule under test changes at 18 and the band runs 10-19, no case can be written that sits one day either side of it. ### Blanking The value is removed and the field left empty. It preserves nothing except the column's presence, and it silently converts every downstream read into an exercise of the missing-value path - worth knowing, because that path is usually the one nobody meant to be testing. ## Choosing, and writing the choice down | Transformation | Preserves | Destroys | Good fit | |---|---|---|---| | Substitution | type, length, format, plausibility | agreement with other fields, real distribution | fields a parser or a screen reads | | Shuffling | the column's exact distribution | the row's own value, cross-field rules | aggregate and reporting cases | | Banding | ordering, rough magnitude | precision, boundary values | range filters, grouped reports | | Blanking | nothing but the column itself | everything | fields no case reads | Choosing well means asking three things about the field, in this order: - **What do the cases read from it?** A total, a sort order, a threshold, a format, or nothing at all. - **What is the least a replacement can keep and still serve that?** Keeping more than the tests need is not free. - **Who else consumes this dataset?** A column another team's cases depend on is not yours to empty quietly. Two practical rules follow: 1. **Record the transformation next to the field**, as a checked-in map rather than steps buried in an operations runbook. That map is the artefact a reviewer reads and the artefact a failing test is diagnosed against. 2. **Re-run the transformation; never hand-patch its output.** A dataset corrected in place is no longer reproducible, and the next refresh silently loses the correction. ## Where the transformation runs changes what is available There is a real split between estates that transform **while the extract is being taken** - each row rewritten as it streams out, so no unmasked copy ever lands outside the protected system - and estates that **land a full copy first** and transform it in place. The privacy argument favours the first: there is no window in which a complete real dataset sits in a weaker environment. But the choice also decides which transformations are even possible. Substitution, banding and blanking are **row-local**: they need nothing but the current row, so they work in a streaming pass. Shuffling is **column-global** - you cannot redistribute a column's values until you hold all of them - and so is anything that fits a distribution to the real data before replacing it. An estate that transforms in flight has quietly given up shuffling and distribution-fitted substitution; an estate that lands a copy first keeps them and pays with a real, if brief, exposure window. The usual resolution is to do the row-local work in flight and accept a narrow landed copy for the few columns that genuinely need the global pass.

  • When is shuffling a column the right choice, and when is it clearly the wrong one?
    Shuffling suits a case that reads the column in aggregate - a total, a spread, a range filter - because the set of values is unchanged. It is wrong whenever a case reads the row: the value now belongs to a different record, so any rule comparing two fields of one record, or checking that a value matches its owner, is testing nonsense. It also leaves every real value in the file, which some reviewers refuse outright.
  • A field is blanked and the suite still passes. Why is that not reassuring?
    Because the cases that read the field may now be exercising the missing-value path rather than the behaviour they were written for, while still asserting a result that happens to match. A green suite over an emptied column tells you the code tolerates absence, not that the logic works. Check that the branches those cases were meant to reach are still being entered against the masked dataset.
  • Why record the per-field choice as a checked-in map rather than as steps in a runbook?
    Because the map is read by two audiences who never meet: a reviewer asking what was removed, and an engineer diagnosing a case that behaves oddly over the dataset. A checked-in map is diffable, reviewable and reproducible, so a change to a field's treatment arrives as a visible change rather than as a mysterious shift in behaviour after the next refresh.

It is like retouching a group photograph: you can paint a different face onto each body, swap the faces between people, blur everyone into age brackets, or cut the heads out. Each hides the person, and each ruins a different thing you might have wanted the photograph for.

saying these in an interview costs you the question

  • Applies one transformation to the whole dataset instead of choosing per field
  • Assumes an emptied field is harmless because nothing can leak from it
  • Thinks shuffling anonymises rows, forgetting every real value stays in the file
  • Bands values without checking whether a case tests a boundary inside the band
  • Substitutes plausible values and expects agreement with other fields to survive
  • Hand-edits the masked output instead of re-running the transformation
open as a page

How do you mask a field whose format the code under test parses and validates?

level: middleimportance: should knowfreq 52%

basics

~20 s

Replace it with a value that satisfies the same rules - same length and character classes, same internal segments, a recomputed check digit - then run the replacement through the same validation routine the product uses before publishing the dataset.

open as a page

Why should a field's masking rule in a test dataset be chosen from what the tests assert, not its name?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A field's name says what it holds, not what the tests read from it. Choosing by name routinely destroys the property a case depends on - a boundary, an ordering, a format - while leaving untouched fields needlessly realistic.

open as a page

Why does replacing every value in a column with one placeholder pass privacy review yet break tests?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A constant column is unarguably private and destroys the column's shape: distinct values, spread, lengths, how many rows a filter returns. Cases that sort, group, deduplicate or search that field then pass or fail for reasons unrelated to the code.

open as a page