skip to content

Why should a field's masking rule in a test dataset be chosen from what the tests assert, not its name?

level: seniorimportance: should knowfreq 44%

answer

  1. the column list is not the input
  2. read the assertions, not the field names
  3. classify what each case depends on
  4. weakest transformation that still satisfies it
  5. a green suite can mean the branch vanished

basics

~20 s

A field's name says what it holds, not what the tests read from it. Choosing by name routinely destroys the property a case depends on - a boundary, an ordering, a format - while leaving untouched fields needlessly realistic.

solid answer

~50 s

Two fields with the same kind of name can need opposite treatment. A date of birth that no case reads can be emptied; the same field in a product whose pricing changes at a fixed age must keep enough precision for a case to sit either side of it. So the input to the decision is the **assertions**, not the column list. The method is to walk the cases that touch each field and classify the dependence: does anything parse it, order it, count distinct values, compare it against another field, or branch on a threshold inside it? Then pick the weakest transformation that still satisfies that dependence and the privacy requirement. Where the two conflict, the case moves to hand-built data rather than the transformation being weakened. The dangerous outcome is not a red suite but a green one, where the branch the case aimed at is no longer reachable.

code

pseudocode · 15 lines
pseudocode
for field in fields_requiring_transformation:
    dependences = empty_set
    for case in cases_reading(field):
        dependences += classify(case)
        # parsed | ordered | bounded | distinct | related | present | none

    rule = weakest_rule_satisfying(dependences, privacy_requirement(field))
    if rule == NOTHING_SATISFIES_BOTH:
        move_to_hand_built_data(cases_reading(field))
        rule = blank()

    masking_map[field] = rule

# the check that matters, after the dataset is built
assert branches_entered(masked_dataset) == branches_entered(hand_built_dataset)

go deeper

for a junior

Know that what happens to a field depends on how the tests use it, and that the field's name does not tell you that. If you are unsure, find the cases that read the field and say what each one checks about it.

for a middle

Be able to walk a field's cases and name what they depend on - a format, an ordering, a threshold, a count of distinct values - then pick a transformation that keeps exactly that and nothing more.

for a senior

Show that you hunt the silent failure: a case that still passes because the branch it aimed at is no longer reachable. Be ready to describe a check that catches it, and to say when a case should move to hand-built data instead.

for a principal

Own the boundary between derived and hand-built data across the test estate, and hold the line that a red test is never a reason to weaken a transformation. Decide who arbitrates when a case and a privacy requirement genuinely cannot both be served.

## Why the field's name is a weak signal A masking plan written from the column list is written from names: anything that looks like a person's name gets a substitution, anything that looks like a date gets emptied, anything numeric gets a band. It is fast and it is auditable, and it is wrong about as often as it is right, because the name describes what the field *holds* while the decision depends on what the tests *read*. Two examples make the point. - A date of birth in a product that only displays it can be emptied without consequence. The same field in a product whose pricing changes at a fixed age is load-bearing: band it into ten-year ranges and no case can be written that sits one day either side of the boundary. - A free-text note field looks harmless and is often the most dangerous column in the extract, because anything at all may have been typed into it. Nothing in the name tells you that either. ## Classify the dependence, then pick the weakest transformation For each field already marked as needing a transformation, walk the cases that touch it and ask what kind of dependence they have. There are only a few kinds: 1. **Parsed or validated** - something splits, matches or verifies the value, so the replacement must be a legal member of the same set. 2. **Ordered** - a case sorts, ranges or compares. Order must survive; precision may not. 3. **Bounded** - a rule changes behaviour at a specific value, so that value must stay expressible. 4. **Distinct-counted** - a case groups, deduplicates or asserts how many different values exist, so distinct values must stay distinct. 5. **Related** - the field must agree with another field, or with a row elsewhere in the same dataset. Both sides move together or neither moves. 6. **Present or absent** - the case only cares whether the field is filled, so almost anything is permissible, including emptying it. 7. **Untouched** - no case reads the field at all. Empty it and stop paying to keep it realistic. Then choose the weakest transformation that satisfies the dependence - weakest meaning the one that keeps the least. Keeping more than the tests need is not free: every property retained is a property a reviewer has to think about, and one more thing that might carry signal out of the protected system. | Dependence | Usually fits | Usually does not | |---|---|---| | Parsed or validated | format-preserving substitution | emptying, or a fixed placeholder | | Ordered | banding, or order-preserving substitution | shuffling within the column | | Bounded | substitution with the boundary populated | banding across the boundary | | Distinct-counted | substitution that keeps distinct values distinct | banding, or a fixed placeholder | | Related | paired substitution from legal combinations | independent per-field substitution | | Present or absent | anything, including emptying | - | | Untouched | emptying | anything more expensive | ## The failure that does not look like a failure A transformation that breaks a case usually announces itself: the suite goes red, somebody reads the masked value and repairs the plan. The expensive outcome is the opposite. The case still passes, but the branch it was written to reach is no longer reachable - every row now falls on one side of the threshold, or every value now sorts identically, so the assertion holds for a reason with nothing to do with the behaviour under test. Two habits catch it: - **Compare branch execution counts** between a run over hand-built data and a run over the masked dataset. A branch that used to be entered and now never is has been masked out of existence. - **Prove the transformation can still fail.** For each field whose dependence is bounded, keep a case that must fail when the boundary is misread. If it passes over the masked dataset no matter what the code does, the dataset has stopped testing the rule. ## When the requirement and the case cannot both be met Sometimes no transformation both keeps what the case needs and removes what the reviewer needs removed, because the precision the rule turns on is the precision that identifies the subject. The answer is not to weaken the transformation. It is to move that case off derived data entirely and build the record by hand - legal, precise and referring to nobody - while the extract treats the field as harshly as the reviewer wants. Derived data is for breadth and awkward realism; hand-built data is for the cases that assert on exact values. Keeping that line clear is what stops a masking plan being renegotiated every time a test goes red.

  • How do you find out which cases actually depend on a field before you transform it?
    Start from the assertions rather than the schema: search the suite for the field and read what each case does with it - parse, order, compare, count distinct, branch on a value. Where the suite is too large to read, transform that one field, re-run, and watch what changes, including which branches stop being entered. The output is a short dependence note per field, which is also what a reviewer wants to see.
  • A case goes red the day a new masking rule ships. How do you tell a dataset problem from a code problem?
    Re-run the same case against a record built by hand for it. If it passes there and fails only over the masked dataset, the transformation removed something the case depends on and the plan is wrong. If it fails against both, the code changed. Making that a two-minute check is worth more than any amount of arguing about whose change broke the build.

saying these in an interview costs you the question

  • Writes the masking plan from the column list alone
  • Weakens a transformation because a test went red
  • Treats a green suite over masked data as proof nothing broke
  • Empties a field that a rule branches on at a threshold
  • Keeps realistic values in fields no case ever reads