skip to content

How do you mask a field whose format the code under test parses and validates?

level: middleimportance: should knowfreq 52%

answer

  1. some fields are parsed, not merely stored
  2. the replacement must still pass validation
  3. same length, same positions, same separators
  4. recompute the check digit, never copy it
  5. run masked rows through the real validator

basics

~20 s

Replace it with a value that satisfies the same rules - same length and character classes, same internal segments, a recomputed check digit - then run the replacement through the same validation routine the product uses before publishing the dataset.

solid answer

~50 s

When a field's **format is load-bearing** - the product splits it into parts, matches it against a pattern, or verifies a check digit - a replacement that ignores the format never reaches the logic you meant to exercise. The record is refused at the input boundary, and the case quietly becomes an input-validation test. Format-preserving masking replaces the value with a different value of the same shape: same length, same character classes in the same positions, the same separators, the same internal segments, and a check digit recomputed over the new body rather than copied from the old one. Where one field's format must agree with another column, the pair is replaced together. The check is mechanical: feed the masked rows through the same validation routine the product uses on input, and fail the masking job, not the suite, when a row does not pass.

code

pseudocode · 11 lines
pseudocode
mask_account_number(original):
   prefix    = original[0..3]                 # region segment the product parses
   body      = random_digits(length(original) - 5)
   candidate = prefix + body
   return candidate + check_digit(candidate)  # recomputed, not copied

verify_masked_dataset(masked_rows, source_rows):
   for row in masked_rows:
      assert product_input_validator.accepts(row.account_number)
      assert row.account_number not in real_account_numbers
   assert count(masked_rows) == count(source_rows)

go deeper

for a junior

Know that some fields are parsed and validated rather than simply stored, so a replacement has to look like a legal value of that field. Replacing an identifier with a row of asterisks usually means the record never gets past the input check at all.

for a middle

Explain what has to be held constant: length, character positions, separators, internal segments and a recomputed check digit. Be ready to say why copying the original check digit onto a new body is worse than data that is obviously wrong.

for a senior

Demonstrate that the masking job verifies its own output against the routine the product uses on input and fails there, rather than letting a malformed dataset surface as a mysterious suite failure. Mention drawing replacements from a reserved range so a masked value can never be genuine.

for a principal

Own the position that dataset validity is part of the test estate's contract rather than each team's courtesy, and decide when a field whose format is itself identifying should be dropped from the extract and covered by hand-built data instead.

## When a field's format is load-bearing Most fields in a dataset are read as opaque text: a name is rendered and compared, nothing more. Some are not. An account identifier is split into a prefix and a body; a long numeric identifier carries a final check digit computed over the digits before it; a postal code is matched against a pattern before the address is accepted; a reference is parsed into a date part and a sequence part that the product then sorts on. For those fields the format is not decoration - it is an input contract, and the code under test enforces it. Replace such a value with a row of `XXXXXXXX` or a random string of the wrong length, and the record never reaches the behaviour the case was written for. It is refused at the edge: the load fails, or the screen rejects the row, or the parse throws before the interesting branch runs. What is left is a suite that exercises input validation very thoroughly and the business rule not at all - and because the failure looks like a broken test rather than a broken dataset, it often gets closed by loosening the assertion. ## What the replacement has to hold Format-preserving masking means the substitute is a member of the same set as the original. In practice that decomposes into: - **Length and character classes.** Same number of characters, digits where digits were, letters where letters were, and the same case pattern if anything compares case. - **Separators and fixed positions.** Punctuation, spacing and any fixed marker stay where they were, because the parser either counts offsets or splits on the separator. - **Internal segments.** If the value decomposes into meaningful parts, each part is replaced with a legal member of its own part - a region segment with another real region segment, a date part with a real date. - **Derived digits.** Any check digit is recomputed over the replaced body. Copying the original digit onto a new body produces a value that fails validation in the most annoying way possible, because it looks correct to a human reading it. - **Agreement with other fields.** Where another column must agree with part of this one, both move together, or the pair is drawn from a table of legal combinations. - **Declared constraints.** If the column is declared unique or is used as a key inside the dataset, the replacement has to respect that or the load itself fails. | Property of the original | Kept? | Why | |---|---|---| | The exact value | no | that is the entire point of the exercise | | Length and character pattern | yes | parsers and fixed-width readers depend on it | | Check-digit validity | yes, recomputed | validators refuse an inconsistent value | | Real-world referent | no | the substitute must not denote a real subject | | Sortability of a date-like part | usually | cases frequently order on it | ## Verify the dataset, not the suite The check that matters is to run the masked output through **the same validation routine the product uses on input** - the routine itself, not a re-implementation of it - and to fail the masking job when a row does not pass. A re-implementation drifts from the real rules, and the drift surfaces as a test failure weeks later, in a team that has no idea a masking rule changed. Two further checks are cheap and catch most of the rest: a row count that matches the source, and a sample of masked rows walked through the product's own parse-and-render path. It is also worth asserting the negative: that no masked value is itself a real one. A generator producing legal values from the same space will occasionally emit a value that belongs to a genuine subject. Drawing replacements from a range the real world does not issue - a reserved prefix, a block set aside for documentation and testing - removes that problem rather than managing it. ## When the format cannot be preserved Sometimes the format is itself the sensitive part: a rare pattern only a handful of subjects carry. Preserving it preserves the giveaway, and no amount of care about check digits helps. The honest answer then is that the field is not a masking problem at all - the case that needs it moves to hand-built data, constructed to be legal rather than derived from anything real, while the extract empties the column. Recognising that at planning time is far cheaper than discovering it in a review after the dataset has already been copied into three environments.

  • What goes wrong if the masked value keeps the original's check digit?
    The digit no longer verifies the body it accompanies, so every row is refused - but only by the validator. In a listing or a screenshot the value looks entirely plausible, so the diagnosis usually starts at the wrong end, with someone assuming the validation rule has regressed. Recomputing the digit over the replaced body removes the whole class of confusion.
  • How would you stop a masked value from accidentally being a real one?
    Draw replacements from a range the real world does not issue - a reserved prefix or a block set aside for testing - so a collision with a genuine value is impossible rather than merely unlikely. Where no reserved range exists, check the output against the source set and redraw on a hit. Either way, make it something the masking job asserts, not something you hope for.

It is the difference between forging a plausible passport number and scribbling over the page: one still goes through the reader, the other is turned away at the desk before anyone looks at the traveller.

saying these in an interview costs you the question

  • Replaces a structured value with a fixed string of asterisks
  • Copies the original check digit onto a freshly generated body
  • Assumes any random string of the same length will validate
  • Masks one half of a paired code and leaves the other untouched
  • Learns the format broke from a suite failure rather than at masking time