skip to content

How can an approval test's normalisation hide a real regression, such as a stale-cache read serving an out-of-date price?

level: seniorimportance: should knowfreq 30%

answer

  1. Stability is bought with sensitivity
  2. A pattern matches more than you aimed at
  3. Sorting cures noise and cures detection
  4. Break the value on purpose, expect red
  5. An artefact of placeholders asserts nothing

basics

~20 s

Every normalisation removes information from the comparison. A pattern broad enough to tame a varying value also erases everything else it matches, so a wrong price behind a masked amount, or a wrong order behind a sort, no longer produces any difference to fail on.

solid answer

~50 s

Normalisation trades sensitivity for stability, and a rule that takes too much makes the artefact green on broken output. The common ways it happens: a pattern matched by shape rather than by field, so masking one variable amount erases every amount; sorting a collection before comparison when the rendered order is itself behaviour; eliding a whole block or line rather than the value inside it; and rounding or tolerating a numeric difference. The defences are the same in every case. Prefer injecting determinism over masking. Anchor each pattern to a named field. Audit a scrubber by deliberately corrupting the value it protects and confirming the test goes red — a normalisation nobody has tested is an untested assertion. Keep a handful of hand-written assertions on the invariants that must never drift, and treat an approved artefact that is mostly placeholder text as evidence it asserts almost nothing.

code

pseudocode · 10 lines
pseudocode
# too broad: erases every price, tax line and total as well
received = replace_pattern(received, DECIMAL_NUMBER, "<amount>")

# narrow: only the value that genuinely varies between runs
received = replace_pattern(received, "shipping: " + DECIMAL_NUMBER, "shipping: <amount>")

# audit the rule: corrupt what it must still protect, expect red
fixture.set_price("hardback-9781", 18.47)   # the stale figure
run_pack()
expect_failing_cases(at_least = 1)

go deeper

for a junior

Recall the core idea: anything a normalisation rule replaces is no longer compared, so a rule that covers more than it must leaves real faults invisible to the test.

for a middle

Be able to walk through the mechanism — a shape-matched pattern catching more than the varying value — and name the narrower alternative of anchoring the pattern to the field it belongs to.

for a senior

Demonstrate the audit habit: corrupt the protected value on purpose and require a red run. Talk about keeping written assertions on money and ordering invariants alongside the approvals.

for a principal

Own the risk framing. Decide which invariants may never be normalised away, make the audit a standing practice rather than a heroic one-off, and treat a pack that has never failed as a finding.

### The trade every normalisation makes An approval test compares produced output against an approved artefact byte for byte. Every normalisation rule you add — a scrubber, a sort, a rounding, an elision — removes some of those bytes from the comparison in exchange for stability. That is a legitimate trade, but it is a trade, and the failure mode of the discipline is making it carelessly: the suite stays green, run times stay fast, and the thing it stopped checking is exactly the thing that breaks. ### How masking happens in practice **Matching by shape rather than by field.** A varying value is tamed with a pattern describing what it looks like — any decimal number, any long token, any date-shaped string. The pattern then also matches everything else of that shape. Masking a variable shipping surcharge with a "decimal number" pattern also masks every unit price, every line amount, every tax figure and the total. **Sorting an ordered thing.** Sorting a rendered collection before comparison is the standard cure for an unspecified iteration order — and if the displayed order is part of the behaviour, it also cures the test of ever detecting a wrong sort. **Eliding structure, not values.** Removing an entire line or block because "that part varies" means the approved artefact no longer records that the block exists. A change that drops it altogether now compares equal. **Tolerance.** Comparing numbers with a tolerance, or rounding before comparison, hides drift smaller than the tolerance — which is often precisely the size of a rounding regression. **Granularity too coarse the other way.** Approving a summarised projection is good practice, but if the projection omits the field under test, the approval is silent about it. ### A worked incident The receipt renderer behind an online bookstore checkout is guarded by a **340-case regression pack** of approved receipts. Early on, a per-order shipping surcharge varied by run, and the fastest fix was a scrubber replacing every decimal-number token with `<amount>`. It worked: the pack went stable. Later a caching layer was introduced in front of the pricing service. A **stale-cache read** began serving a snapshot for one publisher's catalogue: a hardback that had been repriced from 18.47 to 21.95 kept rendering at the old figure on any request that hit the warm entry. Every receipt in the pack still contained the same words, the same lines, the same fields — and every amount on every line had been replaced by `<amount>` before comparison. All 340 approvals stayed green. The defect surfaced 11 days later from a finance reconciliation, not from the suite that existed to catch exactly this. Nothing about the incident is exotic. It is the ordinary consequence of a scrubber that removed a whole class of values in order to stabilise one of them. ### Defending against it *Prefer control to masking.* If the surcharge had come from an injected pricing stub with fixed values, no pattern would have been needed and the amounts would have stayed under comparison. *Anchor the pattern.* Match `shipping: <decimal>` rather than any decimal. The narrowest rule that stabilises the run is the correct one. *Audit each normalisation by deliberately breaking what it covers.* Change the price in the fixture, run the pack, and confirm it goes red. A normalisation whose failure has never been observed is an assertion nobody has tested. This is a cheap, repeatable check and the single most useful thing to say in an interview. *Read the approved artefact as a reviewer, not as a diff.* Ask what proportion of it is placeholder text, and ask the sharper question: if the behaviour were wrong, would this file still be identical? An artefact where the answer is "yes" is not protecting anything. *Keep targeted assertions beside the approvals.* Money invariants — total equals the sum of lines, tax derived from the taxable subtotal, currency consistent — belong in written assertions that no scrubber can reach. *Watch for tests that have never failed.* An approval test that has been green through every change to the code it covers is either covering nothing or masking everything. Both are worth a look. ### The framing that lands Normalisation is redaction. Redacting the date on a contract makes it reusable; redacting the amounts makes it impossible to tell you were overcharged. Redact the fewest characters that make the run reproducible, and prove by experiment that what remains can still fail.

  • How do you prove a normalisation rule is not too broad?
    By experiment. Change the protected value in the fixture — a price, an ordering, a field's presence — run the pack and require at least one case to fail. Keep that corruption check as a test of the test where the value is high-risk. A rule whose failure has never been observed is an assertion nobody has verified.
  • If narrow rules are better, why not remove normalisation entirely?
    Because genuinely environmental values would then fail the comparison on unchanged code, and a pack that reddens for no cause trains the team to re-approve without reading. That destroys more detection than any scrubber does. The target is the narrowest set of rules that makes the run reproducible, not zero rules.
  • A team asks to compare amounts with a small tolerance because a currency conversion drifts. What do you push back on?
    A tolerance hides every regression smaller than itself, and rounding faults are usually that size. Fix the source instead: pin the conversion rate in the test so the value is exact. If the drift is genuinely outside your control, isolate that one field behind a narrow rule and assert the invariant it feeds — for example that the total still equals the sum of the lines.
  • What signal in the suite would make you re-audit a set of approval tests?
    An approval test that has never failed across many changes to the code it covers. Either it is exercising nothing or its normalisation has removed everything that could differ. Pair that with an artefact-level check: how much of each approved file is placeholder text, and has that proportion been growing?

Normalisation is redaction: blacking out the date on a contract makes it reusable, but blacking out the amounts means nobody can tell you were overcharged.

saying these in an interview costs you the question

  • Widens a scrubber pattern until the suite goes green
  • Assumes a green approval pack means no regression escaped
  • Sorts rendered output without asking whether order is behaviour
  • Never verifies that a normalised test can still fail
  • Elides whole lines rather than the varying value inside them
  • Adds a numeric tolerance instead of pinning the varying input

context