Why must an approval test normalise timestamps and generated ids before comparing output?
answer
- Byte equality is maximally sensitive
- Values that vary between runs
- Red on unchanged code teaches reflex approval
- Control the value before rewriting it
- Numbered placeholders keep identity visible
basics
~20 sBecause values that change every run make the received artefact differ from the approved one even when the code has not changed. The test then fails nondeterministically, so teams re-approve reflexively and the approval stops meaning anything.
solid answer
~50 sAn approval test compares produced output against a stored artefact, so any value that varies between runs — a clock reading, a generated identifier, a temporary path, an unspecified collection order — guarantees a difference on unchanged code. There are two cures, and the better one comes first. **Control** the value: inject a fixed clock and a deterministic identifier source, and sort collections whose order is not itself behaviour. Where you cannot control it, **scrub** it: rewrite the received text before comparison, replacing each varying value with a stable placeholder. Good scrubbing is narrow and anchored to a named field rather than to a broad pattern, and it numbers repeated values (`<order-id-1>`, `<order-id-2>`) so the artefact still records whether two ids were the same. Sloppy scrubbing is worse than none: it hides the very values the test is meant to guard.
code
pseudocode · 15 linesfunction scrub(text):
# anchored to the labelled field, not to every date-shaped string
text = replace_pattern(text, "placed at: <ISO_INSTANT>", "placed at: <timestamp>")
seen = empty_map()
for value in find_pattern(text, "order id: <TOKEN>"):
if value not in seen:
seen[value] = "<order-id-" + text(size(seen) + 1) + ">"
text = replace_literal(text, value, seen[value])
if size(seen) == 0:
fail_test("no order id found: the output shape changed")
return text
assert_equal(read_file(approved_path), scrub(received))go deeper
Be ready to say why a timestamp in the output makes the comparison fail on every run, and to name the fix in one line: replace it with a stable placeholder or feed the code a fixed clock.
An interviewer expects the mechanics: the categories of varying value, the received-text rewrite step, and why numbering repeated identifiers preserves information that a single constant placeholder throws away.
Show that you weigh each normalisation against the sensitivity it costs, prefer injecting determinism to masking, and can name what a too-broad pattern stops detecting in a long-lived pack of artefacts.
Own the systemic angle: unstable approvals train a team to re-approve without reading, so the normalisation policy is really a policy about whether the suite's human review survives.
### Why non-determinism is fatal here An approval test's assertion is byte equality between the received output and the approved artefact. That makes it maximally sensitive: any character that differs fails the test. So any value produced from outside the logic under test — the wall clock, an identifier generator, a random source, the process environment — turns the test into one that fails on unchanged code. That is the textbook definition of a flaky test, except here it is not an accident of timing; it is designed in. The damage is not the red build. It is what the red build teaches. When a test fails for reasons nobody caused, the fastest way to green is to re-approve the received artefact. Do that a few times and re-approval becomes reflex, at which point the human review that gives an approval its meaning is gone and every future regression is one keystroke away from being blessed. ### The usual sources - **Clock readings** — creation, placement and print timestamps; relative phrasing such as "2 minutes ago"; durations measured in the run. - **Generated identifiers** — order numbers, correlation ids, surrogate keys, idempotency keys. - **Ordering that is not specified** — records iterated from a container with no defined order, results merged from concurrent work, or a set rendered without a sort. - **Environment** — absolute paths, host and user names, locale-driven decimal separators and date formats, line-ending conventions, trailing whitespace. - **Randomness** — sampling, shuffling, jitter, anything seeded from the clock. - **Version and build markers** stamped into rendered footers. ### Two strategies, in order of preference **Control the value.** Inject a clock the test fixes to a chosen instant, and an identifier source that hands out a deterministic sequence. Sort collections in the production code where the order is genuinely part of the contract. Controlled output needs no rewriting at all: the artefact contains the real value, so the artefact stays readable and nothing is masked. **Scrub the value.** When the varying value is produced somewhere you cannot reach — a downstream service's correlation id, a formatter's locale handling — rewrite the received text before comparing. The rule of thumb is that a scrubber should be as narrow as you can make it: match a labelled field rather than a general shape, so that "placed at: 2026-04-17T09:14:02Z" becomes "placed at: <timestamp>" without touching anything else on the line. ### Scrubbing well Three details separate a scrubber that preserves the test's power from one that guts it. *Preserve identity.* Replacing every order id with one constant `<id>` destroys the information that two ids were equal, or different. Number them in order of first appearance and reuse the number for repeats. Then a bug that reuses one order's identifier for a second order still shows as a diff. *Preserve structure.* Scrub the value, not the field. Removing the whole line means the approved artefact no longer records that the field exists, so a regression that drops the field entirely passes. *Fail loudly when the pattern does not match.* If a scrubber expected one placed-at value and found none, the output shape changed — that deserves a failure, not a silent pass. ### A worked case The receipt renderer behind an online bookstore checkout emits, per order: a placed-at timestamp, an order id, a per-item list, a payment reference from the payment gateway, and a rendered footer with the build marker. Its **340-case regression pack** was unusable at first: every case differed on every run. The fix was mostly control — a fixed clock at a chosen instant and a counting order-id source in the test wiring — plus two narrow scrubbers for the values that come back from outside: the payment reference and the build marker. Item order was left alone deliberately, because the receipt's line ordering *is* behaviour, and sorting it before comparison would have hidden a wrong sort. ### The trap to name in an interview The moment a scrubber's pattern is broader than the values it must neutralise, the test stops guarding whatever else that pattern happens to match. Replacing "every decimal number" to tame a variable shipping surcharge also erases every price and every total, and from then on the pack of 340 receipts cannot fail on a wrong amount. Normalisation buys stability by giving up sensitivity; the whole craft is giving up as little as possible.
- Why prefer injecting a fixed clock over scrubbing the timestamp out of the output?Because control leaves the artefact honest. The approved file still shows a real formatted instant, so it keeps guarding the format, the field's presence and the time zone rendering. A scrubber replaces all of that with a placeholder, and every property of the value it hides is a property the test no longer checks.
- The output renders a set of contributor names in an unspecified order. Do you sort it before comparison?Sort it only if the order is genuinely not part of the contract — then sorting removes noise and hides nothing. If the rendered order is behaviour a user sees, sorting before comparison would mask a wrong sort. In that case fix the determinism at the source by giving production code a defined order.
- How do you decide whether a scrubber is too broad?Ask what a bug in that region would look like, then cause one: change the value deliberately and confirm the test goes red. A scrubber that keeps the test green while the value is wrong is too broad. Counting how much of the approved artefact is placeholder text is a fast second check.
saying these in an interview costs you the question
- Says re-approving on every run is normal practice
- Replaces all values matching a broad pattern to be safe
- Uses one constant placeholder for every generated id
- Deletes the whole varying line instead of the value
- Believes scrubbing is preferable to injecting a fixed clock
- Sorts output before comparing without asking if order is behaviour