skip to content

Why must an approval test normalise timestamps and generated ids before comparing output?

level: middleimportance: must knowfreq 41%

answer

  1. Byte equality is maximally sensitive
  2. Values that vary between runs
  3. Red on unchanged code teaches reflex approval
  4. Control the value before rewriting it
  5. Numbered placeholders keep identity visible

basics

~20 s

Because values that change every run make the received artefact differ from the approved one even when the code has not changed. The test then fails nondeterministically, so teams re-approve reflexively and the approval stops meaning anything.

solid answer

~50 s

An approval test compares produced output against a stored artefact, so any value that varies between runs — a clock reading, a generated identifier, a temporary path, an unspecified collection order — guarantees a difference on unchanged code. There are two cures, and the better one comes first. **Control** the value: inject a fixed clock and a deterministic identifier source, and sort collections whose order is not itself behaviour. Where you cannot control it, **scrub** it: rewrite the received text before comparison, replacing each varying value with a stable placeholder. Good scrubbing is narrow and anchored to a named field rather than to a broad pattern, and it numbers repeated values (`<order-id-1>`, `<order-id-2>`) so the artefact still records whether two ids were the same. Sloppy scrubbing is worse than none: it hides the very values the test is meant to guard.

code

pseudocode · 15 lines
pseudocode
function scrub(text):
    # anchored to the labelled field, not to every date-shaped string
    text = replace_pattern(text, "placed at: <ISO_INSTANT>", "placed at: <timestamp>")

    seen = empty_map()
    for value in find_pattern(text, "order id: <TOKEN>"):
        if value not in seen:
            seen[value] = "<order-id-" + text(size(seen) + 1) + ">"
        text = replace_literal(text, value, seen[value])

    if size(seen) == 0:
        fail_test("no order id found: the output shape changed")
    return text

assert_equal(read_file(approved_path), scrub(received))

go deeper

for a junior

Be ready to say why a timestamp in the output makes the comparison fail on every run, and to name the fix in one line: replace it with a stable placeholder or feed the code a fixed clock.

for a middle

An interviewer expects the mechanics: the categories of varying value, the received-text rewrite step, and why numbering repeated identifiers preserves information that a single constant placeholder throws away.

for a senior

Show that you weigh each normalisation against the sensitivity it costs, prefer injecting determinism to masking, and can name what a too-broad pattern stops detecting in a long-lived pack of artefacts.

for a principal

Own the systemic angle: unstable approvals train a team to re-approve without reading, so the normalisation policy is really a policy about whether the suite's human review survives.

### Why non-determinism is fatal here An approval test's assertion is byte equality between the received output and the approved artefact. That makes it maximally sensitive: any character that differs fails the test. So any value produced from outside the logic under test — the wall clock, an identifier generator, a random source, the process environment — turns the test into one that fails on unchanged code. That is the textbook definition of a flaky test, except here it is not an accident of timing; it is designed in. The damage is not the red build. It is what the red build teaches. When a test fails for reasons nobody caused, the fastest way to green is to re-approve the received artefact. Do that a few times and re-approval becomes reflex, at which point the human review that gives an approval its meaning is gone and every future regression is one keystroke away from being blessed. ### The usual sources - **Clock readings** — creation, placement and print timestamps; relative phrasing such as "2 minutes ago"; durations measured in the run. - **Generated identifiers** — order numbers, correlation ids, surrogate keys, idempotency keys. - **Ordering that is not specified** — records iterated from a container with no defined order, results merged from concurrent work, or a set rendered without a sort. - **Environment** — absolute paths, host and user names, locale-driven decimal separators and date formats, line-ending conventions, trailing whitespace. - **Randomness** — sampling, shuffling, jitter, anything seeded from the clock. - **Version and build markers** stamped into rendered footers. ### Two strategies, in order of preference **Control the value.** Inject a clock the test fixes to a chosen instant, and an identifier source that hands out a deterministic sequence. Sort collections in the production code where the order is genuinely part of the contract. Controlled output needs no rewriting at all: the artefact contains the real value, so the artefact stays readable and nothing is masked. **Scrub the value.** When the varying value is produced somewhere you cannot reach — a downstream service's correlation id, a formatter's locale handling — rewrite the received text before comparing. The rule of thumb is that a scrubber should be as narrow as you can make it: match a labelled field rather than a general shape, so that "placed at: 2026-04-17T09:14:02Z" becomes "placed at: <timestamp>" without touching anything else on the line. ### Scrubbing well Three details separate a scrubber that preserves the test's power from one that guts it. *Preserve identity.* Replacing every order id with one constant `<id>` destroys the information that two ids were equal, or different. Number them in order of first appearance and reuse the number for repeats. Then a bug that reuses one order's identifier for a second order still shows as a diff. *Preserve structure.* Scrub the value, not the field. Removing the whole line means the approved artefact no longer records that the field exists, so a regression that drops the field entirely passes. *Fail loudly when the pattern does not match.* If a scrubber expected one placed-at value and found none, the output shape changed — that deserves a failure, not a silent pass. ### A worked case The receipt renderer behind an online bookstore checkout emits, per order: a placed-at timestamp, an order id, a per-item list, a payment reference from the payment gateway, and a rendered footer with the build marker. Its **340-case regression pack** was unusable at first: every case differed on every run. The fix was mostly control — a fixed clock at a chosen instant and a counting order-id source in the test wiring — plus two narrow scrubbers for the values that come back from outside: the payment reference and the build marker. Item order was left alone deliberately, because the receipt's line ordering *is* behaviour, and sorting it before comparison would have hidden a wrong sort. ### The trap to name in an interview The moment a scrubber's pattern is broader than the values it must neutralise, the test stops guarding whatever else that pattern happens to match. Replacing "every decimal number" to tame a variable shipping surcharge also erases every price and every total, and from then on the pack of 340 receipts cannot fail on a wrong amount. Normalisation buys stability by giving up sensitivity; the whole craft is giving up as little as possible.

  • Why prefer injecting a fixed clock over scrubbing the timestamp out of the output?
    Because control leaves the artefact honest. The approved file still shows a real formatted instant, so it keeps guarding the format, the field's presence and the time zone rendering. A scrubber replaces all of that with a placeholder, and every property of the value it hides is a property the test no longer checks.
  • The output renders a set of contributor names in an unspecified order. Do you sort it before comparison?
    Sort it only if the order is genuinely not part of the contract — then sorting removes noise and hides nothing. If the rendered order is behaviour a user sees, sorting before comparison would mask a wrong sort. In that case fix the determinism at the source by giving production code a defined order.
  • How do you decide whether a scrubber is too broad?
    Ask what a bug in that region would look like, then cause one: change the value deliberately and confirm the test goes red. A scrubber that keeps the test green while the value is wrong is too broad. Counting how much of the approved artefact is placeholder text is a fast second check.

saying these in an interview costs you the question

  • Says re-approving on every run is normal practice
  • Replaces all values matching a broad pattern to be safe
  • Uses one constant placeholder for every generated id
  • Deletes the whole varying line instead of the value
  • Believes scrubbing is preferable to injecting a fixed clock
  • Sorts output before comparing without asking if order is behaviour

context