skip to content

Why does a regenerated WireMock mapping set diff noisily against the committed meter-API set?

level: seniorimportance: should knowfreq 47%

answer

  1. noise first, signal underneath
  2. identity churn, not behaviour change
  3. every regenerated mapping gets a fresh id
  4. generated file names under mappings/ and __files/
  5. read both through the admin API, then normalise

basics

~20 s

Most of the diff is churn, not change. WireMock gives every regenerated mapping a fresh id and a generated file name, and volatile response values differ per call. Normalise both sets, then compare request and response pairs.

solid answer

~50 s

Almost none of the noise is a change in the district heating meter API. In WireMock every mapping the recorder produces gets a freshly generated `id`, and the file written under `mappings/` is named from the request plus a generated suffix, so a re-run renames and re-identifies every file even when behaviour is byte-identical. Extracted bodies under WireMock's `__files/` get generated names too. On top of that sit genuinely volatile values — a `Date` response header, a `readingTakenAt` stamp, the current `flowTemperatureC` — which differ on every call. The fix is to stop diffing files and start diffing behaviour: load each set into its own WireMock instance, read both through `GET /__admin/mappings`, delete `id` and `name`, sort by method and URL, and compare that. What survives normalisation is a real difference worth a reviewer's time.

code

bash · 9 lines
bash
# dump one WireMock instance's mappings, stripped of per-run identity, in stable order
norm() {
  curl -s "http://localhost:$1/__admin/mappings" | jq -S '[.mappings[] | del(.id, .name) | {request, response}] | sort_by(.request.method, (.request.url // .request.urlPath // ""))'
}

norm 8080 > committed.json    # instance rooted on the committed mappings/
norm 8081 > refreshed.json    # instance rooted on the regenerated set

diff -u committed.json refreshed.json

go deeper

for a junior

Know that a re-recording does not overwrite the old files neatly: WireMock writes new files with generated names and fresh ids, so a raw directory diff shows everything as changed.

for a middle

Explain the mechanism behind the churn — per-run mapping ids, generated file names under WireMock's mappings/ and __files/, and volatile response values — and why comparing the admin API's JSON is more useful than comparing files.

for a senior

Demonstrate the normalisation you would actually run: two instances, both sets read through the admin API, identity fields deleted, entries sorted, volatile fields scrubbed symmetrically, and only then a diff a reviewer can read.

for a principal

Own the question of who reads this diff and how often, and what stops it becoming a report nobody opens. Decide whether normalisation lives in a shared script, and where the line sits between automated production and human acceptance.

## The problem: a diff that is 100% churn You re-record the district heating meter API, put the regenerated set beside the committed one, and run `diff -r`. Every file is reported as changed or renamed, including the mappings for endpoints nobody has touched in a year. The instinct is that the upstream has changed enormously. It has almost certainly changed hardly at all. The reason is that a WireMock mapping file carries two quite different kinds of content: the **behaviour** — which request it matches and what it answers with — and the **identity** the recorder minted for it on this particular run. A file-level diff cannot tell those apart, so it reports the identity churn at the same volume as a real change, and the real change drowns. ## The sources of noise, in rough order of volume 1. **A fresh mapping id per run.** Every regenerated mapping is a new mapping as far as WireMock is concerned, so it carries a newly generated `id`. That alone changes every file. 2. **Generated file names.** WireMock's recorder derives the file name under `mappings/` from the request and adds a generated suffix, so the second run writes differently named files rather than overwriting the first run's. 3. **Body file names.** When WireMock's recorder extracts a response body into `__files/`, that file gets a generated name too, and the mapping's reference to it changes with it. 4. **Volatile response headers.** A `Date` header, or any per-request correlation header the meter API returns, differs on every single call. 5. **Volatile payload values.** `readingTakenAt`, a `generatedAt` stamp, the current `flowTemperatureC` for a substation — real data that legitimately moves between two captures minutes apart. Only a sixth category matters: the same request now producing a **structurally different** response — a new field such as `tariffBand`, a field that changed type, a status code that moved from `200` to `404`, a header the committed matchers depend on that has disappeared. ## Normalising before you compare The reliable technique is to compare the two sets as *data*, not as directory trees: 1. Start one WireMock instance rooted on the committed set and a second rooted on the regenerated one, on different ports. 2. Read both through WireMock's `GET /__admin/mappings`, which hands you the mappings as JSON regardless of how they are laid out on disk. 3. Delete the per-run identity fields — the mapping `id` and its `name` — from each entry. 4. Sort the entries by method and URL so the two lists line up positionally. 5. Diff the results. That one pass removes categories 1 through 3 entirely. Categories 4 and 5 need a second decision: either scrub the known-volatile fields and headers out of both sides before comparing, or leave them in and accept that a reviewer will skim past them. Scrubbing is better when the refresh is frequent, because an unscrubbed diff trains people to ignore the output. ## Which diff entries mean what | what you see | what it usually means | |---|---| | same request, same response, different id or file name | pure churn — normalisation should have removed it | | same request, response differs only in a timestamp or a temperature | volatile data, not a change | | same request, response gained or lost a field | a real upstream change worth reviewing | | a mapping only in the regenerated set | the client made a call the committed set never covered | | a mapping only in the committed set | usually a gap in your replay, occasionally a removed endpoint | The last row is the one people misread most often. A missing regenerated mapping is far more likely to mean the replay never made that call than that the meter API dropped the endpoint, and the two conclusions lead to opposite actions. ## Keeping the noise down next time - **Replay from the client, deterministically.** The same calls, in the same order, with the same inputs, so two captures are comparable in the first place. - **Pin what you can.** A fixed `meterId` and a fixed readings window remove a whole class of incidental difference. - **Keep the normalisation as a script in the repository**, not as something each engineer reinvents when they run a refresh. - **Diff behaviour, not bytes.** Two mappings that match the same request and return the same status, body shape and relevant headers are equivalent, whatever their files look like. - **Review the diff, do not import it.** The output is a list of questions for a person, not a patch to apply. Once normalised, a refresh of a stable API should produce a diff you can read in a minute or two. If it does not, the normalisation is incomplete — and that is worth fixing before the next refresh, because a noisy diff is one that nobody reads twice.

  • Your normalised diff is still large. What do you check before concluding the meter API changed?
    Check that the replay was deterministic — same meter identifiers, same readings window, same call order — and that you scrubbed the volatile fields on both sides, not just one. A large diff after a comparable replay and symmetric scrubbing is worth believing; before that, it is more likely a difference in how the two captures were driven.
  • Should the refresh diff be produced automatically in CI?
    Producing it can be automated; acting on it should not be. The diff needs a reader who can tell a new `tariffBand` field from a moved timestamp. A useful middle ground is a scheduled job that regenerates, normalises and publishes the diff as a reviewable artefact, while leaving the decision to accept any part of it with a person.

It is like diffing two photographs taken from a handheld camera: almost every pixel has moved, so you cannot see the one object that actually changed until you line the frames up first.

saying these in an interview costs you the question

  • Reads a churn-heavy diff as evidence the upstream changed
  • Diffs directory trees instead of normalised mapping data
  • Forgets that regenerated mappings carry freshly generated ids
  • Scrubs volatile fields on one side of the comparison only
  • Treats a mapping missing from the regeneration as a removed endpoint
  • Applies the regenerated set as a patch instead of reviewing it