skip to content

A test suite replays one recorded response from its product's generative step: how do you keep that recording honest?

level: seniorimportance: should knowfreq 37%

answer

  1. Nothing fails when a fixture rots
  2. A text difference proves nothing here
  3. Stamp it with its configuration
  4. Judge a scheduled real call by properties

basics

~20 s

Stamp the recording with its capture date and the configuration it was captured under, and refresh whenever any of that changes. Comparing it word for word against a fresh real answer proves nothing, because the real step never repeats itself.

solid answer

~50 s

Do not try to detect staleness by comparing the recording with a fresh real answer: the step words its answer differently every time, so the comparison is red on every run and proves nothing. Instead store the recording with the configuration it was captured under - prompt version, context assembly, the identifier and settings of the step called, and a capture date - and let any change to those fail a cheap check that forces a re-capture. For the case where nothing on your side changed and the step moved underneath, make a scheduled real call and judge it against the same properties the recording was chosen for: it carries a total line, it parses, it sits under the length limit. When a real answer stops satisfying those, the recording is no longer representative. Add a maximum age and a named owner for everything else.

code

yaml · 13 lines
yaml
recording: order-summary-happy-path
capturedAt: 2026-03-11
capturedUnder:
  promptVersion: 7
  contextAssembly: order-lines-v3
  stepIdentifier: pinned-in-config
  generationSettings: { variability: low, maxLength: 400 }
satisfiedProperties: [ hasTotalLine, underLengthLimit, parsesCleanly ]
refreshWhen:
  - anyValueUnder capturedUnder changes
  - capturedAt older than 90 days
  - scheduled real call stops satisfying satisfiedProperties
owner: orders-team

go deeper

for a junior

Know that a stored answer from a generative step goes out of date, and that nothing in the suite fails when it does. A replayed case stays green whatever the real step is doing now.

for a middle

Explain why comparing the recording against a fresh real answer is not a staleness test: the real step words its answer differently every time, so the comparison would be red on every run and would mean nothing.

for a senior

Show the mechanisms you would actually run - a refresh triggered by a change in the recorded configuration, a scheduled real call judged against the properties the recording was chosen for, and an age limit reported to an owner.

for a principal

Own the question of how much fixture upkeep a product is worth, and how refresh work stays funded when nothing is failing, nobody is asking for it, and the suite looks healthy.

## Why a recording rots quietly A recording is a real answer, captured once, then frozen. Everything that produced it keeps moving: the wording of the prompt, the way the supplied context is assembled, the identifier and version of the step the product calls, its generation settings, and the step's own behaviour, which changes without asking the team. None of that movement makes a replayed case fail. That is the whole problem. A replayed case is green by construction, so the day its recording stops resembling reality it is exactly as green as the day it was captured, and the suite has quietly become a suite about a product that no longer exists. Teams discover this the same way every time: an incident, a look at the input that broke, and the realisation that the fixture the suite has been defending is two years old. ## Why a difference is not a signal The obvious idea - call the real step, compare its answer to the recording, refresh when they differ - does not work here, and knowing why is what separates a real answer from a plausible one. The real step's output legitimately varies. Ask it the same thing twice and the wording differs; ask it tomorrow and it differs again. A text comparison is therefore red on every single run, and a check that is always red is not a check. Loosening it until it passes - normalise the whitespace, compare lengths, score similarity against some cut-off - trades one problem for a worse one: an arbitrary number nobody can defend, which fails for reasons unrelated to the product. The honest reframe is this: **a recording is not there to be a copy of today's answer. It is there to be an answer with the properties the code has to handle.** Staleness is therefore not "the text differs". Staleness is "something that shaped this recording has changed", or "an answer produced today no longer has the properties this recording was chosen for". ## Refresh triggers you can check Store the recording together with the configuration it was captured under, and the triggers become mechanical rather than a matter of someone remembering: | What changed | Why the recording is suspect | |---|---| | The prompt text or its version | The answer was produced by different instructions | | The context assembly | The step was given different material to work from | | The identifier or version of the step called | A different producer, with different habits | | The generation settings | Different length, structure and variability | | Nothing, but months have passed | The step's behaviour moves on its own | A recording whose stored configuration no longer matches the product's current configuration should fail a cheap check in the pipeline - not because the recording is definitely wrong, but because nobody has looked. ## Noticing staleness anyway Configuration triggers miss the case that matters most: nothing on the team's side changed and the step changed underneath. Two mechanisms cover it, and neither is a text comparison. 1. **A scheduled real call judged by the same properties the recording satisfies.** If the recording was chosen because it carries a total line, parses cleanly and sits under the length limit, then make a real call on a schedule and assert exactly those properties of the answer that comes back. When a real answer stops satisfying them, the recording is no longer representative and the parsing or guarding built around it is at risk. That is a signal about the product's assumptions, not a difference between two pieces of text. 2. **An age stamp with a stated maximum.** Every recording carries its capture date, and anything past the agreed age is reported - to an owner, on a schedule, not as a merge block. Age is a weak signal, but it is the only one that fires when the team has stopped paying attention, which is exactly when it is needed. ## What an honest recording carries - The **capture date** and the configuration it was captured under. - The **properties it was chosen for**, written down as the checks the case asserts, so a refresh can be judged rather than eyeballed. - Enough about **how it was captured** that someone can re-capture it deliberately rather than approximately. - An **owner**. A fixture with no owner is refreshed after it has already mattered. Refreshing is then a small, boring job: re-run the capture, confirm the new answer satisfies the recorded properties, commit the recording and its stamp together. The failure mode is never that the refresh is hard. It is that nobody is ever told it is due.

  • What has to be stored alongside a recording for a refresh trigger to be checkable at all?
    The capture date and the configuration the capture ran under - prompt version, how the supplied context was assembled, which step was called and with which generation settings - plus the properties the recording was chosen for. Without that stamp, a trigger is a person remembering, and people stop remembering as soon as the suite is green.
  • Why not simply re-capture every recording on every pipeline run?
    Then it is not a recording, it is a live call with extra steps. You pay the cost and inherit the ambiguity of every real call, and the suite loses the determinism that made a red result accuse your own code. Re-capture is a deliberate act with a review, not something a pipeline does silently.

A photograph of a street is accurate on the day it is taken and never stops being a valid photograph. It stops describing the street long before anyone notices, because nothing about the photograph changes when the street does.

saying these in an interview costs you the question

  • Proposes comparing the recording with a fresh real answer word for word
  • Refreshes recordings only after an incident forces someone to look
  • Assumes a green replayed case means the recording is still representative
  • Stores the recording with no capture date and no configuration
  • Wants every recording re-captured silently on every pipeline run