For a generative step in a product, what defect does each test seam - replay, script, or live call - let through?
answer
- Every seam buys something with blindness
- Name the defect, not the caveat
- A frozen answer cannot notice change
- Live only sees what happened today
basics
~20 sReplay hides any change in what the real step now returns, including a shape the parser cannot handle. An authored answer hides everything the real step actually does. A live call hides shapes it happens not to produce that day.
solid answer
~50 sName the defect, not the caveat. A **replayed recording** is frozen, so the suite cannot notice the real step's output shape moving: the day it starts wrapping its answer in an extra sentence, the extraction that worked on the recording breaks in production while every case stays green. An **authored answer** is bounded by the author's imagination, so any real shape nobody thought of - a different structure, an unexpected length, characters the renderer treats specially - is untested. A **live call** only shows what the step happened to produce that day, so rare or unprovokable outputs such as an empty answer, a refusal or a truncation never run the branches that handle them, and a green result is not evidence those branches work. The repair is to arrange the seams so their blind spots do not overlap.
code
pseudocode · 10 lines# captured 2026-01-04, replayed on every run since
recorded = "TOTAL: 42 units"
parseTotal(text) = numberAfter(text, "TOTAL:")
assert parseTotal(recorded) == 42 # green forever
# what the real generative step returns today
today = "Here is the summary. The total is 42 units."
parseTotal(today) # no marker -> production defect
# no case in the suite ever sees thisgo deeper
Know that every stand-in for a generative step buys stability by hiding something. Be able to say that a stored answer cannot notice the real step changing, because it is the same text on every run.
Name a concrete defect per seam rather than a caveat: a parser that breaks when the real output gains a sentence, a shape the author never imagined, a rare answer the live call never happened to produce.
Demonstrate that you arrange a suite so the seams' blind spots do not overlap, and that for any case put in front of you, you can state what it would fail to catch and which other case covers it.
Be ready to argue about the residual risk the whole arrangement leaves, and which of those blind spots the product can accept given what a wrong or unhandled answer costs it.
## The right question to ask of a seam Every seam buys stability with blindness. The useful way to compare the three is not "which is most realistic" but "name the specific product defect that ships because this case is wired this way". A general caveat - *recordings can get out of date* - is not an answer. The defect is. ## What a replayed recording stops catching A recording is a constant. The real generative step is not. - **The output shape moved and the parser no longer copes.** The recording says `TOTAL: 42 units`; the real step now says `Here is the summary. The total is 42 units.` The extraction that keys off a fixed marker breaks in production and never in the suite. - **The step began refusing a category of input.** Real users meet an apology where the product expects an answer; every replayed case still gets the answer. - **Length.** The recording is three hundred characters; the real step now returns three thousand, and the storage column, the rendering box or a downstream limit cannot take it. - **Anything about the request.** Depending on where the stand-in sits, the case may never inspect what the code actually sent, so an omitted piece of the supplied context or a wrong instruction can be broken for months behind a green case. ## What an authored answer stops catching An authored answer is bounded by the author's imagination, so it stops the suite catching **everything the author did not think of**. In practice that is a short list of recurring surprises: - answers that are correct but arrive in a different structure than the one the author assumed; - text that is well formed but carries characters the renderer or the storage layer treats specially; - an answer that is plausible, well shaped and wrong, which is a situation no invented fixture ever raises; - the real distribution of lengths and shapes, which nobody invents accurately. It also stops the suite catching **the step changing at all**, because there is nothing there that could change. ## What the real call stops catching This is the one candidates miss, because live feels like the honest option. A live case sees exactly one answer: whatever the step produced for that input at that moment. It therefore stops the suite catching: - **the rare shape.** If the step returns an empty answer once in five hundred calls, no live case has ever seen one, and the branch that handles it has never run. - **the shape you cannot provoke.** A refusal, a truncation, a malformed response, an answer in the wrong language - you cannot ask the real step for these on demand, so live cases never exercise them. - **regressions in the team's own handling**, because a red live case is ambiguous. Teams that lean on live cases loosen the checks until they stop failing, and a loose check catches nothing. | Seam | Specific defect that ships unnoticed | |---|---| | Replayed recording | Parsing, guards or limits break when the real output shape moves | | Authored answer | Any real shape the author never imagined | | Real call | Rare or unprovokable shapes, and regressions hidden by loosened checks | ## Covering the gaps on purpose Because each seam's blind spot is different, the repair is to arrange the seams so that no blind spot is shared: 1. Write authored answers for the shapes you cannot provoke - empty, refusal, over-long, malformed, wrong structure - because that is the only seam capable of producing them. 2. Keep at least one replayed recording per feature so the shape of a real answer is under test rather than an invented one. 3. Run a small live set somewhere its ambiguity is affordable, and read its failures rather than gating on them. 4. Say out loud, in review, what each new case's seam is blind to. If the answer is "nothing", the case is wired to a seam nobody has thought about. The test of whether a team understands its own suite is whether it can answer, for any green run, the question "what would still be broken?". With three seams and three different blind spots, that answer is available; with one seam used everywhere, it is not.
- After an outage a team adds twenty more recorded cases. Why might that not reduce the risk?Every one of those recordings was captured at the same moment from the same version of the step, so they share a single blind spot: none of them can notice the step changing afterwards. More recordings widen the range of shapes under test only if they came from genuinely different inputs, and even then they age together.
- Which output shapes should an authored answer always cover for a feature backed by a generative step?The ones you cannot provoke on demand: an empty answer, a refusal, an answer truncated part way, one far longer than expected, and one in a structure the parser does not expect. Those are precisely the shapes a recording and a live call are both unlikely to hand you, so authoring them is the only way the handling gets exercised at all.
saying these in an interview costs you the question
- Answers with a caveat instead of naming a defect that would ship
- Thinks calling the step live removes the need for authored answers
- Believes more recordings automatically widen what the suite catches
- Assumes a green live result means rare output shapes are handled
- Says a stored answer is realistic, so nothing important is missed