In a product feature backed by a generative step, what can a test case put in that step's place, and what does a failing case prove?
answer
- Where the case takes control
- Three choices at one boundary
- Stored answer, authored answer, real call
- Determinism decides what a red build accuses
basics
~20 sThree seams exist: replay a response recorded from a real call, return an answer the test author wrote, or make the real call. The first two make a failure point at your own code; the third does not.
solid answer
~50 sA test case for a feature whose output comes from a generative step has three places to cut. **Replay** returns a response captured earlier from a real call. **Script** returns an answer the test author wrote by hand, chosen to drive a particular branch. **Live** lets the case reach the real step. Replay and script both make the case deterministic, so a red result points at code the team owns: the prompt assembly, the parsing, the guards, the rendering. A live case is the only one that exercises the real behaviour, but a red result there could be the team's code, the step's legitimate variation, or the provider refusing the call, and the case alone cannot say which. Most suites use all three - authored answers to drive each branch of the handling, a replayed recording for a real-shaped happy path, and live calls in a small separate set.
code
pseudocode · 10 lines# one product case, wired three ways
case "order summary renders as a list":
input = "summarise order 4471"
seam = REPLAY -> generativeStep = replay(recording "summary-2026-03-11")
seam = SCRIPT -> generativeStep = returns("- first line, - second line")
seam = LIVE -> generativeStep = realCall()
answer = generativeStep.run(input)
assert render(answer).isBulletList()go deeper
Be able to name the three things a test case can put where a generative step would be: a stored real answer, an answer the author wrote, and the real call itself. Know that the first two return the same text every time.
Explain what each seam makes deterministic and what that costs. An interviewer expects you to say which parts of the code - prompt assembly, parsing, guards, rendering - are still under test once the step's answer is held fixed.
Show that you read a red build differently under each seam. Under a fixed answer it accuses your own code; under a real call it could be your code, the step's variation, or the provider, and you say so before you start debugging.
Own the argument that a suite uses all three seams deliberately rather than drifting into one. Be ready to justify the split to a team that wants everything live for realism, or everything stood in for speed.
## The seam and what it separates A **seam** is the place where a test case takes control of something the code would otherwise reach for on its own. In a product feature whose answer text comes from a generative step, that place is the boundary between the code the team wrote and the step itself. Everything on the team's side is testable and ought to be tested: assembling the prompt and the supplied context, sending the request, reading what comes back, parsing it, validating it, guarding it, storing it, rendering it. The step is not the team's code and is not what an ordinary product suite is trying to certify. Choosing a seam is choosing how much of the step to let into the case. ## The three seams - **A replayed recording.** Someone made a real call once and stored what came back; the case returns that stored text. It carries the real shape of a real answer - the trailing sentence nobody planned for, the field order, the length - which is what makes it valuable and also what makes it perishable. - **An authored answer.** The test author writes the text by hand and chooses it to drive one branch: an empty answer, a refusal, an answer twice the expected length, one in the wrong structure, one perfectly ordinary. It never happened. It is a hypothesis about what could happen. - **The real call.** The case reaches the live step and takes whatever comes back today. | Seam | Same result every run | Exercises the real step | Cost per suite run | What a failure implicates | |---|---|---|---|---| | Replayed recording | Yes | As it behaved on the capture date | Negligible | The team's code, or a stale recording | | Authored answer | Yes | No | Negligible | The team's code, for a shape the author chose | | Real call | No | Yes | Time and money, every run | The team's code, the step's variation, or the provider | ## What a failing case proves under each This is the part candidates skip, and it is the whole reason the choice matters. A red build is only useful if it points somewhere. Under a **replayed** or **authored** seam the answer is a constant. Nothing outside the repository can change the result, so a case that turns red turned red because the team's code changed behaviour: the prompt assembly, the parsing, a validation rule, the rendering. That is a clean defect signal and it is safe to block a merge on. Under the **real call**, three quite different things produce the same red: the team's code broke; the step returned a legitimately different answer that the case's checks happened not to accept; or the provider refused, timed out, or was slow. The case cannot distinguish them. A red live case is therefore a prompt to investigate, not a verdict on the change under review, and treating it as one is how teams learn to ignore their own pipeline. The mirror of that is the useful part: a **green** replayed case says nothing about whether the real step still behaves that way, and a **green** live case says only that the code handled one answer produced once. ## Where the stand-in sits Two arrangements are common and they are not interchangeable. In one, the stand-in lives **inside the same process** as the code under test: the collaborator that would call the step is replaced before the code runs, so no request is ever produced. In the other, the stand-in sits **at the network boundary** the process talks over: the code assembles and sends a real request, and something in front of the step answers it. The difference decides what a recording is replayed *against*. An in-process stand-in is usually handed whatever arguments the code computed and returns the recording regardless of them, so a case can pass while the prompt assembly is quietly broken - nothing ever looked at the prompt. A boundary stand-in sees the actual request and can key the recording to it: send a different request and no recording matches, and the case fails. The boundary arrangement puts more of the team's code under test - serialisation, headers, deadlines, the wrapper that retries - for a heavier setup. If the case is about parsing and rendering, the closer seam is enough; if the request itself is the thing at risk, the seam has to sit further out. ## Choose per case, not per suite The mature answer is that all three live in one suite and each case picks the one that matches what it is trying to prove: 1. **Authored** for every branch of the team's handling. This is most of the suite, and it is cheap and stable. 2. **Replayed** for the ordinary path, so at least one case is shaped like a real answer rather than an author's idea of one. 3. **Live** for a small, named set that runs where its ambiguity cannot block unrelated work. A suite wired entirely to one seam is either fast and blind or honest and untrustworthy.
- Why keep answers written by the test author at all, once real recordings are available?A recording shows one shape: whatever the step produced that day. Authored answers are the only seam that can produce the shapes you cannot ask for on demand - an empty answer, a refusal, a truncation, a wrong structure - so they are what drives each branch of the handling. Realism is not the point of those cases; branch reach is.
- Does a live case that passes prove the feature is correct?No. It proves the code handled one answer, produced once, for one input. The next answer to the same input may differ in wording, length or structure, and nothing about the pass says the code copes with that. A live green is evidence the path is connected, not evidence the handling is complete.
- How do you decide the seam per case rather than per suite?Ask what the case is trying to prove. If it is about the team's handling of a particular output shape, author the answer. If it is about a real-shaped answer flowing through end to end, replay a recording. If it is about the step actually being reachable and behaving, make the real call - and accept that the case's failures then need reading rather than blocking.
It is the difference between rehearsing against a recording of your co-star, against a stand-in reading lines you wrote yourself, and against the co-star in person. Only the last shows what they will actually do; only the first two let you rehearse a scene on demand.
saying these in an interview costs you the question
- Says a replayed recording proves the feature still works end to end
- Treats every red live case as a defect in the team's own code
- Picks one seam for the whole suite and never revisits it
- Cannot say what a failure under each seam actually points at
- Dismisses authored answers as unrealistic and therefore worthless