When a test of a feature that answers from a generative step fails and will not reproduce, what must that test run itself have recorded?
answer
- The failure may not happen twice
- Write it while the attempt is alive
- Input, supplied material, instruction version
- Raw answer plus a correlation id
basics
~20 sWrite the evidence at failure time, because a repeat may never reproduce it: the exact input, the material and instruction text given to the generative step, the model identifier and generation settings, the raw output, and a correlation id.
solid answer
~50 sA feature that answers from a generative step need not fail the same way twice, so the failing attempt is the only chance to collect evidence. At the moment the assertion fails the run should write: the exact input the test supplied; the material and instruction text handed to the generative step, with the version of that instruction text; the model identifier and the generation settings; the raw output that came back; the derived value the assertion actually compared; the assertion with its expected and actual values; and a correlation id shared with the product's own log lines so the two can be lined up afterwards. Write it unconditionally on failure, one record per attempt rather than per test case, to a durable place that outlives the build machine, with personal data removed on the way in.
code
pseudocode · 15 lineson assertion_failed(testCase, attempt):
write_record(durableStore, {
correlationId: attempt.correlationId,
testCase: testCase.name,
attemptNumber: attempt.index,
input: redact(testCase.input),
suppliedMaterial: redact(attempt.suppliedMaterial),
instructionTextId: attempt.instruction.id,
instructionVersion: attempt.instruction.version,
modelIdentifier: attempt.modelIdentifier,
generationSettings: attempt.generationSettings,
rawOutput: redact(attempt.rawOutput),
derivedValue: attempt.derivedValue,
assertion: { name: ..., expected: ..., actual: ... }
})go deeper
Be ready to say why you cannot simply run the failing test again when the feature's answer comes from a generative step, and to name the input and the returned text as things the run itself has to save.
Explain the whole record and why each part is there: input, supplied material, instruction text and version, model identifier, generation settings, raw output, derived value, the failed assertion, and a correlation id written at failure time.
Show you have designed one: writing unconditionally on failure, sampling passes for comparison, one record per attempt, a durable location off the build machine, and structured output so two attempts can be compared rather than re-read.
Own the trade-off between what every failing attempt costs to store and how much of it is ever read, and set the capture standard once, centrally, so every team shipping a feature backed by a generative step gets it without deciding.
## Why the failing attempt is the only evidence you get A product feature that answers from a **generative step** is not guaranteed to behave the same way twice. The same user input can produce different wording, a different structure, or a refusal, and the difference can come from the step's own sampling, from the material the product assembled and handed over on that attempt, from an edit to the instruction text, or from a change to which model the product calls. That makes a failing test over such a feature categorically different from a failing test over deterministic code. Over deterministic code, a failure is a fact you can walk back to whenever you like: the test name and a stack trace are enough, because the machine will do it again on demand. Over a generative feature, the attempt **is** the fact. Once it is over, the in-memory objects are gone, the build machine may have been destroyed, and a repeat is a *new* attempt that produces its own evidence and settles nothing about the old one. The rule follows: **write the evidence while the attempt is alive, not when someone gets round to investigating.** ## What the failing run must write Each item below earns its place by answering a question an investigator would otherwise have to guess at. | Recorded at failure | The question it answers | |---|---| | The exact input the test supplied | Was the failing case the one you think it was? | | The material supplied to the generative step | Did the product assemble the right supporting content? | | The instruction text, and its version | Which wording produced this, and is that wording still in effect? | | The model identifier | Was this the model the release was built against? | | The generation settings | Would a later attempt even be the same call? | | The raw output, before any handling | Was the returned text wrong, or the product's handling of it? | | The derived value the assertion saw | What exactly was compared? | | The assertion, with expected and actual | Which property failed, and by how far? | | The attempt number within the test case | Was this the only failing attempt of several? | | A correlation id | Can this be lined up with what the product itself logged? | The **correlation id** is the cheapest item on that list and often the most valuable. One identifier, generated by the test and carried on the call, lets an investigator put the test's record beside whatever the product logged for the same attempt. Without it, both sets of evidence exist and nobody can prove they describe the same moment. ## Writing it so it survives A few habits separate a capture that works from one that only looks like it works. 1. **Unconditional on failure.** No flag to switch on, no advice to re-run with more output enabled. The whole premise is that the second attempt may not fail. 2. **Sampled on success.** A passing attempt from the same build is the comparison an investigator reaches for first: same input, same test case, different outcome, so which recorded fact differed? Keeping every passing attempt is usually too expensive; keeping the most recent one per test case rarely is. 3. **One record per attempt.** When a test case calls the generative step several times, folding those into a single record destroys the thing you need — which attempt failed, and how it differed from the ones that did not. 4. **Durable and off the box.** A build machine's local disk is temporary, and console output is usually truncated and rotated away exactly when you want to read it. 5. **Structured rather than prose.** A machine-readable record lets a reader compare two attempts field by field; a paragraph of formatted text makes them read both in full. 6. **Redacted on the way in.** Personal data that is never written cannot escape from the store afterwards. ## What this record is not Three neighbouring things are easy to confuse with it, and the confusion produces either duplicated effort or a gap. - It is **not the defect write-up**. A person composes that afterwards, out of this material; the record is raw evidence, not a narrative aimed at a reader. - It is **not a measurement of the generative model's quality**. Nothing here scores an answer or grades the feature against a standard; it explains one failure. - It is **not production monitoring**. This is the suite's own evidence about its own attempt, produced under a test's control on a build machine. The discipline is small and the payoff is out of proportion to it. Teams that skip it accumulate a class of failure that gets discussed rather than investigated: it failed on the build machine, nobody could make it happen again, and after a week the failure is forgotten rather than explained. Capturing the attempt is what turns that into an ordinary defect with an ordinary explanation.
- Should a passing attempt be recorded too?Sampled, yes. A passing attempt from the same build is the comparison an investigator reaches for first: same input, same test case, different outcome, so which recorded fact differed? Keeping every passing attempt is usually too expensive; keeping the most recent passing record per test case rarely is, and it turns a lone failure record into a difference.
- Where should the record be written so it is still readable a week later?Somewhere that outlives the machine that ran the test, keyed by the correlation id. Not the build machine's local disk, and not only console output, which is normally truncated and rotated away exactly when someone wants it. Machine-readable structure matters more than presentation, because the common operation is comparing two attempts field by field.
- A test case calls the generative step five times. How should that be recorded?One record per attempt, each carrying its own attempt number under the test case's correlation id. Folding five attempts into one record loses precisely what an investigation needs: which attempt failed, and how its recorded facts differed from the four that did not.
Treat the failing attempt like a flight you cannot fly again: whatever the recorder did not write down is not recoverable afterwards, however carefully you ask.
saying these in an interview costs you the question
- Says re-running the failing test is enough to investigate it
- Records only the assertion message and the test case name
- Plans to add capture later, once the failure is understood
- Keeps evidence only on the build machine's local disk
- Writes one record per test case even when it makes several attempts