Why must a failing test keep the generative step's raw answer, not just the value the product parsed from it?
answer
- Two different things can be wrong
- Bad text, or good text mishandled
- Capture at the boundary, before handling
- Store the pair and the transformation between
basics
~20 sOnly the raw answer separates a bad answer from good text the product mishandled. With just the derived value, a deterministic parsing, trimming or fallback bug in the team's own code reads as variation in the generative step.
solid answer
~50 sA product rarely shows a generative step's output untouched: it parses a structure out of it, trims and normalises, checks a shape, drops unrecognised fields, or substitutes a fallback when the check fails. The assertion then runs on that derived value. So two very different things can fail. The step returned genuinely wrong text, or it returned acceptable text that the product's own handling lost or mangled. The second is an ordinary deterministic bug, fixable today. A record holding only the derived value cannot tell them apart, and the cheaper story wins: an empty result looks like the step returning nothing when a shape check rejected a good answer. Keep the pair — the returned text captured at the boundary before any handling, and the derived value the assertion saw — plus which transformations ran and whether a fallback fired.
code
pseudocode · 14 linesanswer = generative_step(suppliedMaterial, instruction)
record.rawOutput = answer.text // captured at the boundary
record.rawLength = length(answer.text)
parsed = parse(answer.text) // may throw, or quietly mangle
if parsed is invalid:
parsed = FALLBACK
record.fallbackUsed = true
record.transformations = ["parse", "trim", "shape check"]
record.derivedValue = parsed
assert_on(parsed)go deeper
Be ready to say that the value your assertion compares is not what came back, and that keeping both is what lets anyone tell a bad answer from good text your own code mishandled.
Explain where the capture has to sit — at the boundary the answer arrives on, before parsing — and what else belongs beside it: the derived value, the transformations that ran, and whether a fallback was substituted.
Show the misattribution this prevents in practice: deterministic handling bugs closed as variation, and real failures of the step waved away as formatting, plus how you handle answers too large to keep in full.
Push for the capture to live in the shared boundary every feature calls through, so no team can ship a generative feature whose failures are only ever describable in terms of what its own code made of the answer.
## Two different things can be wrong A feature that answers from a generative step almost never ships the step's output straight to the user. Something in the product takes the returned text and turns it into a value: it parses a structure out of it, trims and normalises it, checks it against a shape, drops fields it does not recognise, applies a fallback when the check fails, or rewrites it into the product's own format. The assertion in the test then runs against that derived value. So there are two distinct places a failure can come from. - The generative step returned text that is genuinely wrong, or in a shape the product never agreed to accept. - The generative step returned perfectly acceptable text, and the product's own handling of it lost, mangled or rejected it. The second is ordinary deterministic code with an ordinary deterministic bug. It is fixable today, testable forever, and has nothing to do with variation. The first is a different kind of problem entirely. **A record that keeps only the derived value cannot tell them apart** — and when it cannot, the cheaper story wins. ## What the derived value hides | Symptom the assertion reports | With the raw answer stored | With only the derived value | |---|---|---| | Empty result | The step returned a well-formed answer that the shape check rejected | Looks like the step returned nothing | | A field is missing | The field was present under a name the parser does not recognise | Looks like the step omitted it | | Text is truncated | The answer was complete and the trimming rule cut it | Looks like the answer stopped early | | A default value appears | The fallback ran, and you can see the answer that triggered it | Looks like the step produced the default | Every row on the right is the same misreading: a deterministic bug in code the team owns, recorded as variation in a step the team does not control. It is a comfortable misreading, because it moves the problem somewhere unfixable, and it is why suites over generative features accumulate failures nobody ever gets to the bottom of. The same table also runs the other way — a genuine failure of the step gets closed as a formatting nuisance, and the feature ships with a real hole in it. ## Store the pair, and the transformation between them The rule is small: **write the raw answer before the product touches it, then write the derived value the assertion actually saw, and enough about what happened in between to explain the difference.** 1. Capture the returned text at the boundary where it arrives, not after the handling code. If the capture happens inside the parser, a parser that throws takes the evidence with it. 2. Capture the derived value at the point the assertion reads it, not where the handling code believes it produced it. 3. Record which transformations ran and in what order, and record explicitly when a fallback or a repair path was taken. An assertion that compares against a substituted default looks, in a record without this, exactly like an assertion against the step's own work. 4. Where the raw answer is too large to keep for every attempt, truncate deliberately, mark the truncation in the record, and store the original length. Silently keeping the first part is how a team ends up certain the answer stopped early. ## The reflex this defends against "The model is non-deterministic" is true, available, and a complete conversation-stopper. It explains any failure, which is exactly what makes it worthless as an explanation. The raw answer is the single artefact that makes the claim checkable: with it in the record, anyone can read what came back and see for themselves whether the text was defensible. Without it, the claim cannot be argued with and so it is never argued with. There is a second, quieter benefit. Once raw answers are stored for failures, they accumulate into a picture of what the step actually tends to return — which shapes recur, which phrasings the handling code was never written for, which of the product's assumptions about the format were optimistic. That picture is a by-product of ordinary failure evidence rather than a separate exercise, and it usually produces more fixes in the product's own handling code than in anything else. ## The one-line version Keep the input, keep what came back, and keep what you made of it. A record with the middle term missing can describe that something went wrong, but it cannot say which half of the feature did it.
- The raw answer is too large to keep for every failing attempt. What do you do?Truncate deliberately rather than by accident: mark the truncation in the record and store the original length, so nobody reads a cut record as an answer that stopped early. Keep failures in full where you can afford it and sample passes; the failing attempts are the ones that have to be explainable.
- The product substituted a fallback value when its shape check rejected the answer. What must the record show?That the fallback ran, what the answer was that triggered it, and which value was substituted. Without that flag the assertion's actual value looks like the generative step's own work, and an investigation spends its time asking why the step produced a default it never produced.
Keeping only the derived value is like keeping the cropped print and throwing away the negative: when the picture is wrong you can no longer tell whether the subject or the crop was at fault.
saying these in an interview costs you the question
- Stores only the parsed value because the returned text is noisy
- Blames variation for every failure of a generative feature
- Logs the answer after trimming and normalising it
- Captures inside the parser, so a throwing parser takes the evidence
- Records a substituted fallback value as though the step produced it