skip to content

Suite Mechanics

Living with a step that legitimately varies inside an ordinary product suite: standing the model in or calling it, repeats and pass rules, and what a failure records. Such suites lose trust fast.

on this pageshow

questions

12

When a test of a feature that answers from a generative step fails and will not reproduce, what must that test run itself have recorded?

level: middleimportance: must knowfreq 62%

answer

  1. The failure may not happen twice
  2. Write it while the attempt is alive
  3. Input, supplied material, instruction version
  4. Raw answer plus a correlation id

basics

~20 s

Write the evidence at failure time, because a repeat may never reproduce it: the exact input, the material and instruction text given to the generative step, the model identifier and generation settings, the raw output, and a correlation id.

solid answer

~50 s

A feature that answers from a generative step need not fail the same way twice, so the failing attempt is the only chance to collect evidence. At the moment the assertion fails the run should write: the exact input the test supplied; the material and instruction text handed to the generative step, with the version of that instruction text; the model identifier and the generation settings; the raw output that came back; the derived value the assertion actually compared; the assertion with its expected and actual values; and a correlation id shared with the product's own log lines so the two can be lined up afterwards. Write it unconditionally on failure, one record per attempt rather than per test case, to a durable place that outlives the build machine, with personal data removed on the way in.

code

pseudocode · 15 lines
pseudocode
on assertion_failed(testCase, attempt):
    write_record(durableStore, {
        correlationId:      attempt.correlationId,
        testCase:           testCase.name,
        attemptNumber:      attempt.index,
        input:              redact(testCase.input),
        suppliedMaterial:   redact(attempt.suppliedMaterial),
        instructionTextId:  attempt.instruction.id,
        instructionVersion: attempt.instruction.version,
        modelIdentifier:    attempt.modelIdentifier,
        generationSettings: attempt.generationSettings,
        rawOutput:          redact(attempt.rawOutput),
        derivedValue:       attempt.derivedValue,
        assertion: { name: ..., expected: ..., actual: ... }
    })

go deeper

for a junior

Be ready to say why you cannot simply run the failing test again when the feature's answer comes from a generative step, and to name the input and the returned text as things the run itself has to save.

for a middle

Explain the whole record and why each part is there: input, supplied material, instruction text and version, model identifier, generation settings, raw output, derived value, the failed assertion, and a correlation id written at failure time.

for a senior

Show you have designed one: writing unconditionally on failure, sampling passes for comparison, one record per attempt, a durable location off the build machine, and structured output so two attempts can be compared rather than re-read.

for a principal

Own the trade-off between what every failing attempt costs to store and how much of it is ever read, and set the capture standard once, centrally, so every team shipping a feature backed by a generative step gets it without deciding.

## Why the failing attempt is the only evidence you get A product feature that answers from a **generative step** is not guaranteed to behave the same way twice. The same user input can produce different wording, a different structure, or a refusal, and the difference can come from the step's own sampling, from the material the product assembled and handed over on that attempt, from an edit to the instruction text, or from a change to which model the product calls. That makes a failing test over such a feature categorically different from a failing test over deterministic code. Over deterministic code, a failure is a fact you can walk back to whenever you like: the test name and a stack trace are enough, because the machine will do it again on demand. Over a generative feature, the attempt **is** the fact. Once it is over, the in-memory objects are gone, the build machine may have been destroyed, and a repeat is a *new* attempt that produces its own evidence and settles nothing about the old one. The rule follows: **write the evidence while the attempt is alive, not when someone gets round to investigating.** ## What the failing run must write Each item below earns its place by answering a question an investigator would otherwise have to guess at. | Recorded at failure | The question it answers | |---|---| | The exact input the test supplied | Was the failing case the one you think it was? | | The material supplied to the generative step | Did the product assemble the right supporting content? | | The instruction text, and its version | Which wording produced this, and is that wording still in effect? | | The model identifier | Was this the model the release was built against? | | The generation settings | Would a later attempt even be the same call? | | The raw output, before any handling | Was the returned text wrong, or the product's handling of it? | | The derived value the assertion saw | What exactly was compared? | | The assertion, with expected and actual | Which property failed, and by how far? | | The attempt number within the test case | Was this the only failing attempt of several? | | A correlation id | Can this be lined up with what the product itself logged? | The **correlation id** is the cheapest item on that list and often the most valuable. One identifier, generated by the test and carried on the call, lets an investigator put the test's record beside whatever the product logged for the same attempt. Without it, both sets of evidence exist and nobody can prove they describe the same moment. ## Writing it so it survives A few habits separate a capture that works from one that only looks like it works. 1. **Unconditional on failure.** No flag to switch on, no advice to re-run with more output enabled. The whole premise is that the second attempt may not fail. 2. **Sampled on success.** A passing attempt from the same build is the comparison an investigator reaches for first: same input, same test case, different outcome, so which recorded fact differed? Keeping every passing attempt is usually too expensive; keeping the most recent one per test case rarely is. 3. **One record per attempt.** When a test case calls the generative step several times, folding those into a single record destroys the thing you need — which attempt failed, and how it differed from the ones that did not. 4. **Durable and off the box.** A build machine's local disk is temporary, and console output is usually truncated and rotated away exactly when you want to read it. 5. **Structured rather than prose.** A machine-readable record lets a reader compare two attempts field by field; a paragraph of formatted text makes them read both in full. 6. **Redacted on the way in.** Personal data that is never written cannot escape from the store afterwards. ## What this record is not Three neighbouring things are easy to confuse with it, and the confusion produces either duplicated effort or a gap. - It is **not the defect write-up**. A person composes that afterwards, out of this material; the record is raw evidence, not a narrative aimed at a reader. - It is **not a measurement of the generative model's quality**. Nothing here scores an answer or grades the feature against a standard; it explains one failure. - It is **not production monitoring**. This is the suite's own evidence about its own attempt, produced under a test's control on a build machine. The discipline is small and the payoff is out of proportion to it. Teams that skip it accumulate a class of failure that gets discussed rather than investigated: it failed on the build machine, nobody could make it happen again, and after a week the failure is forgotten rather than explained. Capturing the attempt is what turns that into an ordinary defect with an ordinary explanation.

  • Should a passing attempt be recorded too?
    Sampled, yes. A passing attempt from the same build is the comparison an investigator reaches for first: same input, same test case, different outcome, so which recorded fact differed? Keeping every passing attempt is usually too expensive; keeping the most recent passing record per test case rarely is, and it turns a lone failure record into a difference.
  • Where should the record be written so it is still readable a week later?
    Somewhere that outlives the machine that ran the test, keyed by the correlation id. Not the build machine's local disk, and not only console output, which is normally truncated and rotated away exactly when someone wants it. Machine-readable structure matters more than presentation, because the common operation is comparing two attempts field by field.
  • A test case calls the generative step five times. How should that be recorded?
    One record per attempt, each carrying its own attempt number under the test case's correlation id. Folding five attempts into one record loses precisely what an investigation needs: which attempt failed, and how its recorded facts differed from the four that did not.

Treat the failing attempt like a flight you cannot fly again: whatever the recorder did not write down is not recoverable afterwards, however carefully you ask.

saying these in an interview costs you the question

  • Says re-running the failing test is enough to investigate it
  • Records only the assertion message and the test case name
  • Plans to add capture later, once the failure is understood
  • Keeps evidence only on the build machine's local disk
  • Writes one record per test case even when it makes several attempts
open as a page

In a product feature backed by a generative step, what can a test case put in that step's place, and what does a failing case prove?

level: middleimportance: must knowfreq 58%

basics

~20 s

Three seams exist: replay a response recorded from a real call, return an answer the test author wrote, or make the real call. The first two make a failure point at your own code; the third does not.

open as a page

How many repeats should a test case run against a generative product feature, and what does a four-in-five pass rule claim?

level: middleimportance: must knowfreq 60%

basics

~20 s

Repeat count follows the success-rate difference you need to detect; three to five repeats separate only coarse differences. A four-in-five rule claims the feature's per-attempt success rate is high, not that any single user attempt will succeed.

open as a page

Why must a failing test keep the generative step's raw answer, not just the value the product parsed from it?

level: juniorimportance: should knowfreq 44%

basics

~20 s

Only the raw answer separates a bad answer from good text the product mishandled. With just the derived value, a deterministic parsing, trimming or fallback bug in the team's own code reads as variation in the generative step.

open as a page

Why is re-running a test case until the build turns green not a repair when the product's answer comes from a generative step?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Re-running draws another sample from the same varying feature; it changes the reported result, not the feature. Under a pass rule that already tolerates the expected variation, a repeated red is evidence the success rate is genuinely below the bar.

open as a page

Failing tests of a generative feature save real customer text as evidence. How should that evidence be redacted, kept and read?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Apply per-field redaction in the writer, so the unredacted form never reaches the store; keep enough shape that the record still explains the failure; set and enforce a retention window; and default read access to the people investigating that feature.

open as a page

A failing test saved the generative step's input and output but nothing about the call itself. What can an investigation not conclude?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Without the supplied material, the instruction version, the model identifier and the generation settings, an investigation cannot say which wording or model produced the failure, whether a repeat is the same call, or whether a later edit fixed anything.

open as a page

A test suite replays one recorded response from its product's generative step: how do you keep that recording honest?

level: seniorimportance: should knowfreq 37%

basics

~20 s

Stamp the recording with its capture date and the configuration it was captured under, and refresh whenever any of that changes. Comparing it word for word against a fresh real answer proves nothing, because the real step never repeats itself.

open as a page

For a generative step in a product, what defect does each test seam - replay, script, or live call - let through?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Replay hides any change in what the real step now returns, including a shape the parser cannot handle. An authored answer hides everything the real step actually does. A live call hides shapes it happens not to produce that day.

open as a page

In a generative product feature, which evidence shows a varying test result comes from the feature rather than the test?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Pin the generative step to one fixed recorded answer and repeat: if the result still moves, the test is the source. Then read which assertion moved - generated wording varies by design, a stored identifier does not.

open as a page

Calling the real generative step is slow, costly and sometimes fails for reasons outside your team - what seam policy do you set?

level: principalimportance: should knowfreq 40%

basics

~20 s

Keep the merge gate deterministic and put the rich live cases on a schedule and before a release, with a named owner. Add one minimal real call to the gate so a total outage of the step cannot pass unnoticed.

open as a page

A generative product feature's test case met its criteria in four of five repeats - what does a single pass mark hide?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A single pass mark collapses a rate into a certainty. It hides which repeat missed and on which criterion, how badly it missed, whether four-in-five is normal for this case, and whether the misses cluster on one input.

open as a page