skip to content

In a product feature whose answer text comes from a generative step, which parts can a test still assert exactly?

level: middleimportance: must knowfreq 62%

answer

  1. Most of the feature is ordinary software
  2. Draw the boundary before writing cases
  3. Dispatch, fields, citations, states, permissions
  4. Only the wording needs tolerant checks

basics

~20 s

Everything except the generated wording stays exactly assertable: the dispatch path and its inputs, the response's non-text fields, citation rendering, the empty and cut-short states, and permission refusals. Map that surface first, then write looser checks over the text.

solid answer

~50 s

A feature built around a generative step is mostly ordinary software, so draw the boundary before writing anything. On one side sits the wording, which will not repeat itself. On the other sits everything with a single right answer: **which internal capability or retrieval path the request was dispatched to and with which inputs**, the non-text fields and shape of the response, whether a shown citation resolves to the document it names, the empty state when nothing matched, the marker when a length ceiling cut the answer short, and the refusal a caller without permission receives. Those get exact expected values and plain equality assertions. Only the remaining wording needs the looser, property-shaped work. Teams that skip the map end up with one broad end-to-end case asserting a whole recorded answer string: it reddens on every harmless rephrasing and stays green while the dispatch silently breaks.

code

pseudocode · 13 lines
pseudocode
surface_map(feature = "answer_with_sources"):

  exact:
    dispatch.capability      == "search_documents"
    dispatch.args.query      == request.query
    response.sourceIds       == ["doc-14", "doc-22"]
    response.state           == "complete"
    citation("doc-14").target resolves
    render_for(guest_caller)  == "refused"
    stored.conversationTurns == 1

  variable:
    response.answerText      -> tolerant property checks only

go deeper

for a junior

Be ready to name parts of such a feature that behave like any other software: the refusal an unauthorised caller gets, the empty state, the fields beside the answer text. Knowing that only the wording varies is the point.

for a middle

Explain the mechanics of drawing the boundary: walk the feature hop by hop, mark each as fixed or variable, and name the seam where a test observes each fixed hop. Expect to justify why equality is right for those hops.

for a senior

Show the production consequence: a suite that asserts whole answer strings is muted within a month and stops catching dispatch and rendering regressions. Demonstrate how the map keeps failures localised and the tolerant checks small.

for a principal

Own the split as policy. Decide where the boundary sits for a whole product area, what the team is allowed to assert loosely, and how much investment the tolerant side justifies against the exact side that catches most real defects.

## The feature is not the generative step A product feature built around a generative step is ordinary software with one component whose output is not fixed. A request arrives, the caller is authorised, inputs are validated, a decision is made about what to fetch or which action to invoke, a call is made, a result comes back, it is shaped into fields, rendered, persisted, logged and accounted for. Exactly one hop in that chain refuses to repeat itself. Every other hop has a single right answer and can be asserted with plain equality, exactly as it would be if no generative step existed at all. **The deterministic surface** is the name for that set of hops. Mapping it is a deliberate act performed before any case about the wording is written, and it is usually the difference between a suite that localises a defect and a suite that only ever says "something is different". ## What sits on the exact side - **Dispatch and routing** — which internal capability, retrieval path or downstream action the request was sent to, and with which input values. Observed at the hand-off, compared against one expected value. - **The response's non-text fields** — identifiers, counts, status values, the list of source references attached to the answer, the overall shape of the payload the caller receives. - **Citation rendering** — that a reference the interface displays resolves to the document it names, is reachable, and is not a dangling pointer. Whether that document actually supports the sentence beside it is a separate subject with a separate home. - **Presentation and lifecycle states** — pieces arriving over time and assembling cleanly, the marker shown when a length ceiling cuts the answer short, the defined empty state when nothing matched, the control disabled while a request is in flight. - **Permissions** — a caller without access receives the refusal and no fragment of the answer, whatever the generative step produced. - **Persistence and accounting** — what the feature stored against the conversation, what it recorded for billing, what it wrote to its logs. ## What is left over Only the wording. Its desirable properties — a required fact present, a source attached, a tone bound — and the rate at which they hold belong to acceptance work, not to equality assertions. That work is real and necessary, but it is slow, noisy and expensive, and the map exists partly to keep it confined to the smallest possible piece of the feature. | Part of the feature | Expected value | Assertion style | | --- | --- | --- | | Dispatched capability and inputs | one, known before the case runs | equality | | Non-text response fields | fixed for a fixed request | equality | | Citation target resolves | always | equality or reachability | | Empty, cut-short, refused states | one defined rendering each | equality | | Generated wording | none | property-shaped, tolerant | ## Why the map is drawn first 1. **Most defects live in the plumbing.** A wrong capability selected, an argument dropped, a citation pointing at the wrong document, an empty state that renders as a blank panel, a refusal that leaks the first sentence — these are the failures that reach users, and every one of them is exactly assertable. 2. **Without the map, one case swallows everything.** The tempting shortcut is a single end-to-end case that asserts a recorded answer string. It fails whenever anything at all changes in the wording, which trains the team to ignore it, and it passes when the dispatch quietly regresses, because the wording still reads fine. 3. **It makes the tolerant checks affordable.** Once the exact surface is carved out, the variable part is small, and the expensive judgement-shaped checks are aimed only there. ## Drawing it in practice 1. Walk the feature end to end and write down every hop, including authorisation, persistence and rendering. 2. Mark each hop *fixed* or *variable*. Nearly everything is fixed; be suspicious of a second variable hop. 3. For every fixed hop, name the seam where a test can observe it — a returned field, a recorded hand-off, a rendered element, a stored row. 4. Write exact cases against those seams first, and only then decide what the wording itself must hold. ## Where the map stops The map does not tell you whether the answer is any good; it tells you what "the software worked" means independently of that judgement. Both are needed for a release decision, but keeping them apart is what lets a red suite mean something specific instead of meaning that the text changed again.

  • The feature returns a structured result with several fixed fields and one free-text field. Where does the boundary go?
    Around the free-text field only. Every other field keeps an exact expected value for a fixed request, including the field that says how many sources were used and the field that says whether the answer was cut short. Assert those by equality, and let the free-text field carry the tolerant checks. A single string field being variable does not make its siblings variable.
  • What goes wrong when a team asserts the whole answer string as a shortcut for checking everything at once?
    The case becomes a change detector rather than a defect detector. Any rewording turns it red, so the team lowers its tolerance or mutes it, and by then it no longer notices that the wrong capability ran or that a citation points at the wrong document. One broad assertion also gives no localisation: it says the answer differs, never which hop moved.

One room in the building has furniture that moves every day; you still check the doors, the wiring and the locks against exact expectations.

saying these in an interview costs you the question

  • Claims nothing is testable because the output varies each call
  • Asserts the whole answer string equals one recorded example
  • Treats permission refusals as part of the variable output
  • Skips the map and writes a single broad end-to-end case
  • Assumes quality scores for the generative step replace product cases