When a feature's generative step is swapped and its wording shifts, which test cases should fail and which must not?
answer
- Two changes, one red suite
- Exact cases never read the answer text
- Wording failures must show before and after
- Break the contract deliberately to prove separation
basics
~20 sNothing on the exact-assertion surface should move: dispatch, response fields, citation rendering, states and permissions keep passing. Only the wording-reading checks may fail, and their output must name the text as what changed rather than the surrounding behaviour.
solid answer
~50 sArrange the suite so the two kinds of change produce visibly different failures. Keep every exact assertion — dispatch and its inputs, non-text fields, citation rendering, streaming, cut-short, empty and permission states — in cases that never read the answer text, so a pure rewording leaves them green. Keep the tolerant checks over the wording separate and clearly labelled, so their failure reads as "the text changed" and carries the before and after. Then the diagnosis is mechanical: **only wording checks red** means the prose moved and someone judges whether that is acceptable; **exact cases red** means the contract broke and the change is not shippable regardless of how the text reads; **both red** means the exact failures are the ones to read first. Prove the arrangement rather than assuming it, by breaking the contract deliberately and confirming the wording checks stay green.
code
pseudocode · 13 linessuite contract_cases: # must stay green when only wording moves
assert dispatch.capability == "search_documents"
assert response.sourceIds == ["doc-14", "doc-22"]
assert view.shows("incomplete_marker") when result.stoppedAtCeiling
assert render_for(guest) == "refused"
suite wording_checks: # may go red; failure is a judgement call
report(request, previous_text, current_text)
check answerText carries required_fact
check answerText attaches at least one source reference
proof_run:
break contract on purpose -> expect contract_cases red, wording_checks greengo deeper
Know that changing what produces a feature's text should not change how the feature behaves, and that separate cases exist for each. Be able to say why one case asserting both is a problem.
Explain the arrangement: structural cases that never read the answer text, tolerant checks kept apart and reporting before and after, and the diagnosis each combination of failures gives.
Show how you prove the separation rather than assume it — a deliberate contract break that must leave the wording checks green — and how you stop those checks eroding across successive swaps.
Own the release rule this enables: which failures block a swap, who judges an accepted wording shift, and what has to be recorded so a later reader can tell text drift from a behaviour regression.
## Two changes that look identical from a distance Swap the generative step behind a product feature — a new instruction set, a different provider, a newer version of the same one — and two very different things can happen at once. The **wording** changes, which is expected and usually harmless. The **contract** may also change: a different capability gets selected, an argument is dropped, a field arrives with a new name, citations stop rendering, the cut-short marker disappears. From the outside both changes present the same way, as a suite going red. A suite that cannot separate them is worth very little at exactly the moment it is needed most. If a routine swap turns the whole thing red, nobody reads it; the team scans the diff, decides it looks fine, and ships. Every real regression riding along with that swap ships too. | What changed | Exact cases | Wording checks | What it means | | --- | --- | --- | --- | | Wording only | green | some red | judge the new text and accept or reject | | Contract only | red | green | a defect, unshippable regardless of the prose | | Both | red | red | read the exact failures first, they localise | | Nothing | green | green | the swap is behaviour-neutral so far | ## Arranging the assertions - **No exact case may read the answer text.** This is the load-bearing rule. One assertion on a phrase inside an otherwise structural case destroys the whole arrangement, because that case now reddens for both kinds of change. - **No tolerant check may carry the contract.** If the only thing verifying that a citation renders is a tolerant check over the answer, then relaxing that check to absorb new wording quietly deletes the citation assertion. - **Separate them physically.** Different files, different suites, different reports. The separation must be visible in the failure output, not only in someone's head. - **Make the wording failures informative.** A wording check that fails should show the request, the previous text and the current text, so the judgement can be made from the report rather than by rerunning things by hand. - **Label the swap.** Record which version of the generative step and which instruction set produced a given run's results, so a later reader can tell a wording shift from a code change. ## Proving the arrangement instead of assuming it The arrangement is a claim about the suite, and claims about suites should be tested. Two cheap exercises do it: 1. **Break the contract on purpose** in a scratch branch — rename the dispatched capability, drop an argument, remove the cut-short marker — and confirm that the exact cases fail and the wording checks do not. If a wording check goes red, it was secretly asserting the contract. 2. **Change only the wording** — by pointing the feature at a deliberately paraphrasing configuration — and confirm the exact cases stay green. If any of them fails, it was secretly reading the text. Run both after any large restructuring of the suite. They take an afternoon and they are the only evidence that the separation is real. ## Living with the wording failures A wording check going red is not automatically a defect; it is a request for a judgement. Someone reads the before and after and decides whether the new phrasing still satisfies what the product promised. That judgement is easier when the checks are shaped around a small number of required properties rather than around whole strings, and it is much easier when the exact cases are green, because then the only open question is the text itself. The failure mode to watch is **erosion**: each time a wording check is inconvenient it is loosened slightly, and after several swaps it asserts nothing. Loosening should be a deliberate, reviewed act with a reason recorded, not a reflex during a release. ## Why this pays The payoff is a specific sentence a team can say with confidence after a swap: *nothing about how this feature behaves has changed; only what it says has changed, and here is exactly how.* No amount of overall greenness or overall redness supports that sentence. Only a suite whose two halves fail independently does, and arranging it that way costs almost nothing if it is done before the first swap rather than after the third confusing one.
- A single end-to-end case asserts both the dispatched capability and a phrase from the answer. What do you do with it?Split it. The structural half moves into the exact cases and keeps its equality assertions; the phrase moves into the wording checks, if it is worth keeping at all. Left as one case it reddens for both kinds of change, which means its failure carries no information and it will eventually be muted.
- How do you keep the wording checks from being loosened into meaninglessness over successive swaps?Make loosening a reviewed change with a recorded reason, the same as any other contract edit, and keep the checks few and property-shaped so each one is worth defending. Reviewing the set periodically against what the product actually promises catches the ones that have been relaxed until they pass anything.
saying these in an interview costs you the question
- Accepts a fully red suite after a swap as normal
- Mixes a phrase assertion into an otherwise structural case
- Relies on tolerant checks to verify the surrounding contract
- Loosens a wording check whenever it is inconvenient
- Never verifies that the two halves fail independently