How do you write test cases so an eloquent generated answer that misses the product's promise still fails?
answer
- Write criteria before reading responses
- Fluency is not an obligation
- Derive from the requirement, not the output
- Check obligations one at a time
- Keep beautiful-but-incomplete counter-examples
basics
~20 sDerive every criterion from the product's stated promise before any response is read, then adjudicate each response against that list only. Fluency, extra detail and pleasing tone are not criteria, so a well-written answer that omits a promised item fails.
solid answer
~50 sThe trap is writing criteria after reading a batch of generated responses: the responses are articulate, the criteria drift toward describing them, and the suite ends up certifying the generative step's current habits instead of the product's promise. Break the order. Write the criteria from the requirement - what the feature undertakes to tell the user, on what backing, within what bounds - and freeze them before any response is read. Then adjudicate each response against that list and nothing else. Two disciplines make it stick: check obligations one at a time rather than forming an overall impression, and keep a set of counter-example responses that read beautifully yet omit a required item, so a reviewer's instinct stays calibrated. Content the promise never asked for is neither credit nor failure, unless a prohibition covers it.
code
pseudocode · 14 linesWRONG ORDER
1. call the feature 20 times
2. read the replies
3. write criteria describing what those replies did
-> the suite certifies today's wording, not the promise
RIGHT ORDER
1. read the requirement:
"tell the customer the delivery window on their order"
2. derive and freeze:
MUST reply states order.deliveryWindow
MUST reply states no date order does not carry
SHOULD reply is at most 60 words
3. call the feature; adjudicate each reply against step 2 onlygo deeper
Be ready to say that a reply reading well is no evidence it kept the promise, and to point at the requirement as the source of what a case checks.
Explain the mechanics: derive each criterion from the stated requirement, freeze the list before responses are read, and adjudicate obligation by obligation instead of forming a single impression.
Show how you detect a suite that has drifted into describing the generative step's habits, and how counter-examples that read well yet omit a required item keep a reviewer's instinct calibrated.
Own where the promise is written down and who may change it, so that criteria across teams trace to one statement of what the feature undertakes rather than to whatever each team happened to observe.
## Two possible sources for a criterion Every acceptance criterion on a generated reply comes from one of two places, and the choice decides what the suite is able to detect. | Source | How it is written | What it certifies | How it fails | |---|---|---|---| | The product's promise | Read the requirement, state the observable it implies, freeze it before any response is read | That the feature does what was undertaken | Hard to write when the promise itself is vague | | The observed responses | Call the feature, read what comes back, describe it | The generative step's current habits | Passes forever, including when the promise breaks | The second is the default failure mode of a team under time pressure, and it is seductive precisely because the responses are articulate. You call the feature twenty times, the replies are fluent and well organised, and the criteria written afterwards describe what you just admired. The suite becomes a mirror. It will happily certify a release in which the feature stopped stating the delivery window, because no criterion ever said it had to — every response that was read happened to include one. ## Why fluency reads as correctness The trap is not laziness. It is a reliable perceptual effect. - Fluent prose signals competence, and a reader transfers that signal to the content. - A confident sentence about a fact is read as a checked fact. - Extra detail feels like generosity, so a reply that answers a question nobody asked reads as better rather than off-target. - An absence is much harder to notice than an error. A wrong figure jumps out; a missing figure inside a well-structured paragraph does not. - Reviewing a reply as a whole yields one overall impression, and impressions are dominated by prose quality. Together these mean a reviewer forming a global judgement passes promise breaks at a rate they would not believe of themselves. ## Working in the right order 1. **Read the requirement, not the output.** Write down what the feature undertakes to tell the user, on what backing, within what bounds. 2. **Derive the observables.** Each undertaking becomes a property with a locus and a decision rule. 3. **Freeze the list.** Record that it was written before any response was read, and who agreed it. 4. **Only then call the feature.** Adjudicate each response against the frozen list, item by item. 5. **Change the list only through the requirement.** If a response reveals a genuine gap, the repair is a change to what the product promises, agreed with whoever owns it, and the criterion follows from that. Step four is where the discipline is actually applied. Checking obligations one at a time — is the figure present, does it match the record, is the date the one on the order, is the forbidden statement absent — defeats the impression effect in a way that reading the reply and asking "is this good?" never does. ## Counter-examples keep instinct calibrated Beside the criteria, keep a small set of worked responses that read beautifully and fail. A polished two-paragraph reply that omits the promised figure. A warm, well-structured answer that quotes a figure taken from the wrong record. A helpful reply that answers a broader question than the one asked. New reviewers read these before their first session and learn what a promise break looks like when the prose is good. Without them every reviewer discovers the effect alone, usually after shipping one. ## Content the promise never asked for Extra content is out of scope by default. It earns no credit — the case turns on the promise, not on generosity — and it is not a failure unless it breaches a prohibition, such as inventing a commitment or advising outside the product's remit. If the extra content is consistently valuable, that is a proposal to change what the product promises, made in the open, and not a criterion someone adds quietly because the responses happened to contain it. ## How to tell a suite has drifted Ask, of any passing case, which sentence of the requirement it verified. If the answer is a description of the reply rather than a clause of the promise, that criterion came from the output. Two more symptoms are worth watching for: criteria whose wording mirrors the phrasing of the responses, and a suite that has never failed for any reason other than an outage or a formatting change.
- A reviewer passes an answer because it "clearly understood the question". What do you challenge?That the sentence names no obligation. Ask which listed criterion the reply met and where in the reply it is read. Comprehension is inferred rather than observed, and it is exactly the impression a fluent reply produces most strongly. If the promise really does include something the current criteria miss, add it as an observable property rather than letting the impression stand as the outcome.
- How do you handle a generated answer that gives more than the promise required?Extra content is out of scope unless a criterion covers it. It earns no credit, because the case turns on the promise, and it is not a failure unless it breaches a prohibition such as inventing a commitment or advising outside the product's remit. If the extra content is consistently valuable, propose a change to what the product promises rather than adding a criterion quietly.
saying these in an interview costs you the question
- Writes criteria after reading a batch of generated replies
- Passes a response because it reads well overall
- Treats extra unrequested detail as evidence of correctness
- Forms one overall impression instead of checking each obligation
- Cannot say which promise a passing case actually verified