In acceptance criteria for an AI-powered feature's answer, why separate must, should and must-never?
answer
- Not every expectation is equal
- Promise, preference, prohibition
- Three obligations, three consequences
- A missed should does not fail
- One breached prohibition stops everything
basics
~10 sBecause the three carry different obligations. A must failing means the promise was not kept and the case fails; a should failing is recorded but does not block; a must-never failing stops everything regardless.
solid answer
~50 sFlattening every expectation into one list makes each response either perfect or broken, which is unusable when wording varies. Split them by what a failure licenses. **Must** carries the product's promise: if the required observable is absent, the case fails and a defect is raised. **Should** carries preference - a tone bound, a preferred ordering, an offered next action; a miss is recorded as an observation, feeds the quality trend, and does not fail the case on its own. **Must-never** carries prohibition: an invented commitment, advice the product is not licensed to give, disclosure of something another user owns. A breach ends the case immediately and is escalated regardless of how well every must was satisfied. The split also tells a reviewer what a mixed result means: many shoulds missed is a trend, one must-never breached is an incident.
code
pseudocode · 12 linesCRITERIA for feature "order summary reply"
MUST reply states order.deliveryWindow
MUST every amount in reply appears in order
SHOULD reply is at most 80 words
SHOULD reply opens with the outcome, not the reasoning
MUST_NEVER reply states or implies a refund decision
MUST_NEVER reply promises a date order does not carry
adjudicate(reply):
if any MUST_NEVER holds -> stop, escalate, block release
if any MUST fails -> fail case, raise defect
if any SHOULD fails -> record observation, case passesgo deeper
Be ready to name the three obligations and what each does when unmet: a must failing fails the case, a should failing is noted, a must-never failing stops everything. One example of each on a real feature is enough.
Explain the mechanics of the split: where each obligation is recorded, how a mixed result is reported, and why promoting every preference to a must makes a varying response impossible to ship.
Show judgement about which promises earn prohibition status, how a breach is escalated beyond an ordinary defect, and how you keep the should list from becoming a graveyard nobody reads.
Own the policy: who may add a prohibition, what evidence promotes a preference into the promise across releases, and how the three lists stay honest as the feature's promise changes.
## Why one flat list fails Gather everything anyone wants from a generated reply into a single list of expectations and every response becomes either perfect or broken. That works when output is fixed. It collapses when wording varies, because some of those expectations are the promise the product made and others are how the team would prefer the reply to read. Flattened together, a reply that keeps every promise fails because it opened with the reasoning instead of the outcome, and the suite starts blocking releases over taste. The split into must, should and must-never is not a taxonomy exercise. Each obligation exists because it licenses a different action when it is not met. ## Three obligations, three consequences | Obligation | What it carries | Example on a reply summarising an order | When unmet | |---|---|---|---| | **Must** | The product's promise | The reply states the delivery window recorded on the order | The case fails and a defect is raised against the feature | | **Should** | A preference | The reply opens with the outcome before the reasoning | The case still passes and the miss is recorded as an observation | | **Must-never** | A prohibition | The reply states that a refund has been approved | The case stops and is escalated beyond an ordinary defect | - **Must** is what the feature undertook to do. If the required observable is absent, the promise was not kept, and quality elsewhere does not compensate for it. - **Should** is preference: tone, ordering, an offered next action, brevity beyond the hard bound. A miss is real information — it is how you see quality drifting — but it does not block, and pretending it does makes acceptable variation look like breakage. - **Must-never** is prohibition. It covers behaviour whose consequences reach outside the session: an invented commitment a customer could act on, advice the product is not licensed to give, disclosure of something another user owns. A breach ends the case immediately and travels a different reporting path, because the question it raises is not "does this build ship" but "has this already reached a user". ## What each licenses the tester to do 1. **Must unmet** — record the case as failed, keep the response verbatim, raise a defect against the feature, and stop counting that response toward the promise. 2. **Should unmet** — record an observation against the case, let the case pass, and let the accumulated observations argue for a change once the trend is clear. 3. **Must-never breached** — stop that path, preserve the response, notify whoever owns the risk, and treat release as blocked until the behaviour is understood. ## Common mistakes in the split - **Everything becomes a must.** Preferences promoted to blocking obligations mean a reply that keeps the promise in unusual wording now fails. The suite loses the ability to distinguish "the product broke its word" from "the product phrased it oddly". - **Nothing becomes a should.** With no preference list, quality drift is invisible. Nobody records that replies have grown steadily more verbose across three releases, because no criterion covers it. - **Prohibitions nobody can observe.** "The reply must never mislead" has no observable in it. Rewrite it as the concrete statement the reply must not make. - **A prohibition recorded as an observation.** The most damaging inversion: a reply asserts a commitment the product cannot honour, and it is logged as a tone note because the rest of the reply read well. ## The split as a reporting instrument Once the three lists exist, a mixed result reads immediately. Twelve shoulds missed across forty responses with every must met is a quality trend for the next planning session. One must missed is a defect. One must-never breached is an incident, whatever the other numbers say. Without the split all three arrive as "the suite is at eighty-eight percent", which tells the person deciding nothing about what kind of trouble the feature is in. Two boundaries are worth keeping straight. First, a must-never is a test obligation; the enforcement that stops a prohibited statement reaching a user while the product is running is a separate design concern with its own failure modes. Second, promoting a should to a must is a product judgement backed by the recorded observations, not something the person writing the cases decides quietly in order to make a suite look stricter.
- A stakeholder wants every should promoted to a must. What do you tell them?That the promotion has a price: each promoted item now blocks a release, so a reply that keeps the whole promise in unusual wording will fail. Ask which of them the product would genuinely refuse to ship without. Whatever survives that question becomes a must; the rest stay as recorded observations, so the quality trend is still visible without holding the feature hostage to phrasing.
- How is a must-never different from a must phrased in the negative?A negatively phrased must is still one obligation weighed alongside the others, and its breach produces an ordinary defect. A must-never is absolute: it ends the case whatever else the response did, and it carries a reporting path rather than a backlog entry. Reserve it for behaviour with consequences outside the session, such as an invented commitment a customer could act on.
saying these in an interview costs you the question
- Treats every expectation as equally blocking
- Fails a case because a stated preference was missed
- Records a prohibited statement as a minor observation
- Writes prohibitions nobody can observe in the reply
- Keeps no shoulds at all, so quality drift goes unrecorded