skip to content

Before writing test cases for an AI-powered feature, why agree how often each acceptance criterion may miss?

level: middleimportance: must knowfreq 57%

answer

  1. Usually right is not always right
  2. Attach an allowance to each criterion
  3. Agree the number before any case exists
  4. An unstated allowance reads as absolute
  5. First miss becomes an argument, not data

basics

~20 s

A criterion with no stated allowance is read as absolute, so the first response that misses it forces an unplanned argument. Agreeing the allowance up front makes a rare miss an expected, recorded outcome rather than a crisis.

solid answer

~50 s

A feature built on a generative step is usually right, not always right, and the acceptance criteria have to say which. Write each criterion with its allowance attached: this must hold on every reply; this must hold on at least nine of ten replies to the same request; this may miss on an unusual request provided the degraded text appears. Agree those numbers with whoever owns the product's promise **before** any case is written, because afterwards the number is negotiated under pressure with a failing result on the table. A criterion silent on the point is applied as absolute by everyone who reads it. The first miss then splits the team three ways - repeat until green, quietly reword the criterion, or quarantine it as unstable - and all three destroy the record of how often the feature actually keeps its promise.

code

pseudocode · 11 lines
pseudocode
CRITERION "delivery-window-stated"
  obligation: MUST
  check:      reply states order.deliveryWindow

  allowance:  holds on at least 9 of 10 replies
  sample:     10 replies to the SAME request, kept verbatim
  counting:   case fails when misses > 1

  agreed_by:  owner of the product promise
  agreed_on:  2026-03-04
  relaxed:    (empty - every change recorded here with a reason)

go deeper

for a junior

Be ready to say that a feature whose reply is generated is usually right rather than always right, so a criterion has to state how often it may miss before anyone runs a case.

for a middle

Explain the mechanics: attach an allowance with a sample size and a counting rule to each criterion, and describe the three bad outcomes when the first miss arrives with no allowance agreed.

for a senior

Show how you get the number signed by the owner of the promise, how you stop it creeping upward release after release, and what evidence you attach when you propose changing it.

for a principal

Own the standard for allowances across features: which promises may carry one at all, what the organisation is willing to state publicly about them, and who answers when the agreed rate is not met.

## Usually right is a different contract A feature whose reply comes from a generative step keeps its promise most of the time. "Most" is not a defect to be engineered away before release; it is the contract the feature is being built under. An acceptance criterion therefore has to say what "most" means, in numbers, before anyone runs a case against it. Otherwise the number gets decided by accident, at the worst possible moment, by whoever is in the room when the first miss appears. ## What an unstated allowance does A criterion with no allowance attached is read as absolute. Nobody writes "and this must hold every single time" because it feels implied by the plain obligation, and that is exactly how the person running the case applies it. The first response that does not meet it is a failure nobody planned for, and the team splits three ways. - **Repeat until green.** The case is executed again, a conforming reply appears, and the miss vanishes from the record. The suite now hides the very property it was written to expose. - **Reword after the fact.** Somebody softens the criterion with "generally" or "where appropriate", under pressure, with a failing result on the table. The number that emerges is whichever number makes this particular failure acceptable. - **Quarantine the case.** It is excluded from reported results as unstable, and the promise it protected is no longer checked by anything. All three destroy the same thing: the record of how often the feature actually keeps its promise. That record is the only asset the suite was building. ## Writing an allowance that stays adjudicable An allowance is useful only if the person applying it has nothing left to invent. "Mostly holds" and "usually present" are impressions again. A usable allowance fixes four things. 1. **The sample** — how many responses, and to what. Ten responses to the same request is a different statement from ten responses to ten different requests, and the two answer different questions. 2. **The counting rule** — how many misses are tolerated before the case fails, stated as a count against that sample rather than a feeling about the proportion. 3. **The evidence** — that the responses are kept verbatim, so a miss can be re-read later instead of argued about from memory. 4. **The agreement** — who signed the number, and when. | | No allowance stated | Allowance stated | |---|---|---| | How it reads | The reply must state the delivery window | The delivery window must appear in at least nine of ten replies to the same request | | First miss | An unplanned argument under pressure | An expected outcome, recorded against the allowance | | What it records | Whether one reply conformed | Whether the feature holds the promise at the agreed rate | | Its own failure mode | Quiet rewording after a failure | The agreed number creeping upward release after release | The right-hand column has a failure mode too, and naming it is part of doing this honestly. An allowance is a number people can renegotiate. Guard it by recording every change with its date and its reason, so that a reviewer can see at a glance that a criterion has been relaxed three times in a year. ## Where the number comes from The allowance is a product decision, because it says how often a user may be let down. The person writing the cases proposes it and makes the consequence concrete; the owner of the promise signs it. The inputs are ordinary. - What a miss costs the user. An unhelpful reply is not the same as a wrongly stated commitment. - Whether the user can tell. A miss noticed immediately is cheaper than one the user acts on. - Whether a degraded path catches it. If something usable is rendered when the required item is absent, the same miss costs less. - What the alternative costs. Driving a rate down usually means extra checks, repeated calls or narrower scope, and each has its own price. ## What the allowance is not Two neighbouring decisions get confused with this one. The allowance beside a criterion is a property of that criterion: it makes the criterion decidable for one tester holding one set of responses. It is not the overall failure rate the organisation finally decides a release may carry, which is a separate call taken with every criterion in view and other evidence beside it. Nor is it a measurement programme over a labelled corpus; that machinery answers "how often does this hold across everything we have seen", while the allowance answers "does this case pass today". Keeping the two apart matters, because a criterion that quietly imports a measurement result as its pass rule becomes undecidable for the person actually reading the responses.

  • Who should agree the allowance, and why not the person writing the cases?
    The allowance is a product decision: it says how often the feature may fall short of what a user was promised. The person writing the cases proposes it and makes the consequence concrete, but the owner of the promise signs it. Set by the case author alone, it quietly becomes whatever number makes the suite green, and nobody outside the team learns what was conceded.
  • What makes an allowance adjudicable rather than a vague "mostly"?
    It fixes the sample and the counting rule: how many responses, to which request, kept how, and how many misses are tolerated before the case fails. Without those, "mostly holds" is decided afresh by whoever is reading. State them beside the criterion and keep the responses verbatim, and the outcome stops depending on which session happened to see the miss.

An allowance is the tolerance on a machined part: it belongs on the drawing before anything is cut, not in an argument after the first component is measured.

saying these in an interview costs you the question

  • Writes every criterion as absolute for a varying reply
  • Negotiates the allowance only after the first failing session
  • Repeats the call until green instead of recording the miss
  • States an allowance with no sample size or counting rule
  • Lets the person writing the cases set the number alone