skip to content

For a feature whose answer text comes from a generative step, what makes an acceptance criterion adjudicable?

level: middleimportance: must knowfreq 64%

answer

  1. Fixed expected text stops working
  2. An impression is not a criterion
  3. Describe a property of the response
  4. Observable, locus, decision rule
  5. Two readers must reach one outcome

basics

~20 s

An adjudicable acceptance criterion names one observable property of the response, says where to look for it, and fixes the decision rule in advance, so two testers reading the same response record the same pass or fail.

solid answer

~50 s

"The answer should be good" is not a criterion: it names nothing observable and fixes no decision rule. Rewrite it as a **property of the response** that one person can settle by reading the response alone - a required figure must appear, a stated date must match the record, a length bound must hold, a forbidden statement must be absent. Each property needs three parts: **what** must be true, **where** it is read, and **which rule** decides. Write the check so a colleague with no context could apply it, and attach one worked response that meets it beside one that does not, because the pair fixes the boundary that prose cannot. If two honest testers would split on the same response, the wording is still subjective and the criterion, not the response, is the defect.

code

pseudocode · 11 lines
pseudocode
CRITERION "quoted-balance-matches-record"
  observable: the money figure in the summary sentence
  locus:      reply.summary
  rule:       figure == accountRecord.balance(accountId)

  met_example:     "Your balance is 412.60."
  not_met_example: "Your balance is about 400."

adjudicate(reply):
  figure = read(reply.summary)
  return figure == accountRecord.balance(accountId)

go deeper

for a junior

Be ready to say why a test case for a generated reply cannot pin one exact expected string, and to offer one property you could check instead, such as a required figure appearing in the reply.

for a middle

An interviewer expects the mechanics: name the observable, say where it is read, and fix the decision rule so the outcome does not depend on who ran the case. Show one vague line rewritten into a checkable one.

for a senior

Show that you build the criteria from the product's promise before any case exists, defend the boundary examples attached to each, and explain what you do when two testers still split on the same response.

for a principal

Own the standard: which promises the organisation is willing to state as checkable properties, what it costs to keep them adjudicable across releases, and when a promise is too soft to ship as a criterion at all.

## Why "the answer should be good" is not a criterion A feature whose reply is produced by a generative step returns different words for the same input on different days, and both replies can be entirely acceptable. That is the nature of the component, not a defect in it. What it breaks is the ordinary habit of writing a test case whose expected result is a fixed string — the comparison that made the case decidable is gone. The tempting substitute is an impression: "the answer should be good", "the summary should be helpful", "the tone should be right". None of those is a criterion. Each describes how a reader feels about a response rather than something the response contains. Two honest testers reading the same paragraph will land differently, so the outcome of the case depends on who happened to run it. A suite whose results depend on the reader is not evidence about the product. ## The three parts of an adjudicable statement An acceptance criterion is adjudicable when one person, reading one response and nothing else, can record met or not met without asking anyone. That takes three parts. 1. **The observable** — the concrete thing the obligation is about: a figure, a named entity, a date, a phrase, a structural element, or the absence of one. If you cannot point at it in a printed response, there is nothing to adjudicate. 2. **The locus** — where in the response it is read. "Somewhere in the reply" is weak; "in the field the interface renders as the summary" is strong, and it keeps working as replies get longer. 3. **The decision rule** — the exact test applied, written to leave no discretion: present or absent, equal to a value computed independently, within a stated bound, or one of an enumerated set. Miss any of the three and the statement drifts back toward taste. ## Rewriting impressions into properties | Impression | Adjudicable property | |---|---| | The answer should be accurate | The balance stated in the reply equals the balance the account record holds | | The answer should be current | Every figure quoted matches the record as read when the request was made | | The tone should be professional | The reply contains no exclamation and no second-person command | | The answer should be concise | The reply is at most 120 words | | The answer should not overpromise | The reply states no date the order record does not contain | Each right-hand statement names a thing, a place and a rule. Notice that none of them says anything about how well the reply reads. That is deliberate: the criteria describe the product's promise, and a response either carries it or does not. ## Boundary examples do the work prose cannot Even a carefully worded property leaves a band where readers split. The cheapest repair is a pair of worked responses attached to the criterion — one that meets it and one that does not, chosen as close to each other as you can make them. The pair fixes where the author drew the line far better than another sentence of qualification, and it survives handover: a tester who joins next quarter calibrates from the examples instead of re-deriving intent from the wording. When two testers still disagree about a real response, treat that as a defect in the criterion rather than in the testers. The wrong repairs are all tempting. Appoint an arbiter whose reading settles it, and the ambiguity survives for everyone else. Repeat the call until a reply everyone likes appears, and the evidence is gone. Quarantine the case as unstable, and the promise it protected is no longer checked by anything. The right repair is to sharpen the wording and add the disputed response as a new boundary example. ## What to fix before the first case exists - The list of properties, derived from what the feature undertakes to do rather than from responses already seen. - The decision rule for each, in words a colleague with no context could apply. - A met and a not-met example beside each property. - Which properties block and which are merely recorded when missed. - How often each property is allowed not to hold, agreed with whoever owns the promise. The last two are separate obligations in their own right and both have to be settled before anyone runs a case. A criterion silent on either is applied as absolute and blocking by everyone who reads it. ## What adjudicable does not mean Adjudicable is not the same as automated. Plenty of good properties are checked by a person reading a response, and that is fine as long as any person reaches the same result. It is also not the same as narrow: "the reply names the delivery window recorded on the order" admits thousands of wordings and still decides cleanly. The test is never how tight the criterion is. It is whether two readers of the same response can honestly disagree.

  • A criterion says the reply must be "appropriately detailed". How would you rewrite it?
    Replace the impression with a countable observable and a bound: the reply must contain the three items the request asked for and be at most 150 words. If detail genuinely varies by request type, split the criterion per type rather than hedging one statement into vagueness. Whatever survives has to let one person reading one response record met or not met without consulting whoever wrote it.
  • Why attach a met example and a not-met example to each acceptance criterion?
    The pair fixes the boundary that prose cannot. Wording that reads unambiguous to its author usually leaves a band where readers split, and two worked responses either side of the line show exactly where the author drew it. They also survive handover: a tester who joins later calibrates from the examples instead of re-deriving intent, and a disagreement becomes an argument about which side an example falls on.

It works like the checklist a building inspector carries: it never says the building should look nice, it says which things must be present and exactly where to find each one.

saying these in an interview costs you the question

  • Treats "the answer looks right" as an acceptance criterion
  • Pins the expected result to one exact response text
  • Leaves the decision to whoever happens to run the case
  • Writes properties nobody can check without asking the author
  • Assumes any difference in wording means the case failed