What is acceptance testing, and what decides whether a change passes it?
answer
- Judged against the ask, not the code
- Someone agreed the wording beforehand
- The criteria are the oracle
- Observable outcome, never internal structure
- Silence in the criteria is itself a defect
basics
~20 sAcceptance testing judges a change against the stated acceptance criteria of the requirement it implements, not against the shape of the code. It passes when every criterion is demonstrably met, and those criteria are agreed before the work starts.
solid answer
~50 sAcceptance testing asks a different question from tests derived from the implementation: not *does the code do what the code says*, but *does the delivered behaviour meet what was agreed*. The thing that decides pass or fail -- the oracle -- is the set of acceptance criteria attached to the requirement, agreed before the work starts by the people who asked for it and the people building it. That is why a change can have every structure-derived test green and still fail here: self-consistent code that implements the wrong thing. Criteria have to be written so the result is observable and arguable by nobody: "the countersigned copy reaches every signer within 40 seconds" is checkable, "signing works reliably" is not. Where the criteria are silent -- what happens when two of three signatures land and the third fails -- acceptance cannot decide anything, and that silence is itself a defect to raise.
code
pseudocode · 12 linescriterion "AC-14: every signer receives a countersigned copy":
round = start_signing_round(document = "NDA-4471", signers = ["ana", "bo", "cy"])
for signer in round.signers:
sign(round, signer)
assert round.state == "completed"
for signer in round.signers:
copy = delivered_copy(round, signer)
assert copy.exists
assert copy.signature_count == 3
assert history_of(round).contains("completed")
assert elapsed(round) < seconds(40)go deeper
Be ready to say in one sentence that acceptance testing checks the change against the requirement's agreed criteria, not against the code, and to give one example of a criterion that is checkable and one that is not.
Explain the mechanics: where the criteria come from, why they must be agreed before implementation, and how a fully green lower-level suite can coexist with a rejected change. Show you can rewrite a vague criterion into an observable one.
Demonstrate that you read criteria for what they omit. Interviewers expect you to spot the unhandled partial failure or the missing budget, take it back to the requirement owner, and turn the answer into an agreed criterion instead of an assumption inside a test.
Own the question of what makes a criterion set trustworthy across many teams: who is allowed to write and change criteria, how they stay traceable to the requirement, and what evidence a passing acceptance run has to leave behind for a decision to rest on it.
### The level, in one line Acceptance testing is the level at which a change is judged against **what was asked for**, expressed as acceptance criteria on the requirement, rather than against the structure of the implementation. Everything else about it -- who runs it, when, in what surroundings, with what evidence -- follows from that one property. ### Criteria as the oracle An *oracle* is whatever tells you the observed behaviour is right or wrong. Tests derived from code carry their oracle implicitly: the developer knew what the function should return and encoded it. Acceptance testing makes the oracle explicit and external -- it is the agreed criteria, and nothing else. That has three consequences worth being able to say out loud. First, the criteria must be agreed *before* the implementation, otherwise they are a description of what was built rather than a test of it. Criteria reverse-engineered from finished code always pass, and prove nothing. Second, each criterion must name an **observable outcome**. Take a document e-signing flow. "Signing is reliable" cannot pass or fail. "When the last signer completes, every signer receives a countersigned copy, and the completion is visible in the document's history within 40 seconds" can: you can run it, look, and be wrong about nothing. A criterion that requires someone to open the implementation to decide whether it was met is not an acceptance criterion. Third, the criteria are a *finite, enumerated* set. Acceptance produces a statement of the form "all 27 criteria for this requirement were demonstrated on build 4471". That statement is the artefact the level exists to produce. ### Why structure-derived tests cannot substitute The classic interview illustration is a change with a completely green lower-level suite that is still rejected. That is not a paradox. Tests written from the implementation inherit the implementation's assumptions, including its misunderstandings. If the requirement said the countersigned copy goes to every signer and the developer understood it as going to the initiator only, the code is self-consistent, its tests are green, and it is wrong. Acceptance testing is the level whose entire purpose is to catch that class of defect -- a *requirement* defect, not a coding defect -- and it is the reason "my tests pass" is never an answer to "is it done". ### What the criteria fail to say The most valuable thing a person does at this level is notice **silence**. Criteria describe the path someone imagined. In the e-signing flow, a realistic gap: two of three signatures are recorded, the archive write succeeds, and the audit-record write fails. Is the document signed? Is the whole round voided and the two signatures discarded, or held? A partial-failure rollback like this is where real systems hurt, and criteria written from the happy path say nothing about it. Raising that gap before implementation is worth more than running the eventual check, because it is cheaper to answer the question than to unpick a half-signed document later. When you find such a gap, the output is a new or amended criterion, agreed the same way as the others -- not a private decision made by whoever is testing. ### Non-functional criteria belong here too Acceptance criteria are not only about behaviour. A criterion may state a budget: the signing round completes within 4.3 seconds at the 92nd percentile of observed rounds. Stated that way it is checkable, and the percentile matters -- an average hides exactly the slow tail users complain about. A criterion phrased as "fast enough" is the same failure as "works reliably" wearing different clothes. ### Who runs it, and what it leaves behind The level does not dictate a single runner. Criteria may be checked by an automated run the delivery team owns, by a person working through them deliberately, or both -- and, separately, the business may hold its own sign-off. What every variant shares is the record: which build, which criteria, what was observed, who witnessed it, when. That record is what makes acceptance a *decision* rather than an impression, and it is the thing an interviewer is listening for when they ask how you would know a change is done. ### A common trap Acceptance testing is not "run the whole regression pack once more before release". A regression pack answers "did anything that used to work stop working". Acceptance answers "does the new thing meet what we agreed". The two overlap in artefacts and never in purpose, and a candidate who collapses them is telling you they have only seen the level as a calendar slot, not as an oracle.
- Every lower-level test is green and the change is still rejected at acceptance. What kind of defect is that?A requirement defect, not a coding defect. Tests derived from the implementation inherit whatever the implementer understood, so a misread requirement produces self-consistent code with a green suite. Acceptance testing exists to catch exactly that. The fix is to correct the criterion or the shared understanding first, then the code, and to ask why the misunderstanding survived until this point.
- You find the criteria say nothing about what happens when part of a multi-step operation fails. What do you do?Raise it as a gap before implementation rather than deciding privately. Take it back to the people who own the requirement, get an answer -- roll the whole operation back, or hold the partial result for retry -- and have it written as an additional criterion, agreed like the others. An untested silence becomes a production argument later, and by then the half-finished state already exists.
- How would you rewrite the criterion "the signing flow must be fast" so it can pass or fail?Name the operation, the measurement and the threshold: the signing round completes within 4.3 seconds at the 92nd percentile over a stated sample of rounds. That is checkable by anyone, without opinion. Prefer a percentile to an average, because an average hides the slow tail that users actually notice, and state the sample so two people measuring get the same answer.
It is the difference between proofreading a translation against the original and checking that the sentences are grammatical. Perfect grammar is no defence if the paragraph says something the original never said.
saying these in an interview costs you the question
- Calls acceptance testing just one more run of the regression pack
- Treats green lower-level tests as proof the requirement was met
- Writes criteria like 'signing works reliably' with no observable outcome
- Derives acceptance criteria from the finished implementation
- Decides an unstated edge case privately instead of raising the gap
- States a timing criterion as 'fast enough' with no threshold