skip to content

Levels of Testing

Unit, integration and end-to-end scopes, and how to balance their proportions along the test pyramid, trading speed and determinism against realism and confidence. Asked because deciding what level a given behavior should be tested at is a daily judgement call.

on this pageshow

questions

24

What is acceptance testing, and what decides whether a change passes it?

level: juniorimportance: must knowfreq 74%

answer

  1. Judged against the ask, not the code
  2. Someone agreed the wording beforehand
  3. The criteria are the oracle
  4. Observable outcome, never internal structure
  5. Silence in the criteria is itself a defect

basics

~20 s

Acceptance testing judges a change against the stated acceptance criteria of the requirement it implements, not against the shape of the code. It passes when every criterion is demonstrably met, and those criteria are agreed before the work starts.

solid answer

~50 s

Acceptance testing asks a different question from tests derived from the implementation: not *does the code do what the code says*, but *does the delivered behaviour meet what was agreed*. The thing that decides pass or fail -- the oracle -- is the set of acceptance criteria attached to the requirement, agreed before the work starts by the people who asked for it and the people building it. That is why a change can have every structure-derived test green and still fail here: self-consistent code that implements the wrong thing. Criteria have to be written so the result is observable and arguable by nobody: "the countersigned copy reaches every signer within 40 seconds" is checkable, "signing works reliably" is not. Where the criteria are silent -- what happens when two of three signatures land and the third fails -- acceptance cannot decide anything, and that silence is itself a defect to raise.

code

pseudocode · 12 lines
pseudocode
criterion "AC-14: every signer receives a countersigned copy":
    round = start_signing_round(document = "NDA-4471", signers = ["ana", "bo", "cy"])
    for signer in round.signers:
        sign(round, signer)

    assert round.state == "completed"
    for signer in round.signers:
        copy = delivered_copy(round, signer)
        assert copy.exists
        assert copy.signature_count == 3
    assert history_of(round).contains("completed")
    assert elapsed(round) < seconds(40)

go deeper

for a junior

Be ready to say in one sentence that acceptance testing checks the change against the requirement's agreed criteria, not against the code, and to give one example of a criterion that is checkable and one that is not.

for a middle

Explain the mechanics: where the criteria come from, why they must be agreed before implementation, and how a fully green lower-level suite can coexist with a rejected change. Show you can rewrite a vague criterion into an observable one.

for a senior

Demonstrate that you read criteria for what they omit. Interviewers expect you to spot the unhandled partial failure or the missing budget, take it back to the requirement owner, and turn the answer into an agreed criterion instead of an assumption inside a test.

for a principal

Own the question of what makes a criterion set trustworthy across many teams: who is allowed to write and change criteria, how they stay traceable to the requirement, and what evidence a passing acceptance run has to leave behind for a decision to rest on it.

### The level, in one line Acceptance testing is the level at which a change is judged against **what was asked for**, expressed as acceptance criteria on the requirement, rather than against the structure of the implementation. Everything else about it -- who runs it, when, in what surroundings, with what evidence -- follows from that one property. ### Criteria as the oracle An *oracle* is whatever tells you the observed behaviour is right or wrong. Tests derived from code carry their oracle implicitly: the developer knew what the function should return and encoded it. Acceptance testing makes the oracle explicit and external -- it is the agreed criteria, and nothing else. That has three consequences worth being able to say out loud. First, the criteria must be agreed *before* the implementation, otherwise they are a description of what was built rather than a test of it. Criteria reverse-engineered from finished code always pass, and prove nothing. Second, each criterion must name an **observable outcome**. Take a document e-signing flow. "Signing is reliable" cannot pass or fail. "When the last signer completes, every signer receives a countersigned copy, and the completion is visible in the document's history within 40 seconds" can: you can run it, look, and be wrong about nothing. A criterion that requires someone to open the implementation to decide whether it was met is not an acceptance criterion. Third, the criteria are a *finite, enumerated* set. Acceptance produces a statement of the form "all 27 criteria for this requirement were demonstrated on build 4471". That statement is the artefact the level exists to produce. ### Why structure-derived tests cannot substitute The classic interview illustration is a change with a completely green lower-level suite that is still rejected. That is not a paradox. Tests written from the implementation inherit the implementation's assumptions, including its misunderstandings. If the requirement said the countersigned copy goes to every signer and the developer understood it as going to the initiator only, the code is self-consistent, its tests are green, and it is wrong. Acceptance testing is the level whose entire purpose is to catch that class of defect -- a *requirement* defect, not a coding defect -- and it is the reason "my tests pass" is never an answer to "is it done". ### What the criteria fail to say The most valuable thing a person does at this level is notice **silence**. Criteria describe the path someone imagined. In the e-signing flow, a realistic gap: two of three signatures are recorded, the archive write succeeds, and the audit-record write fails. Is the document signed? Is the whole round voided and the two signatures discarded, or held? A partial-failure rollback like this is where real systems hurt, and criteria written from the happy path say nothing about it. Raising that gap before implementation is worth more than running the eventual check, because it is cheaper to answer the question than to unpick a half-signed document later. When you find such a gap, the output is a new or amended criterion, agreed the same way as the others -- not a private decision made by whoever is testing. ### Non-functional criteria belong here too Acceptance criteria are not only about behaviour. A criterion may state a budget: the signing round completes within 4.3 seconds at the 92nd percentile of observed rounds. Stated that way it is checkable, and the percentile matters -- an average hides exactly the slow tail users complain about. A criterion phrased as "fast enough" is the same failure as "works reliably" wearing different clothes. ### Who runs it, and what it leaves behind The level does not dictate a single runner. Criteria may be checked by an automated run the delivery team owns, by a person working through them deliberately, or both -- and, separately, the business may hold its own sign-off. What every variant shares is the record: which build, which criteria, what was observed, who witnessed it, when. That record is what makes acceptance a *decision* rather than an impression, and it is the thing an interviewer is listening for when they ask how you would know a change is done. ### A common trap Acceptance testing is not "run the whole regression pack once more before release". A regression pack answers "did anything that used to work stop working". Acceptance answers "does the new thing meet what we agreed". The two overlap in artefacts and never in purpose, and a candidate who collapses them is telling you they have only seen the level as a calendar slot, not as an oracle.

  • Every lower-level test is green and the change is still rejected at acceptance. What kind of defect is that?
    A requirement defect, not a coding defect. Tests derived from the implementation inherit whatever the implementer understood, so a misread requirement produces self-consistent code with a green suite. Acceptance testing exists to catch exactly that. The fix is to correct the criterion or the shared understanding first, then the code, and to ask why the misunderstanding survived until this point.
  • You find the criteria say nothing about what happens when part of a multi-step operation fails. What do you do?
    Raise it as a gap before implementation rather than deciding privately. Take it back to the people who own the requirement, get an answer -- roll the whole operation back, or hold the partial result for retry -- and have it written as an additional criterion, agreed like the others. An untested silence becomes a production argument later, and by then the half-finished state already exists.
  • How would you rewrite the criterion "the signing flow must be fast" so it can pass or fail?
    Name the operation, the measurement and the threshold: the signing round completes within 4.3 seconds at the 92nd percentile over a stated sample of rounds. That is checkable by anyone, without opinion. Prefer a percentile to an average, because an average hides the slow tail that users actually notice, and state the sample so two people measuring get the same answer.

It is the difference between proofreading a translation against the original and checking that the sentences are grammatical. Perfect grammar is no defence if the paragraph says something the original never said.

saying these in an interview costs you the question

  • Calls acceptance testing just one more run of the regression pack
  • Treats green lower-level tests as proof the requirement was met
  • Writes criteria like 'signing works reliably' with no observable outcome
  • Derives acceptance criteria from the finished implementation
  • Decides an unstated edge case privately instead of raising the gap
  • States a timing criterion as 'fast enough' with no threshold

context

open as a page

In a consumer-driven contract test, what does the recorded contract assert about the provider, and what does it deliberately not check?

level: juniorimportance: must knowfreq 42%

basics

~20 s

A consumer-driven contract records the requests one consumer sends and the response parts it actually reads, then asserts the provider can still produce them. It checks the shape of the boundary, not whether the provider's answers are correct.

open as a page

What makes a test end-to-end rather than a narrower test, and what does it prove?

level: juniorimportance: must knowfreq 78%

basics

~20 s

An end-to-end test drives the fully assembled, deployed system through a real entry point, with nothing inside the system replaced by a stand-in, and checks the observable outcome of a complete journey rather than any internal call or state.

open as a page

What defects does an integration test against a real dependency catch that a fully mocked test cannot?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Integration tests catch defects that live in the boundary itself: wiring and configuration, serialization, schema and type mapping, and real query behaviour. A programmed double returns what you assumed, so it can never contradict that assumption.

open as a page

What is the test pyramid, and why are narrow unit tests the widest layer?

level: juniorimportance: must knowfreq 78%

basics

~20 s

The test pyramid is a shape guideline for an automated suite: many fast, narrow unit tests at the base, fewer tests across a real boundary above them, and a thin end-to-end layer on top. Width means test count.

open as a page

What is a unit test, and what does the term unit actually refer to?

level: juniorimportance: must knowfreq 88%

basics

~20 s

A unit test exercises one small piece of behaviour without touching slow or shared parts of the system, so it runs in milliseconds and fails precisely. The unit is a boundary the team chooses, not a fixed size.

open as a page

How do team-owned automated acceptance checks differ from a business user acceptance sign-off?

level: middleimportance: must knowfreq 60%

basics

~20 s

Automated acceptance checks are owned by the delivery team and run on every build to prove the agreed criteria still hold. A user acceptance sign-off is the business owner's one-off judgement that the release is fit to use.

open as a page

Why does an end-to-end test failure take longer to localise than a unit test failure?

level: middleimportance: must knowfreq 61%

basics

~20 s

An end-to-end failure names only a journey, so the suspect set is every component, configuration value and hop it crossed - and the environment state that produced it is gone unless the run captured evidence while it happened.

open as a page

What is a provider state in a consumer-driven contract, and how does the provider apply it during verification?

level: middleimportance: should knowfreq 29%

basics

~20 s

A provider state is a named precondition attached to one recorded interaction, such as "a meter with three unbilled readings exists". Before replaying that interaction, the provider runs the setup hook registered under that exact name, then answers the request.

open as a page

How do you decide how much of the system an integration test should include?

level: middleimportance: should knowfreq 61%

basics

~20 s

Keep one real boundary in scope and make everything else deterministic or absent. Each extra live component multiplies the suspects behind a failure, so a wide test reports that something is broken rather than what is broken.

open as a page

Why is an ice-cream-cone shaped test suite treated as an anti-pattern?

level: middleimportance: should knowfreq 57%

basics

~20 s

An ice-cream-cone suite inverts the pyramid: a broad layer of end-to-end automation, often with manual checking piled on top, over a thin base of narrow tests. Feedback becomes slow, failures stop naming a cause, and instability spreads.

open as a page

How should you name a unit test so its failure explains itself in a report?

level: middleimportance: should knowfreq 62%

basics

~20 s

Name a unit test for the reader of a red report, who sees only the name: state the unit or behaviour under test, the condition it is in, and the outcome expected. A name needing the word and usually means the test checks two things.

open as a page

What does operational acceptance testing check before a system is accepted into live running?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Operational acceptance testing checks that the team who will run the system can run it: install and upgrade, backup and restore, failover, rollback, monitoring and alerting, access management and the written procedures -- each rehearsed rather than assumed.

open as a page

Provider contract verification passes on every build, yet a provider release still broke a live consumer. How does a stale contract produce that false confidence?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Verification replays the contract that was last published, not the consumer's current code. If the consumer changed without re-recording, or the provider verified a version nobody runs, the build proves compatibility with expectations that no longer exist.

open as a page

Your integration tests pass in isolation but fail when the run order changes. How do you find the coupling?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Replay the failing order from its recorded seed, then bisect: run the failing case last behind ever-smaller subsets until one predecessor reproduces it. That pair names the test leaving residue in the shared dependency and the test depending on it.

open as a page

A 27-minute test suite has 213 end-to-end tests and 46 narrow tests: how do you rebalance the levels?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Rebalance by risk, not by ratio. Freeze new top-level tests, ask of each existing one what it uniquely proves, push duplicated checks down to the level where the rule lives, and delete a top-level case only once something cheaper covers it.

open as a page

A unit test for a subscription renewal job passes by day but fails on the nightly run. How do you diagnose it?

level: seniorimportance: should knowfreq 57%

basics

~20 s

Suspect a hidden ambient input. A test that passes at one time of day and fails at another is reading the wall clock or relying on an unspecified ordering. The fix is to pass the instant in and make the expected order explicit, not to loosen the assertion.

open as a page

Acceptance for a regulated release: how do you satisfy a business sponsor's sign-off and an external regulator's evidence needs at once?

level: principalimportance: should knowfreq 32%

basics

~20 s

Keep one traceable criteria set and produce evidence once. Automated acceptance runs emit dated, build-stamped, criterion-linked records; human sign-off is reserved for judgement. The regulator needs traceability and retention, the sponsor needs confidence -- not two separate test efforts.

open as a page

How do you decide which journeys deserve end-to-end coverage, and keep that set from growing?

level: principalimportance: should knowfreq 44%

basics

~20 s

Select by consequence and uniqueness: journeys whose silent failure is unacceptable and whose outcome no cheaper test can observe. Then govern the set with a published budget, an admission rule for adding one, and a push-down habit after every failure.

open as a page

How do you decide whether a unit test may exercise real collaborators, and what does that choice cost?

level: principalimportance: should knowfreq 44%

basics

~20 s

Decide by what the collaborator costs: keep it real when it is in-process, deterministic and fast; substitute it when it reaches outside the process. A narrow boundary localises failures sharply but breaks on refactoring; a wider one survives change but points at a region, not a line.

open as a page

What is the difference between an alpha phase and a beta phase of a release?

level: juniorimportance: nice to knowfreq 44%

basics

~20 s

An alpha phase exercises a pre-release build inside the developing organisation, on controlled machines, with internal staff or invited users. A beta phase puts that build in front of real external users, on their own machines and data.

open as a page

How do big-bang, top-down, bottom-up and sandwich integration strategies differ in ordering and cost?

level: juniorimportance: nice to knowfreq 24%

basics

~20 s

They differ in the order components are joined. Big-bang assembles everything at once; top-down integrates downward using stubs for lower components; bottom-up integrates upward using drivers above; sandwich works from both ends toward the middle.

open as a page

What version assumptions does an end-to-end run make when components deploy independently?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

An end-to-end run is evidence about one combination of deployed component versions at one moment. Where components release independently that combination may never exist in production, and can even change mid-run, so record every version with the result.

open as a page

When is deviating from the canonical test pyramid shape the right call?

level: principalimportance: nice to knowfreq 29%

basics

~20 s

Deviate when the system's real risk does not live inside components. For thin glue services, data pipelines and configuration-heavy systems, weight the boundary-crossing level instead — the invariant is to catch each risk at the cheapest level that can see it.

open as a page