skip to content

Test Structure

How a single test is laid out and how tests relate to each other: a consistent shape inside, and no dependence on order or shared state between. Asked because a suite nobody trusts is worse than no suite, and structure is what earns that trust.

on this pageshow

questions

13

What are the three phases of the Arrange-Act-Assert test structure?

level: juniorimportance: must knowfreq 78%

answer

  1. A test body reads in ordered beats
  2. Preparation comes before invocation
  3. The middle beat is usually one line
  4. A behaviour vocabulary names the same three

basics

~20 s

Arrange builds the fixture, inputs and collaborators the test needs. Act invokes the one behaviour under test and captures its result. Assert compares that outcome with the expected one. Given-When-Then names the same three phases.

solid answer

~40 s

Arrange-Act-Assert is a layout convention for a single test body. **Arrange** builds the world the behaviour needs: the object under test, input values, seeded data, stand-ins for dependencies. **Act** invokes the behaviour exactly once and captures what comes back — a value, a raised error, or an observable state change. **Assert** compares that outcome against the expected one. The phases are separated by blank lines and kept in that order, so the line number of a failure already tells you which phase broke. Given-When-Then is the same three phases in specification vocabulary: Given is arrange, When is act, Then is assert. The strongest signal a test follows the convention is a one-statement act: when the act runs to six lines, most of it is arrange in disguise, and two genuine acts mean two tests.

code

pseudocode · 10 lines
pseudocode
test "adding a track increases the playlist length"
    # arrange
    playlist = playlistWith(trackCount = 12)
    track    = trackNamed("Halcyon Drift")

    # act
    result = playlistService.add(playlist.id, track.id)

    # assert
    expect(result.trackCount).toEqual(13)

go deeper

for a junior

Be ready to name the three phases in order and say what belongs in each, then point at them in a test you have written. Interviewers use this as a screening question, so a crisp answer matters more than a long one.

for a middle

An interviewer expects you to explain why the act is one statement and what a multi-statement act really contains, and to map the phases onto Given-When-Then without treating them as two different patterns.

for a senior

Show the diagnostic payoff: the phase boundaries are what make a failing line number informative, and you should be able to describe how you steer a team's tests toward that shape in review rather than by decree.

for a principal

Own the consistency argument. A suite where every test reads in the same three beats is navigable and cheap to onboard onto; the tradeoff you defend is uniformity of shape against the cost of reworking existing tests that predate the convention.

## What the convention is Arrange-Act-Assert (AAA) is a layout convention for the body of a single test. It says a test should read top to bottom in three visually separated phases: build the world the behaviour needs, invoke the behaviour exactly once, then check what came back against what you expected. No runner enforces it and nothing fails if you ignore it. Its entire value is what it does to a human reader six months later, at the moment a test goes red and someone has to decide within seconds whether the product broke or the test did. ### Arrange Everything the behaviour needs before it can run: constructing the object or handle under test, building the input values, seeding whatever store or collaborator the code reads from, configuring stand-ins for dependencies, freezing a clock, and registering cleanup. Arrange answers the question "what world does this behaviour live in?" It is normally the longest phase, and when it grows past a handful of lines the usual remedy is to push construction behind a named builder or factory function so the test still reads as one idea. ### Act One invocation of the behaviour under test, with its result captured in a local variable where the behaviour returns one. The result may be a returned value, a raised error, or an observable state change. The act phase should be one statement, or very close to it. If it is six statements, the first five are almost always arrange in disguise — reaching the state you want is preparation, not the thing you are testing. A test with two genuinely distinct acts is two tests. ### Assert The statements that compare the observed outcome against the expected one. The thing that decides the behaviour is wrong is called the oracle, and in a unit test the oracle is these assertions plus the expected values in them. The assertions should express one logical outcome, so a red test names one cause. ## Why the order matters The three phases correspond to the three questions a person debugging a failure asks in order: what was set up, what was done, what was expected. Keeping them in that order and separating them with a blank line means the failing line number alone tells you which phase broke. Interleaving act and assert destroys that property: when the third assertion fails you no longer know whether the state it inspects came from the act you care about or from the two calls you made in between. ## The same shape under another vocabulary Given-When-Then names the identical three phases in specification vocabulary — Given is arrange, When is act, Then is assert. Teams that write executable specifications and teams that write plain unit tests are using one structure with two vocabularies, which is why a scenario translates mechanically into a test body. The difference is register and audience, not shape. There is also a longer relative: the four-phase test described by Meszaros — setup, exercise, verify, teardown. AAA is that pattern with teardown left implicit, because most modern runners handle cleanup through a registered hook rather than trailing statements a failed assertion would skip. ## What the structure buys you in practice - A failure's line number identifies the phase, so triage starts from a fact rather than a read-through. - A reviewer can see at a glance that a test acts more than once, which is the commonest cause of a test whose name no longer describes what it checks. - Arrange bloat becomes visible instead of being spread through the method, which is what prompts extracting a builder. - Tests in a suite start to look alike, and uniform tests are scannable; a suite of 900 tests that all read in three beats is navigable, one where every test invents its own shape is not. ## Common mistakes - Hiding the act inside an assertion, so the call under test is evaluated as an argument. A raised error then surfaces as an assertion problem rather than as the behaviour it is. - Writing the three phase names as comments and then ignoring them, so the comments drift and lie. Blank lines and a one-line act carry the structure without maintenance. - Putting checks in the arrange phase and calling them setup, without deciding whether a failing one should be reported as a broken test or a failed expectation. - Sharing the arrange phase across the whole file so that reading a test tells you nothing about the state it runs against. ## Worked shape ```pseudocode test "adding a track to a playlist increases its length" # arrange playlist = playlistWith(trackCount = 12) track = trackNamed("Halcyon Drift") # act result = playlistService.add(playlist.id, track.id) # assert expect(result.trackCount).toEqual(13) ``` Three beats, one act, one outcome. If this goes red, the line number alone tells you whether the fixture could not be built, the call blew up, or the behaviour returned something else.

  • Where do you create a stand-in for a dependency, and why there?
    In arrange. Building and configuring a stand-in is part of describing the world the behaviour runs in, not part of exercising it. Putting the configuration next to the act blurs the boundary, and a failure in that configuration then reports at a line the reader takes for the behaviour under test.
  • Your act phase needs three calls before the interesting one runs. What does that tell you?
    The first three are arrange. Reaching the state you want is preparation; only the call whose outcome you assert on is the act. Move them up, ideally behind a named builder so the test still reads in three beats. If two of the calls are genuinely behaviours you want to check, that is two tests, not one.
  • Does anything enforce the structure at runtime?
    No. It is a reading convention, carried by ordering and blank lines rather than by any runner feature. That is also why phase-name comments drift: nothing checks them. A one-statement act and a blank line between phases carry the structure with no maintenance cost.

It reads like an experiment write-up: set up the apparatus, perform the one measurement, record whether the reading matched the prediction. Mixing the steps makes the write-up unrepeatable.

saying these in an interview costs you the question

  • Thinks a test runner enforces the three phases
  • Says the act phase may span several unrelated calls
  • Cannot say which phase builds the input data
  • Claims Given-When-Then is a structurally different pattern
  • Puts the call under test inside an assertion argument

context

open as a page

What is a flaky test, and why is it worse than a test that fails every time?

level: juniorimportance: must knowfreq 78%

basics

~20 s

A flaky test gives different verdicts on unchanged code, passing on one run and failing on the next. A consistent failure names a real problem; a flaky one teaches the team to re-run instead of investigate, destroying the suite's signal.

open as a page

What makes a test independent of the other tests in its suite?

level: juniorimportance: must knowfreq 78%

basics

~20 s

An independent test arranges everything it needs itself, writes to no state another test can observe, and leaves the environment as it found it. It gives the same verdict alone, in any position in the run, or beside tests on another worker.

open as a page

What is the assertion roulette test smell, and what refactoring removes it?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Assertion roulette is a test carrying many unnamed assertions, so a failure reports only a line number and you must guess which check broke. Remove it by naming each assertion or splitting the test so one behaviour gets one test.

open as a page

What are the most common causes of a test that fails intermittently on unchanged code?

level: middleimportance: must knowfreq 68%

basics

~20 s

The usual families are timing assumptions, concurrency races, uncontrolled inputs such as the real clock, time zone or random values, calls to real external systems, dependence on the order cases run in, and state leaked between cases.

open as a page

In Arrange-Act-Assert, does one logical assertion per test mean one assert call?

level: middleimportance: should knowfreq 61%

basics

~20 s

No. The guideline is one logical outcome per test, which several assertion statements may describe together. What it rules out is a test making unrelated claims, and above all a test that acts more than once.

open as a page

How do you prove that one test in a suite depends on another test running first?

level: middleimportance: should knowfreq 57%

basics

~20 s

Run the suspect test on its own: failing alone but passing in the full run means it consumes state another test leaves. Then reverse or shuffle the order to reproduce it, and bisect the preceding tests to find the pair.

open as a page

Why is a conditional or loop inside a test body a smell, and what replaces it?

level: middleimportance: should knowfreq 46%

basics

~20 s

Nothing verifies test code, so a branch or loop in a test adds an unchecked path that can report success while checking nothing. Replace a branch with two named tests, and a loop with case-per-input reporting or a whole-value assertion.

open as a page

How do you distinguish an arrange-phase breakage from a genuine assertion failure?

level: seniorimportance: should knowfreq 43%

basics

~20 s

Make each phase self-identifying: end arrange with precondition checks, keep setup where the runner reports it as an error rather than a failed expectation, and keep the act to one captured statement so a failing line names its phase.

open as a page

How do you tell a flaky test from a genuine intermittent defect in the product?

level: seniorimportance: should knowfreq 50%

basics

~20 s

They look identical in a run report, so separate them by cause, not frequency. Reproduce with evidence captured at the moment of failure, then decide whether the nondeterminism sits in the test's assumptions or in the behaviour it observed.

open as a page

Why does a suite of order-coupled tests turn one real defect into dozens of failures, and what does that cost?

level: seniorimportance: should knowfreq 49%

basics

~20 s

Coupled tests inherit each other's state, so when one fails its dependents fail on arrangements that never happened. The red count then measures how far the breakage spread, not how many defects exist, and triage, re-verification and trust all pay for it.

open as a page

How do you remove duplicated setup from a large test suite without creating a mystery guest?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Extract construction, not meaning: a builder or object mother with defaults lets each test override only the value it asserts on, keeping that value visible. A mystery guest is a test whose expectation depends on unseen data.

open as a page

Your team has stopped trusting red builds because most failures turn out to be flakes. How do you fix that?

level: principalimportance: should knowfreq 42%

basics

~20 s

Treat the lost trust as the problem, not the red runs. Decide what a red result must oblige, stop new nondeterministic cases entering the blocking path, fund the cleanup as real work, and keep one gate whose failures the team believes.

open as a page