skip to content

What is the assertion roulette test smell, and what refactoring removes it?

level: juniorimportance: must knowfreq 58%

answer

  1. The failure names a line, not a behaviour
  2. Which of the checks actually broke?
  3. Later assertions never ran
  4. Name it, or split it by behaviour
  5. One composed value asserted once

basics

~20 s

Assertion roulette is a test carrying many unnamed assertions, so a failure reports only a line number and you must guess which check broke. Remove it by naming each assertion or splitting the test so one behaviour gets one test.

solid answer

~50 s

Assertion roulette is a test that ends in a run of assertions, none of which identifies what it checks, so a failure gives you a line number and a bare expected-versus-actual pair and the reader has to reconstruct the intent. It is a diagnosability smell rather than a correctness one: the assertions may all be right, but every future failure costs minutes of reading. There is also hidden evidence, because the first failing assertion stops the test and the checks after it never run. The refactorings, cheapest first: give each assertion a message or replace it with a named domain assertion helper; split unrelated behaviours into separate tests so the failing test name is the diagnosis; compare one composed value in a single assertion so all differences are reported at once; or use an aggregating assertion block when every check for one behaviour really must be reported.

code

pseudocode · 8 lines
pseudocode
test "published timetable is valid":
    timetable = planner.publish(term)
    assertEquals(31, timetable.slotCount)
    assertEquals(12, timetable.roomCount)
    assertEquals(0, timetable.clashes.size)
    assertTrue(timetable.slots[0].start < timetable.slots[1].start)
    assertEquals("B-14", timetable.slots[3].room)
    // ... nine more, none of them named

go deeper

for a junior

Be ready to define the smell in one sentence and give the two cheapest fixes: an assertion message, or one test per behaviour. Knowing that the failure report is what the smell damages is most of the answer.

for a middle

Explain the mechanics: assertions stop at the first failure, so later checks are skipped and evidence is hidden. Show a named assertion helper and a whole-value comparison, and say why splitting increases setup cost.

for a senior

Demonstrate judgement about where the line sits — multiple assertions on one behaviour are fine, unrelated behaviours are not — and describe how you would find and prioritise these tests in a suite you inherited rather than mass-splitting them.

for a principal

Own the trade: a suite is a diagnosis tool, and mean time to understand a failure is the metric worth managing. Talk about assertion vocabulary as a shared asset, and about why threshold-based lint on test code produces compliance rather than legibility.

### What the smell is **Assertion roulette** is Meszaros' name for a test that ends in a run of assertions, none of which says what it is checking. When one fails, the report gives you a line number and a bare expected-versus-actual pair. Whoever picks up the red build has to open the test, count down to that line, and reconstruct from the surrounding setup which behaviour was supposed to hold. Hence the name: the failure tells you that *something in this test* broke, and you spin the wheel to find out what. The smell is about **diagnosability**, not correctness. Every assertion in the test may be right and the test may be catching a genuine defect; the cost is paid by the next person, in minutes of reading that a named assertion would have cost zero. That cost is charged on every failure, forever, which is why it is worth a refactoring rather than a comment. ### A worked example A school timetable planner has one test that publishes a term timetable and then checks it: fourteen assertions covering teacher load, room capacity, slot ordering, and the clash count. The build goes red with `expected 31 but was 30` on line 96. Thirty-one what? Slots, rooms, teachers, minutes? The number carries no domain meaning, the test name (`publishedTimetableIsValid`) carries none either, and the only way forward is to read the whole method. There is a second, quieter cost. Assertions run in sequence and the first failure stops the test, so assertions twelve to fourteen never executed. In this suite the room-mapping change that broke slot counting had also silently mis-assigned rooms — a corruption that assertion thirteen would have caught. The report showed one failure, the team fixed that one, the build went green, and the corrupted rooms shipped. Unnamed sequential assertions do not just delay diagnosis; they **hide later evidence** behind the first failure. ### The refactorings There are four standard moves, and they are not alternatives so much as a ladder. **1. Name the assertion.** The cheapest fix: give each assertion a failure message that states the behaviour, or replace it with a domain assertion helper whose *name* is the message — `assertNoRoomClash(timetable)` instead of a raw equality on a count. The failure now reads as a sentence about the system, and the helper is reusable across the suite. **2. Split by behaviour.** If the assertions cover unrelated behaviours, the test is also an *eager test* — one test exercising several behaviours in sequence. Split it so each behaviour gets a test named after it. Then the failing **test name** is the diagnosis and you never read the body at all. This is the fix that pays best over time. **3. Assert on a composed value.** Where several assertions pick apart one result, compare the whole result to one expected value in a single assertion. A structural comparison reports *all* the differences at once instead of stopping at the first field, which directly fixes the hidden-evidence problem above. **4. Aggregate the assertions.** When you genuinely want every check reported for one behaviour, use an aggregating (soft) assertion block: the checks all run, failures are collected, and the report lists them together. Use this deliberately — it is a reporting tool, not a licence to keep a fourteen-assertion test. ### Where the line actually is "One assertion per test" is a slogan, not the rule, and quoting it is a common way to answer this question badly. Several assertions that together verify **one** behaviour are perfectly healthy: a result object checked field by field is one logical assertion spread over several statements. The smell has two ingredients — assertions that do not identify themselves, **and** a test whose scope spans behaviours that fail for different reasons. Fix either ingredient and the roulette goes away. Conversely, splitting has a cost: fourteen tests each rebuilding a term timetable is fourteen times the setup. That pressure is what pushes teams toward extracted setup, and the extraction has its own failure mode when the extracted data disappears from view. Splitting is right; splitting without a readable way to build the fixture is how one smell is traded for another. ### Detecting it Assertion roulette is one of the few test smells a tool can flag cheaply: count assertions per test method and flag methods above a threshold with no assertion messages and no custom matchers. Treat the count as a *prompt to look*, never as a gate — a thresholded metric on test code reliably produces tests that split into two halves that both still smell. The better signal is behavioural: when a test fails, time how long it takes a person who did not write it to say what broke. If that number is minutes, the suite has this smell whatever the counter says.

  • Is more than one assertion in a test always this smell?
    No. Several assertions that together verify one behaviour are healthy — checking four fields of one result is one logical assertion spread over four statements. The smell needs two ingredients: assertions that do not identify themselves, and a test whose scope spans behaviours that fail for different reasons. 'One assertion per test' is a slogan; the real target is a failure that explains itself.
  • Why does this smell hide defects rather than only slow diagnosis?
    Assertions run in sequence and the first failure ends the test, so every check after it is skipped. A change that broke two things reports one, the team fixes that one, the build turns green, and the second defect ships. Comparing a whole value in a single assertion, or using an aggregating assertion block, makes all differences visible in one run.
  • How would you spot this smell across a large existing suite?
    Count assertions per test method and flag the methods above a threshold that use no assertion messages and no custom matchers. Treat that as a prompt to look, not a gate — a thresholded metric mostly produces tests split into two halves that both still smell. The stronger signal is behavioural: how long it takes someone who did not write the test to say what broke.

A failing test should read like a smoke alarm labelled 'kitchen'. Assertion roulette is an unlabelled panel of fourteen identical alarms: you know something is burning, but not where.

saying these in an interview costs you the question

  • Claims the rule is strictly one assertion per test
  • Thinks the smell is about test runtime or speed
  • Says it does not matter because the test still catches the bug
  • Adds a comment above the assertions instead of naming them
  • Deletes assertions to shorten the test
  • Unaware that later assertions are skipped after the first failure

context