skip to content

How do you prove that one test in a suite depends on another test running first?

level: middleimportance: should knowfreq 57%

answer

  1. Change one thing: the position
  2. Alone first, then reversed, then shuffled
  3. Record what produced the failure
  4. Halve the tests that ran before it
  5. Guard the assumption at the top

basics

~20 s

Run the suspect test on its own: failing alone but passing in the full run means it consumes state another test leaves. Then reverse or shuffle the order to reproduce it, and bisect the preceding tests to find the pair.

solid answer

~40 s

Start with the cheapest experiment: filter the run down to the single test. Passing alone but failing in the suite means something earlier leaves state it cannot tolerate; failing alone but passing in the suite means it consumes state an earlier test creates; failing both ways means the test or the code is simply broken and ordering is a red herring. Next, vary the order deliberately — reverse the declared order, or shuffle with a **recorded seed** so the exposing order can be replayed after the fix. Then bisect: run the first half of the preceding tests plus the suspect, then the second half, and recurse until you have the pair. A guard assertion at the *start* of the suspect test, checking the state it assumes, turns future recurrences into a self-describing failure.

code

pseudocode · 13 lines
pseudocode
test "rate plan import rejects a flattened payload":
    # guard: fail here, with a clear message, if an earlier test leaked state
    assert store.count(property_code: "PROP-4417") == 0,
        "expected no rows for PROP-4417 before arrange; an earlier test leaked state"

    arrange:
        store.insert(property_code: "PROP-4417", plans: 2)

    act:
        result = importer.load(flattened_payload)

    assert result.rejected == true
    assert result.reason == "SCHEMA_MISMATCH"

go deeper

for a junior

Know the first move: run the single test on its own and compare with the full run. Be able to say what each of the three outcomes — passes alone, fails alone, fails both ways — tells you about where the problem lies.

for a middle

Walk through the whole procedure: isolate, reverse, shuffle with a recorded seed, bisect the preceding tests to a pair. Explain why a fix that reorders or pins tests is not a fix, and where file-scoped setup hides state.

for a senior

Show the discrimination step that separates an order dependency from nondeterministic failure before you spend hours on the wrong hypothesis, and describe how you would leave a durable guard so the coupling cannot silently return.

for a principal

Decide how this is prevented rather than hunted: randomised order in the pipeline, seeds captured in run output, and a policy on new tests being runnable alone. Own the cost argument for spending build time on shuffled runs.

An order dependency is a test whose verdict changes with its position in the run. Proving one exists, and then naming the two tests involved, is a routine debugging skill because the failure never announces itself honestly: the symptom is usually "it passes on my machine and fails on the build server", or the reverse, and both are really "it passes in one order and fails in another". ### Start with the cheapest signal: run it alone Filter the run down to the single suspect test and execute it by itself. Three outcomes, three conclusions: - **Passes alone, fails in the suite.** Something earlier in the run leaves state this test does not tolerate — a mutated shared fixture, a record it did not expect to exist, a frozen clock, a changed global setting. - **Fails alone, passes in the suite.** The test relies on state an earlier test creates. This is the chained-test case, and it is the more dangerous of the two because the suite looks green. - **Fails both ways.** Not an ordering problem at all; the test or the code under test is broken, and the ordering hypothesis should be dropped before it eats an afternoon. ### Then make the ordering variable, not fixed A single alternate order proves little. Change the order deliberately and repeatedly: - **Reverse the declared order.** Cheap, deterministic, and it catches the plain "B reads what A wrote" case immediately. - **Shuffle with a recorded seed.** Randomising execution order surfaces dependencies you did not predict, and recording the seed is what makes the discovery actionable: without it you have an unreproducible failure, and with it you can replay the exact order that broke and re-run it after the fix to confirm the fix. - **Run the suspect test first, then last.** A crude two-point version of the same idea, useful when the runner cannot shuffle. ### Then narrow to the culprit pair Once you know the test fails in some orders and not others, the remaining work is a search for the partner. Bisect the set of tests that run before it: run the first half plus the suspect, then the second half plus the suspect, and recurse into whichever half reproduces the failure. Each round halves the candidates, so even a large file converges in a handful of runs. The result is a pair — the test that leaves the state and the test that reads it — and the pair is what you fix. If bisecting is awkward, instrument instead: assert at the *start* of the suspect test that the state it depends on is in the shape it expects. A test that begins by asserting "no records exist for this identifier" turns a mysterious downstream assertion failure into an explicit, self-describing message naming the leak. This start-of-test guard is also a durable regression check: it keeps failing loudly if someone reintroduces the coupling later. ### Reading the evidence Three details separate a real ordering defect from other explanations: **It is deterministic given the order.** Fix the order and the outcome is stable. That is what distinguishes an order dependency from a genuinely nondeterministic failure, which varies run to run with the order held constant. Establishing that stability early is worth one extra run because the two classes of problem have completely different investigations. **The failing assertion is often not the interesting one.** In a coupled pair, the visible failure is downstream of the real event — a count that is one too high, an identifier that already exists, a default that was overwritten. Read the assertion as a symptom and look for who last wrote that state. **Shared state hides in setup, not only in test bodies.** A setup hook that runs once for the whole file rather than once per test is the single most common source, because it looks like ordinary arrangement code while actually being a fixture with a lifetime longer than any one test. ### Turning the proof into prevention A fix that only reorders the two tests, or pins the declared order, is not a fix — it re-creates the dependency with a comment on top. The fix is to give the dependent test its own arrangement and to make the writer clean up after itself. Then keep the property: run the suite in a randomised order in the pipeline so the next dependency is caught by a machine on the day it is introduced, rather than by a person weeks later, when the change that introduced it is no longer obvious.

  • Why record the seed when you shuffle execution order rather than just shuffling?
    Without the seed the failure is unreproducible: you saw a red run, cannot replay the order that caused it, and cannot prove a fix worked. With it, the exact permutation is replayable, so you can reproduce on demand while debugging, confirm the fix against the order that broke, and attach the seed to the ticket. Shuffling finds dependencies; recording the seed is what makes the discovery actionable rather than folklore.
  • How do you tell an order dependency apart from a test that fails nondeterministically?
    Hold the order fixed and re-run several times. An order dependency is deterministic given the order — the same permutation always produces the same verdict — while a nondeterministic failure varies with the order held constant. That one extra experiment is worth doing early, because the two problems have completely different investigations, and searching for a leaking neighbour when the real cause is timing wastes the day.
  • Your bisect points at a pair, but the failing assertion is in a third test. What now?
    Read the failing assertion as a symptom rather than a location. In a coupled chain the visible failure sits downstream of the real event — an unexpected count, an identifier that already exists, an overwritten default. Ask who last wrote the state the assertion reads, and look in setup that runs once per file before you look in test bodies, since a file-scoped hook is the most common owner of long-lived state.

saying these in an interview costs you the question

  • Reorders the two tests and calls it fixed
  • Pins the declared order with an explanatory comment
  • Adds a retry until the suite goes green
  • Never runs the suspect test on its own
  • Assumes any position-dependent failure is a timing race
  • Shuffles order without recording the seed

context