skip to content

Your integration tests pass in isolation but fail when the run order changes. How do you find the coupling?

level: seniorimportance: should knowfreq 55%

answer

  1. Deterministic, not nondeterministic
  2. Reproduce before you investigate
  3. Halve the predecessors, keep the victim last
  4. Who wrote it, who assumed it
  5. Ownership beats a pinned order

basics

~20 s

Replay the failing order from its recorded seed, then bisect: run the failing case last behind ever-smaller subsets until one predecessor reproduces it. That pair names the test leaving residue in the shared dependency and the test depending on it.

solid answer

~50 s

Order dependence means one test's leftovers in a shared stateful dependency have become an input to another, so it is deterministic rather than nondeterministic and can be chased directly. First make it replayable: a randomised order is only useful if the seed is recorded and the runner accepts it back. Then bisect — keep the victim last and halve the preceding set until a single predecessor reproduces the failure. The pair tells you which side is wrong: usually a predecessor that wrote or mutated something it does not own — rows, a queue, a cached entry, a process-wide setting such as a locale or a frozen clock — or a victim asserting on a total count, a generated identifier or a 'first result' that only holds on an untouched dependency. Fix by ownership, not by pinning the order, which hides the coupling until the next test is added.

code

pseudocode · 10 lines
pseudocode
candidates = order_from_seed(8271).before(victim)   # 337 cases
while size(candidates) > 1:
    left, right = split_in_half(candidates)
    if run(left + [victim]).failed:
        candidates = left
    else if run(right + [victim]).failed:
        candidates = right
    else:
        report("needs two predecessors together"); break
report(candidates)   # the writer; victim is the reader

go deeper

for a junior

Be ready to recognise the pattern: a test that passes alone but fails in a suite is usually reading something another test left behind. Know that each test should create the data it asserts on rather than assuming a clean slate.

for a middle

An interviewer at this level expects the procedure: replay the order from a recorded seed, confirm the case passes alone, then bisect the predecessors down to a pair. Be able to list several kinds of leftover state, not just database rows.

for a senior

Show production judgement: distinguish coupling from nondeterminism by reproducibility, decide from ownership which of the pair is defective, and reject the pinned-order, retry and sleep remedies with reasons rather than taste.

for a principal

Own the systemic angle: randomised, seeded ordering as a standing policy so coupling surfaces the day it appears, the ownership convention that makes parallel execution possible at all, and the quarantine rules that keep a pipeline honest without silently shedding coverage.

## What order dependence actually is An integration suite shares something stateful — a data store, a broker, a cache, a filesystem area. Order dependence means one test's leftovers in that shared thing have become an input to another test. Nothing is random about it: the same order gives the same result every time. It only looks mysterious because the order changed for an unrelated reason — a new test was added, the runner sharded differently, someone enabled parallel execution, or the runner deliberately shuffles. This is worth separating from nondeterminism. A case that fails once in forty runs on unchanged code with a stable order is nondeterministic — timing, concurrency, real clocks. A case that always fails in one order and always passes in another is **coupled**, and coupling is far easier to chase because it is reproducible the moment you can replay the order. ## Where the residue hides The candidates, roughly in order of how often they turn out to be the culprit: - **Rows or documents left behind** by a test that inserted more than it removed, or removed nothing. - **Monotonic counters** — generated identifiers, sequence values, version numbers — that a victim asserted on literally. - **Undelivered or unconsumed messages** sitting in a queue when the next test starts consuming. - **Cached state in the process**: an in-memory cache, a lazily built singleton, a connection or session carrying settings from the previous test. - **Mutated global settings**: the runtime locale, default time zone, a fixed clock left frozen, an environment variable changed for one test and never restored. - **Files and directories** in a shared area, including ones a previous test left half-written. The last two categories are the ones teams miss, because everybody looks at the database first. ## Make the failure reproducible before anything else If the ordering is randomised, the run must print a seed and the runner must accept that seed to replay the identical order. A random order with no replayable seed converts a solvable coupling problem into an unsolvable one. If ordering is fixed but sharded, you need the shard assignment as well. The goal of this step is a command that fails every single time. ## Bisect on the predecessors With replay in hand, the search is mechanical: 1. Confirm the victim passes alone. If it fails alone, the problem is not coupling. 2. Run the victim last after the whole preceding set. Confirm it fails. 3. Halve the preceding set, keeping the victim last. Whichever half reproduces the failure contains the culprit; if neither half does alone, the cause needs two predecessors together, which is rare but happens with counters. 4. Repeat until one predecessor plus the victim reproduces it. That pair is the finding. On a 340-case regression pack the search costs about nine runs, which is usually less time than one engineer spends reading logs. ## A worked case A marketplace bidding engine has a 340-case regression pack that turned red on the build server after a shuffle seed changed. Three cases failed, all asserting on the closing price of an auction. Replaying seed 8271 reproduced them exactly; each passed alone. Bisection put the victim at position 338 behind successively smaller sets and landed on a single predecessor: a case that had been added to check that a bid amount is stored with the correct decimal separator, and which set the process locale to a grouping-separator format and never restored it. Every later test that formatted a price for comparison then produced `1.847,50` where it expected 1847.50. The store had nothing to do with it; the shared state was a process-wide setting, and no amount of data cleanup would have found it. ## Decide which side is wrong A pair does not tell you which test to change. Ask what each test owns: - If the predecessor **wrote or mutated something it does not own** — rows outside its own keyed data, a global setting, a shared file — the predecessor is the defect. Confine it: give it uniquely keyed data, and restore anything global it changes even when it fails. - If the predecessor is well behaved and the victim asserted on something it never created — a total row count, "the first result", a specific generated identifier, an empty queue — the victim is the defect. Make it query for its own uniquely-keyed records and assert only on those. Both fixes are about **ownership**: a test asserts on what it created, and leaves behind nothing that was not there when it started. How the shared dependency is reset between tests, and by which mechanism, is a test-data concern that belongs to a different discipline; the ownership rule is what makes the reset strategy nearly irrelevant. ## The fixes that are not fixes Pinning the execution order makes the red go away and leaves a suite that is one insertion away from breaking again, with the coupling now invisible. Retrying the failed case buries a deterministic defect under a mechanism meant for nondeterminism. Adding a wait changes nothing, because the residue is not a timing problem. Quarantining is defensible for an hour to keep the pipeline honest, but a coupled test is a known, reproducible defect, and leaving it quarantined discards the coverage while the coupling stays in the suite. ## Prevention worth stating in an interview Randomise the order on purpose, with a recorded and replayable seed, so coupling surfaces on the day it is introduced rather than months later. Have every test address uniquely-keyed data. Restore any global setting in teardown that runs even after a failure. And treat "passes only in this order" as a defect report about the suite, not a scheduling requirement.

  • How do you tell order coupling apart from a genuinely nondeterministic case?
    By reproducibility. Coupling fails every time in one order and passes every time in another, so replaying the seed reproduces it exactly. Nondeterminism fails intermittently on unchanged code even with the order held constant, and its causes are timing, concurrency, real clocks or external latency. Different diagnosis, different fixes.
  • The pair is found. How do you decide which of the two tests to change?
    By ownership. If the predecessor wrote or mutated something outside its own keyed data — rows it did not create, a shared file, a process-wide setting — the predecessor is at fault and must confine and restore. If the predecessor behaved and the victim asserted on a total count, a specific generated identifier or 'the first result', the victim is at fault and must assert only on records it created.
  • Why is pinning the execution order a bad remedy?
    It removes the symptom and preserves the defect. The suite still contains a test that leaves residue and one that depends on it, but now nothing will reveal it until an insertion, a shard split or a parallel run shifts the order again — and by then the connection to the original change is lost. It also blocks parallel execution permanently.
  • Which shared state do teams overlook when everyone is staring at the database?
    Process-wide state: a mutated locale or default time zone, a clock left frozen, an environment variable changed for one case, an in-memory cache or lazily built singleton, and a connection or session carrying settings forward. Also undelivered messages left on a queue and files left in a shared area.

It is a shared kitchen: every cook works fine alone, and the mess only becomes a bug when the next cook assumes the counter was wiped.

saying these in an interview costs you the question

  • Pins the run order and calls it fixed
  • Retries the case as if it were nondeterministic
  • Adds a wait to let residue disappear
  • Randomises order without recording a replayable seed
  • Only inspects the database, never process-wide state
  • Quarantines the case indefinitely and moves on

context