A suite passes in order but one case fails when run alone. How do you find the state it depends on?
answer
- Which direction did it fail in?
- Read the case against an empty state
- Replay the same order with a recorded seed
- Halve the predecessor list until one matters
- Fix the setup, never the ordering
basics
~20 sRun the case alone and read what it assumes already exists, usually a record, counter or cache an earlier test left behind. Confirm the coupling by shuffling with a recorded seed, then bisect the predecessors to find the one that supplies it.
solid answer
~40 sFailing alone is the useful direction, because it proves the case never sets up something it asserts on. Start by reading the case against a genuinely empty starting state and listing what it takes for granted. Then reproduce the coupling deliberately: run the suite in a shuffled order with the seed recorded so the same order replays, and bisect, running the failing case after the first half of its predecessors, then after a quarter, until one earlier case is the difference. Fix it at the source by giving the case its own setup, not by pinning the execution order, because ordering hints spread and make the suite impossible to shuffle, shard or parallelise. Then close the hole that let the state survive: the reset that never covered that table, counter or cache.
code
pseudocode · 13 linescandidates = orderedRunUpTo(failingCase) # predecessors in the recorded order
while candidates.size > 1:
half = candidates.firstHalf()
resetState()
run(half)
result = run(failingCase)
if result.reproducedTheOriginalOutcome():
candidates = half
else:
candidates = candidates.secondHalf()
report("coupled to", candidates.single())go deeper
Be able to run a single case on its own and notice that it needs data no step inside it creates. That one observation is most of the diagnosis at this level.
Explain the two directions — failing alone means missing setup, failing only in the suite means a predecessor pollutes — and describe shuffling with a recorded seed so a failing order replays exactly.
Demonstrate the search: bisect the predecessor list, name the medium that carries the coupling beyond stored records, and argue why fixing the case's own setup beats pinning the order.
Own the systemic answer — shuffle by default with a printed seed, resets derived from the live schema, and a pipeline that runs sampled cases alone so coupling surfaces the week it is written.
### Two directions, two diagnoses Order dependence shows up in two mirror-image ways, and naming which one you have is most of the work. **Fails alone, passes in the suite.** The case is missing setup. It asserts on something an earlier case happened to create, and the full run supplies it by accident. This is the friendlier direction, because the failure is deterministic and the fix is local. **Passes alone, fails in the suite.** An earlier case leaves behind state this one cannot tolerate: extra rows in a count, an advanced identifier counter, a warmed cache, a mutated global setting. Here the culprit is a *predecessor*, and finding it is a search problem. Both are deterministic given the same order, which is the crucial distinction from a case that is nondeterministic on unchanged code. Confirm determinism first by running the same order twice; a stable result means you are hunting a dependency, not a race. ### Step one: read the case against an empty world Before reaching for tooling, run the case on a freshly reset state and read it line by line, asking of every value it references: which step in this case creates that? In a loyalty-points ledger suite, the answer is frequently "none of them". The case looks up member 4 and awards points; member 4 exists only because an earlier case created it. It asserts that a tier lookup returns "silver"; that reference row was inserted by a reseed the case never invokes. Half of these are found by reading, which is faster than any search. ### Step two: make the order reproducible Suites that always run in the same order, whether that is file-system order or declaration order, hide coupling indefinitely and then break the day someone renames a file. The instrument is a shuffled order **with the seed recorded and printed on failure**, so a failing order replays exactly. A shuffle without a reported seed is worse than no shuffle: it turns a deterministic bug into one nobody can reproduce. ### Step three: bisect the predecessors For the "fails only in the suite" direction, the search is a bisection over the predecessor list. Run the failing case after the first half of the cases that preceded it in the recorded order. If it still fails, the culprit is in that half; if it passes, it is in the other. Over a suite of 1,412 cases that is about eleven rounds, and if a full run takes 27 minutes, each round costs far less because you only run a subset plus the one case. Tooling that automates the halving exists, but the manual loop is short enough to run by hand and is worth being able to describe. The strongest probe in the other direction is running each case in its own fresh process against a reset state. Anything that fails there is missing setup, unambiguously, with no search required. ### What actually carries the coupling When you find it, name the medium, because the fix differs: - **Rows** an earlier case created and never removed, usually because cleanup sat below a failing assertion. - **Identifier counters** that keep climbing, which break any assertion on a literal identifier. - **In-process caches and memoised lookups** that survive a database reset entirely. - **Global configuration, a feature flag, a clock or locale** that one case set and never restored. - **Messages, files and third-party sandbox records** outside the store the reset touches. - **A table the reset never covered**, typically added by a migration after the reset list was written, which reads at first like a schema-drift mismatch rather than a cleanup gap. ### Fix at the source, never by pinning the order The tempting fix is to force the failing case to run after its supplier. It works, and it is debt. The coupling is preserved and hidden in configuration; the suite can no longer be shuffled, sharded or parallelised without carrying the pin; and the case still cannot be run alone, which destroys the fastest debugging loop available. Give the case its own setup instead, then close the hole that let the state survive between cases. ### Parallelism changes the shape With several workers there is no single global order to shuffle, so pin the run to one worker first to recover a deterministic order and bisect there. What remains after that is contention rather than ordering: two cases addressing the same key at overlapping times. Reproduce that by running just the two repeatedly, and fix it by making the key unique per run or per worker, not by serialising the suite. ### The systemic answer One fixed case is worth little if the next one appears next week. The durable posture is mechanical: reset before each case rather than only after, shuffle by default with the seed printed, derive the reset from the live schema rather than a hand-written list, and have the pipeline run a sample of cases individually so a new coupling is caught in the week it is written rather than the quarter it becomes load-bearing.
- A colleague fixes it by forcing the failing case to run after its supplier. What is wrong with that?It preserves the coupling and hides it in configuration. The suite can no longer be shuffled, sharded or run in parallel without carrying the pin, and the next person who reorders anything reintroduces the failure somewhere else. It also means the case still has no setup of its own, so running it alone stays impossible, and running it alone is the fastest debugging loop there is.
- The suite runs eight workers in parallel, so there is no single order to shuffle. Does the technique still apply?Partly. Pin the run to one worker first to recover a deterministic order, because that is where bisecting works. What is left after that is contention rather than ordering: two cases addressing the same key at overlapping times. Reproduce that by running just those two repeatedly, and fix it by making the key unique per run or per worker rather than by serialising the suite.
- How would you stop the whole class of problem rather than this one case?Make the isolation contract mechanical rather than cultural. Reset before each case as well as after, shuffle by default with the seed printed on failure, derive the reset from the live schema instead of a hand-written list, and add a pipeline step that runs a sample of cases individually. Then a new coupling fails in the week it is introduced instead of the quarter it becomes load-bearing.
Bisecting the predecessor list is the same move as bisecting a commit history: halve the range, keep the half that still reproduces, and repeat until one item is left.
saying these in an interview costs you the question
- Calls it unstable without checking the order is deterministic
- Pins the execution order instead of giving the case setup
- Assumes file-name order is a guarantee the suite can rely on
- Shuffles without recording the seed, so nothing replays
- Looks only at stored records, ignoring caches and counters
- Deletes the failing case rather than fixing its setup