skip to content

A test passes alone but later tests in the same run see a stand-in they never registered. What went wrong?

level: seniorimportance: should knowfreq 47%

answer

  1. the replacement outlived its test
  2. a cached graph is shared by key
  3. order-dependence is the tell
  4. the inverse leak stays silent
  5. reset and restore in teardown

basics

~20 s

The replacement outlived the test that asked for it — usually a cached graph built with one test's overrides handed to later tests, or a registry mutated at runtime and never restored. Scope replacements to a wiring; reset stand-in state.

solid answer

~50 s

Two mechanisms cover most of these. First, the harness reuses a booted graph, and the graph it reused was built with the earlier test's replacements — so later tests resolve a stand-in they never asked for, with whatever behaviour it was last programmed with. Second, someone edited a shared registry or a process-wide value at runtime and did not restore it, so the leak is not even scoped to one graph. Diagnose by running the suspect test immediately after the earlier one, and by randomising order until the pairing reproduces. The fixes are structural: declare replacements as configuration the graph is built from rather than mutating anything live, keep tests needing different wiring on different wirings, reset programmed behaviour and recorded calls in teardown, and assert in one test that the boundary still resolved to the real implementation — a canary that fails the moment a stand-in leaks.

go deeper

for a junior

If a test only fails when other tests run first, something was shared between them. Look for state left behind rather than for a defect inside the failing test.

for a middle

Explain the two sharing mechanisms — a cached application graph reused by configuration key, and process-wide state mutated at runtime — and why each makes the symptom depend on execution order.

for a senior

Reproduce the pair rather than the single test, assert what actually resolved to separate a wrong binding from leftover behaviour, and fix structurally: scope the replacement, reset in teardown, randomise order in the pipeline.

for a principal

Set the policy that makes this class extinct: replacements declared as configuration, no runtime mutation of shared state, order varied in continuous integration, and a canary that proves the unmodified wiring still boots somewhere.

## The two leak shapes A replacement that shows up where it was not requested has almost always outlived its test in one of two ways. **Leak through a shared graph.** The harness caches booted applications and hands the same one to every test whose configuration matches. If a replacement was applied to that graph, then every test served by it gets the stand-in — including tests that never mentioned it. This is the common case, and it looks like a leak while being ordinary reuse doing exactly what it says. **Leak through shared mutable state.** Something process-wide was changed at runtime — a registry entry, a statically held collaborator, a global clock or default — and never restored. This form is worse: it is not scoped to a graph at all, so it crosses even between tests that boot separately. ## Why the symptom is order-dependent Both shapes depend on which test ran first, so the suite passes in one order and fails in another. Practical consequences: - Running the failing test alone proves nothing except that it is not the culprit. - The informative run is the *pair*: the suspect test immediately after the candidate polluter. - Randomised or reversed ordering surfaces these deliberately, which is why suites that never vary their order accumulate them silently for years. - The failure often migrates when an unrelated test is added or renamed, because that changes execution order — a pattern frequently misfiled as flakiness. ## The inverse leak, which is more dangerous The same mechanism runs the other way: a test that **needs** a stand-in is served a cached graph that was built without one, so it quietly exercises the real collaborator. Nothing fails; the test passes while reaching out to a live dependency, writing real data, or simply taking a slow path nobody notices. A suite that has one leak shape usually has both, and only the noisy direction gets reported. ## Diagnosing in order 1. Reproduce the pair: suspect test immediately after the candidate. 2. Ask what wiring each test requested, and whether the harness considered them equal — identical keys are precisely why the graph was shared. 3. Inside the failing test, assert what the boundary actually resolved to; that separates "wrong object bound" from "right object, wrong programmed behaviour left over". 4. Look for runtime mutation of anything shared: registry edits, static fields, process-level defaults set in setup. 5. Check teardown: a stand-in that is correctly bound but never reset produces the same symptom as a leaked binding. ## Fixes that hold | fix | what it prevents | |---|---| | declare replacements as configuration read before the graph is built | runtime mutation of a live application | | put tests needing different wiring on different wirings | one graph carrying a replacement into tests that never asked | | reset programmed behaviour and recorded calls in teardown | leftovers on a legitimately shared stand-in | | restore any process-wide value in teardown, unconditionally | cross-graph leaks that survive a reboot | | a canary test asserting the real implementation is bound | the silent inverse leak, where a stand-in reaches a wiring that wanted the real thing | | randomised or reversed run order in continuous integration | order-dependence hiding until an unrelated change reshuffles the suite | Dirty-marking the graph after a test that really did corrupt it is legitimate, and expensive: it discards the cached application so the next test boots afresh. Use it as a targeted remedy, not as a blanket policy, or the suite pays a full startup between every pair of tests. ## Why parallel execution sharpens all of this Running tests concurrently turns an order-dependent leak into a timing-dependent one. Two tests served by the same graph can now program the same stand-in at the same moment, and the failure stops being reproducible by ordering at all — the pair trick above no longer works. Suites that intend to parallelise should settle the sharing question first: which wirings exist, which tests share one, and what a test is allowed to mutate. Concurrency does not create these defects; it removes the one handle you had on them. ## What interviewers listen for A weak answer calls it flakiness and adds a retry. A good answer names the sharing mechanism — a cached graph, or unrestored global state — explains why the symptom depends on order, and proposes a structural fix rather than a reordering that merely hides it. The best answers volunteer the inverse leak unprompted, because that is the one costing the team evidence without ever turning a test red.

  • Why is the inverse leak — a test getting the real collaborator instead of its stand-in — harder to notice?
    Because nothing turns red. The test passes while exercising a live dependency, so the only symptoms are indirect: an unexpectedly slow test, data appearing in a shared environment, or a failure that only shows up when that dependency is unavailable. A canary assertion on what actually resolved is the cheap detector.
  • Is running tests in a fixed order a legitimate fix for this?
    No. It hides the coupling until someone adds, renames or parallelises a test and the order shifts. It also leaves the leak available to every future test. Fix the sharing — scope the replacement to the wiring, reset state in teardown — and vary the order deliberately to prove the leak is gone.
  • When is dirty-marking the graph the right response rather than an expensive habit?
    When a test really did put the application into a state you cannot restore — a component shut down, a pool exhausted, a startup-time value overwritten. Then discarding the cached graph is the honest remedy. Applied by default it converts every test boundary into a fresh startup and hides the underlying coupling.

saying these in an interview costs you the question

  • Calls order-dependent failure flakiness and adds a retry
  • Pins the run order instead of removing the shared state
  • Assumes each test always gets a freshly booted application
  • Resets the stand-in only when a failure has already been reported
  • Ignores the silent case where a test quietly gets the real collaborator