A unittest test passes alone but fails under `python -m unittest discover` — how do you diagnose it?
answer
- The difference is what else ran first
- One process, one import of each module
- Reproduce the pair, not the suite
- Reverse the order to prove it
- Fix the sharing, not the ordering
basics
~20 sDiscovery imports every test module into one process, so the suite shares module-level state. Reproduce by running the two modules together by dotted name, then bisect with -k, -f and -v until the polluting pair is isolated.
solid answer
~40 sThe asymmetry itself is the clue: a whole discovery run is a **single process**, so every discovered module is imported once into `sys.modules` and any module-level or class-level state is shared by everything that follows. In a payroll CSV importer that means one module warming a cached rate table and a later module reading the stale cached value. Reproduce it cheaply: `python -m unittest tests.test_rate_cache tests.test_csv_import -v` runs just that pair in that order. Then bisect, narrowing with `-k` and stopping early with `-f`, and add `--buffer` so only the failing test's output survives. Confirm the direction by reversing the two names — a failure that moves is order dependence, not a bug in the test. The fix is to stop sharing the state, not to pin the run order.
code
console · 2 linespython -m unittest tests.test_rate_cache tests.test_csv_import -v
python -m unittest tests.test_csv_import tests.test_rate_cache -vgo deeper
Learn the key fact behind this: a discovery run is one process, so tests can affect each other through module-level variables. Know how to run two named modules together to see it happen.
Explain why sys.modules caching makes import-time state stick for the whole run, and use -v, -f, -b and -k to narrow a failure to a pair of tests rather than re-running everything.
Show the whole loop: form the order-dependence hypothesis from the symptom, reproduce with two names, prove it by reversing them, then remove the sharing rather than pinning the order. Be explicit that ordering hacks convert a defect into a trap.
Frame it as suite design: shared module-level caches are the root cause, and the durable answer is a convention for building and disposing of expensive state, plus a CI signal that catches order dependence before it lands rather than a quarantine list of flaky tests.
This is a leak diagnosis, and the shape of the evidence tells you that before you read a line of the test. ### Why discovery changes the outcome `python -m unittest discover` does not fork or isolate anything. It walks the tree, imports every matching module into one interpreter, builds a single `unittest.TestSuite`, and runs it in that same process. Three consequences follow: 1. **Module state is process-global.** Anything created at import time or memoised in a module-level or class-level variable lives for the whole run. A payroll importer that caches its rate table on first use serves that same object to every later test. 2. **Modules are imported once.** `sys.modules` caches them, so a later module doing `import payroll.rates` gets the already-warmed object, not a fresh one. "Just re-import it" is not a fix. 3. **Order is deterministic but fragile.** Directory entries are walked in sorted order and method names within a `unittest.TestCase` are sorted too, so the run is repeatable on a given tree — but adding a file changes the order, which is exactly why the failure appeared on an unrelated commit. The usual culprits, roughly in order of frequency: a memoised lookup or connection pool at module scope; monkeypatching that is never undone; a class attribute mutated by an instance; environment variables set for one test; the working directory changed and not restored; and a logging or warnings filter installed once. ### The bisect, concretely Take a payroll CSV import whose loader resolves pay rates across a 17-service dependency graph and caches the resolved table on first import. - **Confirm the baseline.** `python -m unittest tests.test_csv_import.CsvImportTests.test_uses_current_rate` on its own passes. Good — the test is not simply wrong. - **Reproduce the pair.** Test names are positional and honoured in order, so `python -m unittest tests.test_rate_cache tests.test_csv_import -v` runs a candidate polluter followed by the victim. If it fails, you have a two-line reproduction instead of a whole suite. - **Prove it is order, not content.** Swap the two names. A failure that moves or disappears when the order changes is pollution by definition. - **Narrow within the module.** `-k` filters by name and may be repeated, so you can whittle the polluter down to a single method. - **Stop at the first failure.** `-f` ends the run on the first failure or error, which keeps output short and prevents a cascade of secondary failures from obscuring the first real one. - **Silence the passing tests.** `-b` buffers stdout and stderr per test and prints the buffer only for a test that fails or errors. That is doubly useful here: not only is the noise gone, but any output attributed to the failing test that was actually *produced* by setup done earlier stands out. - **Read the ids.** `-v` prints each test's full dotted name, which is both the running order and the exact string to feed back into the command. ### Verifying the hypothesis Once you suspect a specific cached object, inspect it rather than guessing: assert its identity or its contents at the top of the failing test and watch it already be populated. If the object is created at import time, the value will be there before any of your setup ran, which is conclusive. ### Fixing it, and what not to do The right fix removes the sharing. Build the expensive thing inside the test and dispose of it afterwards, or expose an explicit reset the test can call, or inject the cache rather than reaching for a module global. If the object is genuinely expensive, a per-class or per-module fixture with a matching teardown keeps the cost down without leaking across modules. What not to do: rename files to force a convenient order, or pin the run to a sequence that happens to pass. That converts a real defect into a booby trap that fires the next time someone adds a test — and the standard library runner offers no supported ordering knob to lean on anyway. Equally, do not reach for one process per test file as the primary answer; it hides the coupling rather than removing it, slows the suite, and the shared state is usually still wrong in production code that runs the same way. ### What to say in an interview Lead with the mechanism — one process, one `sys.modules`, shared module state — then the two-command reproduction, then the bisect flags, then the fix that removes the sharing. Naming `-f`, `-b` and `-k` in the right roles shows you have actually worked a flaky suite rather than read about one.
- What does `--buffer` actually change about a unittest run?It captures stdout and stderr for the duration of each test and discards the buffer when the test passes, replaying it into the report only when the test fails or errors. Output stops interleaving across tests, so what you read under a failure genuinely belongs to it — and output that appears without a matching test is a strong hint it came from import time or an earlier test's leftovers.
- In what order does a discovery run execute tests?Deterministically but implicitly: directory entries are walked in sorted order, and within a TestCase the loader sorts method names. There is no supported randomisation or ordering option in the standard library runner. That determinism makes a reproduction reliable, but it also means adding or renaming a file can reshuffle the run and expose a dependency that was always there.
- Would running each test module in its own process be a reasonable fix?As a workaround it can unblock a release, but it is not a fix. It masks the coupling instead of removing it, multiplies interpreter start-up across the suite, and leaves the shared state wrong in the production code path that behaves the same way. Use it as a temporary containment measure while the leaking state is made per-test or injected.
It is a shared kitchen: every test cooks in the same room, and the one that fails is simply the first to find an ingredient someone else already used.
saying these in an interview costs you the question
- Calling the test flaky and adding a retry
- Assuming each test module runs in its own fresh process
- Reaching for a rename to force a convenient ordering
- Believing re-importing a module gives fresh module state
- Running the whole suite repeatedly instead of reproducing the pair
- Blaming the failing test without checking what ran before it