Why can a unittest TestCase.subTest loop leak state between its cases?
answer
- same instance, same process, one method
- reporting boundary versus fixture boundary
- setUp does not re-run per case
- passes alone, fails in sequence
- generate real methods for real isolation
basics
~20 sBecause subTest is a reporting boundary, not an isolation boundary: setUp, tearDown and cleanup callbacks bracket the whole test method, so anything a case mutates — instance attributes, module globals, caches, the filesystem — is still there for the next case.
solid answer
~40 s`subTest` splits the *report* into one entry per case; it does not split the *fixture*. `setUp` runs once before the method and `tearDown` once after it, `addCleanup` callbacks fire when the method ends, and every case in the loop shares one `TestCase` instance, one set of module globals and one process. So a case that mutates `self`, warms a module-level cache, writes a file or patches a global runs the following cases against state its predecessors built — and a case that fails part-way leaves that state half-mutated. The tell is a suite where each case passes alone but the loop fails, or where reordering the table changes the verdict. Fix it by building per-case state inside the `with` block, or by generating one real test method per row when true isolation matters.
code
python · 19 linesimport io
import unittest
class SyncRowTest(unittest.TestCase):
def setUp(self):
self.seen = []
def test_rows(self):
for row in ("1.5", "1,5", "1.5"):
with self.subTest(row=row):
self.assertNotIn(row, self.seen)
self.seen.append(row)
result = unittest.TextTestRunner(stream=io.StringIO()).run(
unittest.defaultTestLoader.loadTestsFromTestCase(SyncRowTest)
)
print(len(result.failures), "failure(s); setUp ran once for all three rows")go deeper
Remember one fact: setUp runs once per test method, not once per subtest. If a case changes something, the next case sees the change.
Explain the mechanics — one TestCase instance, one setUp, cleanups deferred to the end of the method — and show the per-case with block or explicit reset that restores independence.
Demonstrate the diagnosis on a real order-dependent failure: run alone, reorder, bisect to the failing pair, then name the shared resource. Say plainly that the traceback points at the victim, not the cause.
Own the suite-wide policy: when a table of rows is allowed to share a method and when the team generates real methods, and how that choice affects CI localisation, parallel execution and the cost of a night-time page.
## Two different boundaries There are two things people mean by "each case is independent": each case gets its own **report entry**, and each case gets its own **fresh world**. `unittest.TestCase.subTest` provides the first and provides nothing at all of the second. Everything below follows from that one sentence. Concretely, for a method that loops over twenty rows inside `subTest`: - `setUp` runs **once**, before the loop, and `tearDown` **once**, after it; - `addCleanup` callbacks registered anywhere in the method fire **once**, after the method completes; - all twenty rows execute against the **same** `TestCase` instance, so anything assigned to `self` persists across rows; - module globals, caches, patched attributes, temporary directories, open connections and clock fakes are shared, because it is all one method in one process. Compare a real test method per row: unittest constructs a fresh `TestCase` instance for every method it runs and brackets each with `setUp`/`tearDown`. *That* is the isolation mechanism; `subTest` is not a substitute for it. ## What the leak looks like in practice Take a scheduled inventory sync between two systems, whose quantity parser must accept several locale-dependent number formats. The parser memoises the format it resolved per source system in a module-level dict — a warm cache measured at an 83% hit rate in production, which is exactly why it exists. The test loops the parser over a table of rows, one per locale, each wrapped in `subTest`. Row one, a plain `1.5`, resolves and caches the dot format. Row two, `1,5` from the other system, hits the warm entry instead of resolving its own, parses to `15.0`, and fails. Every symptom now points at row two's data — but row two passes perfectly when run alone, and passes when the table is reordered to put it first. The defect the test is actually reporting belongs to the *pair*, and the loop, not the parser, is what made the pair possible. The same shape appears with `self.seen`, with a `tempfile` directory reused across rows, with a fake clock advanced by an earlier row, and with any global that a case patches and only restores at the end. ## Diagnosing it The signature is order dependence, and the cheap probes are: 1. **Run the row alone.** Slice the table to one row and rerun the method. Passing alone plus failing in the loop is a leak, not a data bug. 2. **Reorder the table.** If the verdict follows the position rather than the row, state is carried between cases. 3. **Bisect the table.** Halve it until the smallest failing pair is left; the pair names the shared resource. 4. **Assert the invariant, not just the result.** Add an assertion inside the `with` block that the shared thing is in its expected pre-state — the failure then points at the leak instead of at the row. ## Fixing it In rough order of cost: - **Build the per-case state inside the `with` block** and tear it down there — a context manager entered per row, a cache cleared per row, a fresh temporary directory per row. The block is the only per-case scope you have, so use it. - **Keep the case body a pure function of its row.** If the body reads nothing but the row and the code under test, there is nothing to leak. - **Reset the shared resource explicitly at the top of each case** when the code under test insists on a module-level cache. Ugly, honest, and greppable. - **Generate one real test method per row** when isolation genuinely matters — build the functions in a loop and attach them to the class. You regain `setUp`/`tearDown` per case, individually runnable and skippable ids, and parallel-runner friendliness; you pay with a factory the next reader must decode, and names that do not appear literally in the source. ## The judgement call `subTest` is the right default for a table of pure input-to-expected-output rows, which is most tables. The moment a row mutates anything that outlives it, you are choosing between disciplined per-case setup inside the block and generated methods — and the deciding question is not aesthetics but whether a failure in CI will be *localisable* by someone who did not write the test. An order-dependent loop is a test that lies about which row is broken, and a test that lies costs more than the coverage it bought.
- How do you generate one real test method per row without a third-party runner?Build a closure per row in a factory and attach it to the class with `setattr`, naming it `test_<row id>`; a class decorator or `__init_subclass__` can do the same more tidily. The loader then sees real methods, so each gets its own instance, its own `setUp`/`tearDown`, its own id for selection, and its own skip decision. The cost is indirection: the names do not appear literally in the source.
- A case passes alone but fails inside the loop. What is your first move?Reorder or reverse the table and rerun. If the verdict follows position rather than data, it is carried state, and the next step is bisecting to the smallest failing pair — that pair names the shared resource. Only after that do I look at the row's data, because the traceback in an order-dependent loop points at the victim, not the cause.
- Does addCleanup give you per-case teardown inside a subTest loop?No. Cleanups registered during the method run when the *method* finishes, in reverse order, regardless of which subtest registered them — so a callback registered in row one still holds until every row has finished. For per-case teardown, use an ordinary `with` block or `try`/`finally` inside the loop body.
subTest is like separate line items on one invoice: the itemisation tells you which charge is wrong, but it is still a single account, and money one line moved is money the next line sees.
saying these in an interview costs you the question
- Says setUp and tearDown run around each subtest
- Treats subTest as equivalent to separate test methods
- Blames the failing row's data without checking order
- Expects addCleanup to fire between cases
- Assumes a failed case rolls back what it mutated
- Thinks reordering a case table should never change results