Your team wants to turn a directory of 400 JSON fixture files into individual JUnit 5 test cases generated at runtime, one per file. What operational trade-offs would you weigh before committing to that, and how would you keep the suite maintainable?
answer
- Test set defined in data vs in code
- No per-case @Tag / @Disabled / filter selection
- Positional IDs break rerun-failed + flaky history
- Isolation is your job — wrapper with try/finally
- Count guard + static core so "generated nothing" fails loud
basics
~20 sYou gain a case per file with zero code churn as fixtures are added. You lose discovery-time visibility: no per-case tags or disabling, no name-based selection, positional unique IDs that break rerun-failed and flaky history, and no framework-managed per-case setup. Keep it by naming cases after files, sorting, supplying source URIs, and owning isolation explicitly.
solid answer
~1 min**The gain is real:** the test set becomes a function of the data. Drop a fixture in, get a case, no code review of a new test method; delete a fixture, the case disappears. For contract, golden-file and spec-driven suites that is a large maintenance win. **The costs are all downstream of the fact that generated cases do not exist at discovery time:** - No per-case `@Tag`, `@Disabled` or `@Timeout` — quarantining one bad fixture means filtering it in code or moving the file. - Filters and "run this test" select the *generating method*, all-or-nothing. - Identity is positional (`#1`, `#2`, …), so adding a fixture renumbers later cases and any cross-build history — flaky-test tracking, per-test timings, rerun-failed by ID — misattributes. - Jupiter's per-test lifecycle binds to the generator, so isolation between cases is your code's job, not the framework's. - 400 cases from one method means one method's failure aborts the lot. **How I keep it maintainable:** name each case after its file and sort the listing so ordering is stable; supply a `testSourceUri` so failures navigate to the fixture; group with containers mirroring the directory tree; give every case its own fixture inside the executable with `try/finally`; keep a small hand-written `@Test` set for the critical paths so a generator bug cannot silently produce zero cases; and assert the expected case count.
code
java · 17 lines@TestFactory
Stream<DynamicNode> fixturesParse() throws IOException {
List<Path> files = Files.walk(FIXTURES).filter(Files::isRegularFile).sorted().toList();
assertTrue(files.size() >= EXPECTED_MINIMUM, "fixture discovery produced too few files");
return files.stream()
.filter(this::notExcluded)
.map(file -> dynamicTest(FIXTURES.relativize(file).toString(), file.toUri(),
() -> {
Fixture f = Fixture.create();
try {
assertNotNull(f.parser().parse(Files.readString(file)));
} finally {
f.close();
}
}));
}go deeper
Say that generating gives one case per file with no code churn, but you cannot disable or select an individual case.
List the concrete losses — no per-case annotations, filters hit the generating method, no per-case lifecycle — and the naming/grouping practices that keep the report readable.
Add the operational failures: positional IDs versus rerun-failed and flaky history, blast radius of a generator exception, quarantine mechanics, and the count guard.
Frame it as choosing where the test set is defined and what addressability the pipeline needs, then state the engineering investment that makes the trade acceptable and the conditions under which you would refuse it.
## What you are really choosing The decision is *where the test set is defined*: in code, visible at discovery time, or in data, materialised at runtime. Everything else follows. Static `@Test` methods are known before execution, so the whole platform ecosystem — filters, tags, IDE selection, rerun-failed, flaky-test dashboards, per-test timing history, quarantine lists — can address them by stable name. Generated cases are invisible until the generator runs, so none of that machinery can name them individually. ## The case for generating For 400 fixture files the win is genuine and worth stating plainly: - **Zero-friction growth.** A new fixture is a new test with no code change and no reviewer deciding where the method goes. - **No copy-paste drift.** One body, 400 inputs; a change to the assertion logic happens once. - **The suite cannot fall behind the data.** With hand-written tests, someone eventually adds a fixture and forgets the test; with generation that is impossible. - **Independent reporting per fixture.** One broken fixture is one red case, not a whole method aborting at the first bad file. This is the classic sweet spot: golden-file/approval tests, contract tests over a spec, conformance suites, "every registered implementation must satisfy X". ## The costs, and how each one bites **Quarantine is awkward.** A single fixture starts failing for an unrelated reason and you want it off the critical path for a day. With static tests: `@Disabled("JIRA-123")`. With generated cases: you must encode an exclusion list in the generator, or move the file — both of which are code/data changes rather than an annotation, and both of which are easy to forget to revert. Build the exclusion mechanism deliberately (a `skip.txt`, or a naming convention like a `_wip` prefix) rather than improvising. **Selection is all-or-nothing.** A developer debugging one fixture cannot ask the runner for that one case in a fresh process; they can only run the generating method, which runs 400. The mitigation is a system property or environment filter honoured by the generator (`-Dfixture=2024-06-01.json`) — cheap to add, and worth adding on day one. **Identity is positional and therefore unstable.** Generated nodes are keyed by index. Adding a fixture at the front of a sorted listing shifts every subsequent ID. Anything that persists IDs across builds — rerun-only-failed, flaky-test history, per-test duration trends — silently attributes yesterday's data to today's different case. You cannot fully fix this; you can reduce churn by sorting deterministically and by accepting that per-case history is not reliable for this suite. **Isolation is your responsibility.** Jupiter's per-test callbacks bind to the generating method, so all 400 cases share whatever the generator built: one transaction, one mock, one temp directory. Write a wrapper helper that creates and closes per-case state in `try/finally` and route every case through it. Otherwise you get order-dependent failures inside a construct whose report claims 400 independent tests. **Blast radius.** One generator, 400 cases: an exception while generating (unreadable file, bad path, permissions) aborts the container and the remaining cases never exist. A run that reports 40 passing cases instead of 400 looks green-ish at a glance. Guard against this with an explicit count assertion — a small static `@Test` asserting the fixture directory has at least N files, or a generated case that checks the total — so "generated nothing" fails loudly. **Runtime cost visibility.** 400 cases in one method make it hard to see which fixture is slow unless you group and name well; and if the suite is sharded, a single generating method is one indivisible unit of work that cannot be split across agents. ## The maintainable shape 1. **Deterministic source order** — sort the directory listing so the tree and numbering are reproducible. 2. **Names from the data** — the fixture file name is the display name; never `case #12`. 3. **Containers mirroring directories** — group by subdirectory so a failure reads `2024-06 → invoice-negative-total.json`. 4. **`testSourceUri` per case** — a `file:` URI so clicking a failure opens the fixture. 5. **Per-case isolation helper** — create/close state inside the executable with `try/finally`. 6. **A selection escape hatch** — a system property that narrows the generator to one file. 7. **An explicit exclusion mechanism** — documented, greppable, with an issue reference. 8. **A count guard** — fail if the generator produced fewer cases than expected. 9. **A small static core** — a handful of hand-written tests for the critical paths, so the suite has a floor that does not depend on the generator working. ## When I would not generate If the cases need genuinely different setup per case, injected parameters, per-case transactional rollback, or per-case quarantine and history — the framework-managed lifecycle is worth more than the runtime flexibility, and an annotation-driven, discovery-time construct is the better tool. If the input set is small and stable (five cases that rarely change), hand-written tests read better and cost nothing. ## The summary I would give Generate when the *set of cases* is data that changes independently of the code, and accept that you are trading discovery-time addressability for it. Then invest the small amount of engineering — naming, grouping, source URIs, isolation helper, count guard, selection hatch — that turns 400 anonymous nodes into a suite people can actually operate.
- One of the 400 fixtures starts failing and you need it off the critical path today. What do you do?There is no per-case @Disabled, so the exclusion has to live in the generator or the data: an explicit skip list file read by the generator, or a naming convention the filter honours. Whichever you pick, make it greppable and require an issue reference next to the entry, because the usual failure mode is a quarantine that nobody ever removes. Reporting the excluded fixtures as skipped cases, rather than omitting them silently, keeps the report honest.
- Your CI dashboard tracks per-test flakiness by unique ID. Why does that data go wrong for this suite?Generated nodes are identified positionally under the generating method, so adding or removing a fixture renumbers every later case. Yesterday's [dynamic-test:#57] is today's different fixture, and the flakiness history is attributed to the wrong case. Sorting the source keeps numbering reproducible within a fixed set but does not survive insertions, so per-case history for generated suites should be treated as unreliable.
Hand-written tests are a printed roster; generated cases are a turnstile count. The turnstile scales effortlessly but you cannot pull one person out of the queue by name.
saying these in an interview costs you the question
- Claiming generated cases can be tagged or disabled individually.
- Assuming rerun-failed and flaky-test history work per generated case across builds.
- Relying on the framework to isolate generated cases from one another.
- Not noticing that a generator producing zero cases still reports a green run.
- Presenting runtime generation as strictly better than declared tests, with no account of lost addressability.