skip to content

You need one JUnit 5 test case per record in a data file that is read at runtime, generated by a @TestFactory method. How do you build it, and what pitfalls do you guard against?

level: seniorimportance: should knowfreq 25%

answer

  1. classpath resource, not relative path
  2. assert non-empty before mapping
  3. display name carries the record identity
  4. parse inside the case, not the factory
  5. JUnit closes the stream — no try-with-resources

basics

~20 s

Read the file into a stream, map each record to a dynamic test with a descriptive display name, and return the stream — JUnit closes it and consumes it lazily. Guard against an empty file silently producing zero tests, non-descriptive case names, shared state across cases, and reading from the source tree instead of the classpath.

solid answer

~60 s

Shape: open the file lazily (`Files.lines`, a parser producing a `Stream`), `map` each record to a dynamic test whose **display name identifies the record** (`"row 14: invoice-2023-11"`), and return the stream. JUnit consumes it lazily and closes it, so no try-with-resources — closing it yourself hands JUnit a closed stream. Pitfalls I actively guard: 1. **Silent emptiness.** No records → zero tests → green build with no coverage. Assert the record list is non-empty *before* mapping, so the factory itself fails loudly. 2. **Resource location.** Read via the classpath/test resources, not a relative path into the source tree, or it works locally and finds nothing in CI or from a different working directory. 3. **Naming.** Failures report the display name only; "case 7" is useless. Include the identifying field. 4. **Shared state.** All cases run inside one `@BeforeEach`/`@AfterEach` pair on one instance, so per-case setup belongs inside each case's body. 5. **Parse failures.** A malformed record should fail *that* case, not abort the whole factory — parse inside the case body when you want per-record failure isolation. And I sanity-check the alternative: if the file is fixed and checked in, a statically declared data-driven test keeps discovery-time visibility.

code

java · 19 lines
java
@TestFactory
Stream<DynamicTest> everyFixtureRecordValidates() throws Exception {
    List<String> lines;
    try (InputStream in = getClass().getResourceAsStream("/records.csv")) {
        assertNotNull(in, "records.csv missing from test resources");
        lines = new BufferedReader(new InputStreamReader(in, UTF_8))
                .lines().toList();
    }
    assertFalse(lines.isEmpty(), "records.csv contained no rows");

    return IntStream.range(0, lines.size()).mapToObj(i -> {
        String line = lines.get(i);
        return DynamicTest.dynamicTest("line " + (i + 1) + ": " + line, () -> {
            Record record = Record.parse(line);   // per-case parse
            Validator validator = new Validator(); // per-case fixture
            assertTrue(validator.isValid(record));
        });
    });
}

go deeper

for a junior

Show the basic shape: read the records, map each to a dynamic test with a meaningful name, return the stream.

for a middle

Add classpath loading, the stream-closing guarantee, and per-record display names that identify the failing row.

for a senior

Lead with the failure modes — silent empty input, working-directory-dependent paths, factory-level parsing aborting everything, shared state — and the guards for each.

for a principal

Also address scale and reporting (containers, sampling, sharding) and challenge whether runtime generation is genuinely required versus statically declared data-driven tests.

## The canonical shape ```java @TestFactory Stream<DynamicTest> everyRecordIsValid() throws IOException { Path file = Path.of(getClass().getResource("/records.csv").toURI()); List<String> lines = Files.readAllLines(file); assertFalse(lines.isEmpty(), "no records found — fixture missing?"); return IntStream.range(0, lines.size()) .mapToObj(i -> DynamicTest.dynamicTest( "line " + (i + 1) + ": " + lines.get(i), () -> assertTrue(validator.isValid(parse(lines.get(i)))))); } ``` Every decision in that snippet is deliberate. Below is why. ## Loading the data Read through the **classpath**, not a relative filesystem path. A relative path resolves against the JVM working directory, which differs between an IDE run, a Gradle/Maven run, and a CI container — the classic "works on my machine, finds nothing in CI" failure. Test resources are copied into the build output, so `getClass().getResourceAsStream("/records.csv")` is stable everywhere. If you stream lazily (`Files.lines`, `Files.list`, `Files.walk`), return that stream directly: JUnit closes it once the generated cases finish. Do **not** wrap it in try-with-resources — that closes it before JUnit consumes it and the run fails. ## Guarding against zero cases The single most dangerous property of data-driven dynamic tests is that **no data means no tests and a green build**. A misspelled resource name, a fixture that stopped being copied by the build, an over-strict filter — all of them silently delete an entire test suite while the pipeline stays green. Defences, in order of directness: - assert non-emptiness in the factory before mapping (fails the container loudly); - emit a deliberately failing dynamic case when the input is empty; - monitor the executed test count in CI and fail on a drop. At least one of these belongs in any factory whose input comes from outside the code. ## Naming the cases A generated case's identity in reports is its display name. Index-only names ("test 3") force whoever sees the CI failure to re-derive which record broke. Include the discriminating value — a filename, an ID, a line number — and keep it short enough for report tooling. Line numbers plus the key field is a good default: it survives reordering better than an index alone, and it points straight at the file. Be aware the unique ID of a dynamic case is index-based, so inserting a record shifts the IDs of later ones. Do not build rerun-failed automation on those IDs. ## Where to parse Parsing inside the **factory** means a malformed record blows up the factory and *no* cases run — one bad row costs you the whole suite. Parsing inside each **case body** means a malformed record fails only that case, and everything else still reports. Prefer per-case parsing unless the data must be validated as a whole (e.g. checking for duplicate IDs across records, which is legitimately a factory-level concern — or its own dedicated case). ## State between cases Because `@BeforeEach`/`@AfterEach` wrap the factory rather than each case, all generated cases share one test-class instance and one fixture. With per-record data this bites quickly: a case that writes to a shared collection or leaves a database row behind changes what later cases see. Construct per-case fixtures inside the executable, or reset at the start of each case body so a throwing case cannot poison its successor. ## Scale and laziness Lazy consumption means memory stays bounded even for tens of thousands of records, and early failures are reported before later cases are constructed. But consider the reporting side: ten thousand generated cases can overwhelm a CI report renderer and make the run hard to read. If the record count is large, either group cases under containers to keep the tree navigable, or sample/shard the data deliberately rather than accidentally. ## Is a factory even the right tool? The honest check: is the **set of cases** unknowable at authoring time? A checked-in file with a stable shape is arguably static data, and statically declared data-driven tests keep discovery-time visibility, per-case lifecycle and individual selectability. Dynamic generation earns its cost when the data is discovered at runtime — a directory whose contents vary, a registry of implementations, a spec fetched by the build, records whose count nobody controls. ## Summary Stream lazily from the classpath, name cases by their record, assert the input is non-empty, parse per case, keep per-case state local, and verify that runtime generation is genuinely required rather than convenient.

  • Your data-driven factory reports zero tests and the build is green. How do you make that impossible?
    Fail inside the factory before any mapping: assert that the resource stream is non-null and that the parsed record list is non-empty, with a message naming the expected resource. As a backstop, track the executed test count in CI and fail the build when it drops, which also catches inner classes that quietly stopped being discovered. Both together mean a missing fixture is a red build, not a silent coverage hole.
  • Why parse each record inside the generated case rather than inside the factory?
    Parsing in the factory makes one malformed record an exception during generation, which fails the whole container so no cases run at all. Parsing inside each case body isolates the failure to that record, and every other record still reports its own result. The exception is validation that is inherently about the data set as a whole, such as duplicate-ID detection, which belongs at factory level or in a dedicated case.
  • How would you keep a factory that generates thousands of cases readable in CI reports?
    Group related cases under dynamic containers so the report renders a navigable tree instead of one enormous flat list, and make display names short but identifying. If the volume is genuinely excessive, sample or shard the data deliberately — for example run the full set nightly and a representative subset per commit — rather than letting the report become unusable.

saying these in an interview costs you the question

  • Reading the data file with a relative filesystem path instead of a test-classpath resource
  • Wrapping the returned stream in try-with-resources, handing JUnit a closed stream
  • Treating zero generated tests as an acceptable outcome because the build stayed green
  • Naming cases only by index, so CI failures do not identify the offending record
  • Parsing all records in the factory so one bad row aborts every case

context