skip to content

When would you choose runtime-generated JUnit 5 tests from a @TestFactory over statically declared test methods in a large suite, and what does that choice cost you in tooling and CI?

level: principalimportance: nice to knowfreq 18%

answer

  1. justified only when the case set is runtime-derived
  2. no discovery-time plan → no pre-run selection
  3. lifecycle wraps the factory, not the cases
  4. index-based IDs break rerun/flaky history
  5. empty input = green build with zero tests

basics

~20 s

Choose a factory only when the set of cases is unknowable until runtime — files on disk, registered implementations, data fetched by the build. The costs are no pre-run visibility, no per-case lifecycle, index-based identities that break rerun history, and a green build when the input is empty.

solid answer

~60 s

My rule: use `@TestFactory` when **the set of cases, not just their arguments, is discovered at runtime** — one contract test per implementation found through a service registry, one case per fixture file in a directory, one case per entry in a spec the build downloads. In those cases the alternative is a hand-maintained list that silently drifts from reality, which is worse. What it costs at suite scale: - **Visibility.** The cases do not exist in the discovery-phase plan, so IDEs cannot list or select them beforehand and pre-run filters cannot target one. - **Lifecycle.** Callbacks wrap the factory, not the cases, so isolation between cases is your responsibility. - **Identity.** Dynamic case IDs are index-based; inserting a case shifts later ones, breaking rerun-failed and flaky-test history keyed on identifiers. - **Volume and reporting.** A factory can emit thousands of cases and swamp report renderers; grouping and deliberate sampling become design work. - **Silent zero.** Empty input means zero tests and a green build unless you assert otherwise. So I gate it: runtime-derived case sets yes; looping over a checked-in list, no.

go deeper

for a junior

Say that factories are for cases you cannot know when writing the code, and that ordinary tests are the default.

for a middle

Give concrete runtime-derived examples and name at least the lost pre-run visibility and per-case lifecycle.

for a senior

Enumerate the broken tooling contracts — discovery plan, lifecycle, identity, reporting volume — and the empty-input hazard with its guard.

for a principal

Deliver it as policy: the narrow justification, the operational guardrails (non-empty assertions, test-count regression checks, naming rules, grouping/sampling), and why the real risk is silent loss of coverage.

## The decision rule Ask one question: *could I have written these test methods by hand at authoring time and been correct?* - **Yes** — write them, or use a statically declared data-driven mechanism. You keep the full tooling contract. - **No, because the cases come from the environment** — a factory is the right tool, and the maintenance win is real. The strongest cases for dynamic generation share one property: **the list would otherwise be hand-maintained and would silently rot**. - One contract test per implementation discovered through a service registry — add an implementation, it is tested automatically; forget to register it and the registry test catches that separately. - One case per fixture file in a directory — a new fixture becomes a case with no code change, which is exactly what you want from a corpus of regression samples. - One case per entry in a specification the build fetches, so the suite tracks the spec rather than a copy of it. - A matrix over two runtime-derived collections, where the cross product cannot be enumerated in source. ## What you trade away **Pre-run visibility.** Discovery builds the test plan before execution; dynamic cases are created during execution. Practical effects: the IDE shows a single factory node until you run it; you cannot click one generated case to run it in isolation; pre-run filters and unique-ID selection cannot target a case that does not yet exist. Debugging a single failing case means running the whole factory, often with a temporary filter inside the factory itself. **Lifecycle and extensions.** `@BeforeEach`/`@AfterEach` and their extension equivalents wrap the factory method, not each case. All cases share one instance and one fixture, so isolation is hand-rolled inside each case body. Teams that adopt factories broadly reinvent a small per-case setup helper — a sign the mechanism is being pushed past its intent. **Identity stability.** Dynamic cases are identified positionally within the factory. Insert a case in the middle and every later case's identifier shifts. Anything keyed on identifiers — rerun-failed lists, flaky-test dashboards, quarantine lists, historical duration tracking — mis-attributes results. On a suite with flaky-test automation this is a genuine operational cost. **Volume control.** A factory over a large corpus can emit thousands of cases. Lazy consumption keeps memory fine, but reports, dashboards and log scrapers may not cope, and the signal-to-noise of a run degrades. You end up designing grouping (containers), sampling, or a nightly-full/per-commit-subset split — real work that static suites get for free from their fixed size. **Silent zero.** The failure mode that actually bites teams: empty input produces zero tests and a passing build. A fixture directory dropped from the build output, a filter that matches nothing, a registry that fails to load — each quietly deletes a suite. Any factory reading external input needs an explicit non-empty assertion, and ideally a CI check on total test count. **Cognitive load.** A reader of a static test knows exactly what runs. A reader of a factory must mentally execute the generation logic to know what the suite covers, and the answer can differ per environment. Keep generation logic trivial — read, map, name — and push complexity into the assertions. ## Guardrails I would set for a team 1. Factories are for **runtime-derived case sets**; looping over a checked-in list uses static declarations. 2. Every factory reading external input **asserts non-empty** before generating. 3. Display names must **identify the case** (file, ID, key) — never a bare index. 4. Generation logic stays **trivial and side-effect free**; per-case fixtures live inside the case body. 5. CI fails on a **drop in total executed tests**, which catches both empty factories and silently undiscovered classes. 6. Large corpora are **grouped into containers** and, if necessary, split into a fast subset and a full nightly run. ## How to answer this well A principal-level answer does not say "dynamic tests are powerful". It states the narrow condition that justifies them, names the specific tooling contracts that break (discovery-time plan, per-case lifecycle, stable identity), and closes with the operational guardrails that make the tradeoff safe — because the risk here is not a wrong test, it is a suite that quietly stops testing anything.

  • Your team wants to convert a long list of similar static test methods into one @TestFactory. What do you push back with?
    That a checked-in list of cases is not runtime-derived, so the conversion buys concision and loses the discovery-time plan, per-case lifecycle and stable identities. Statically declared data-driven tests give the same concision while keeping individual selectability and per-case isolation. I would only support the factory if the case set genuinely comes from the environment rather than from source.
  • How do you keep flaky-test tracking usable for a suite that generates cases dynamically?
    Do not key tracking on the dynamic case identifiers, since they are positional and shift whenever the input changes. Track at the factory-container level, or have the generation order be deterministic and derive a stable key from the record itself, exporting it via the display name or a reported entry. Otherwise quarantine and history automation will mis-attribute results after any data change.

Static tests are a printed checklist; a factory is a checklist generated from whatever is in the warehouse today — invaluable when stock changes, dangerous when the warehouse door is locked and the list comes back empty.

saying these in an interview costs you the question

  • Reaching for @TestFactory to loop over a hard-coded list that could be static test methods
  • Claiming dynamic tests are equivalent to static ones with no tooling tradeoffs
  • Ignoring that empty input yields a passing build with zero tests
  • Building rerun-failed or quarantine automation on positional dynamic-case identifiers
  • Putting substantial logic and side effects into the generation code, making suite coverage unpredictable

context