Under what circumstances would you implement a custom JUnit Platform TestEngine, rather than solving the problem inside an existing testing framework's programming model?
answer
- New kind of test → engine; new behaviour → extension
- Own discover, execute, UniqueId stability, TestSource
- Cucumber/ArchUnit/jqwik/Spock ship engines
- Try @TestFactory dynamic tests first
- Library-shaped investment, not app-shaped
basics
~20 sWrite an engine only when your tests are not Java methods with annotations — for example specs in files, generated cases, or a different execution model. If tests are still ordinary annotated methods, an extension or a parameterised source is the right level and far cheaper.
solid answer
~50 sAn engine is justified when the *nature of a test* differs from what existing engines model — feature files, spec DSLs, generated or externally-defined cases, or a lifecycle that annotation-driven engines cannot express. That is why Cucumber, ArchUnit, Spock and jqwik ship engines: their tests are not simply annotated Java methods. What you gain: your tests get first-class IDE and CI support, stable unique ids, tags, filtering, and inclusion in the same merged `TestPlan` as everyone else's, with no tool integration work. What you own forever: `discover` and `execute`, a stable `UniqueId` scheme, sensible `TestSource` so IDEs can navigate, correct started/skipped/finished event pairing, parallelism and lifecycle semantics, and compatibility with Platform releases. So the test is: could this be a Jupiter extension, a `@ParameterizedTest` argument source, or a `@TestFactory` returning dynamic tests? If yes, do that. Only when the answer is genuinely no — a new *kind* of test, usually shipped as a reusable library — is an engine warranted.
code
java · 10 linesclass FixtureDrivenTest {
@TestFactory
Stream<DynamicTest> everyFixtureFile() throws IOException {
return Files.list(Path.of("src/test/resources/fixtures"))
.filter(p -> p.toString().endsWith(".json"))
.map(p -> DynamicTest.dynamicTest(p.getFileName().toString(),
() -> assertTrue(Validator.isValid(Files.readString(p)))));
}
}go deeper
It's enough to know engines are how frameworks plug into JUnit, and that you almost never write one.
Contrast an engine with extensions, dynamic tests and parameterised sources, and name real engines like Cucumber or ArchUnit.
Detail the ownership burden — discovery, unique-id stability, test sources, event pairing, parallelism, version maintenance — and the escalation path before reaching for one.
Make it an explicit build-versus-compose judgement: reusable-library scope, ecosystem benefit, long-term maintenance, and a review stance that rejects engines written to dodge the extension model.
## The decision in one line Ask: *is the thing I want to run a new kind of test, or the same kind of test with different behaviour?* New kind → engine. Same kind, different behaviour → extension, dynamic tests, or a parameterised source. ## What an engine costs you Implementing `TestEngine` means owning: **Discovery.** Interpreting whatever selectors make sense for you — `selectPackage`, `selectClass`, `selectFile`, `selectUri`, `selectUniqueId` — plus discovery filters, and returning a `TestDescriptor` tree. Getting `selectUniqueId` right is what makes "rerun this one test" work in an IDE, and it is easy to get wrong. **Identity.** A `UniqueId` scheme that is stable across runs and machines, and that survives reordering and re-generation. Unstable ids break rerun, history, quarantine and flaky-test tracking. If your cases are generated, you need a deterministic naming scheme. **Navigation.** `TestSource` (class, method, file position) so IDEs can jump from a failure to the right line. Without it your engine feels second-class. **Execution semantics.** Correctly paired `executionStarted`/`executionFinished` events for containers and tests, `executionSkipped` for skips, `TestExecutionResult` distinguishing FAILED from ABORTED, and reasonable behaviour when a container setup fails. Also: does your engine support parallelism? Nested containers? Reporting entries? **Compatibility.** The engine SPI is stable but not frozen, and you carry the maintenance across Platform versions. The Platform ships a test kit for exercising engine behaviour, which is a strong hint that engines are non-trivial enough to need dedicated testing. ## What an engine buys you **Universal tooling for free.** Register via `ServiceLoader` and every IDE and build tool that speaks the Platform can discover, display, filter, rerun and report your tests. That is the whole reason the SPI exists — before it, a new framework had to masquerade as a JUnit 4 `Runner`. **Coexistence.** Your engine's tests join the same merged `TestPlan` as Jupiter and vintage tests: one run, one report, one coverage number, tags and engine filters applying uniformly. **Semantic honesty.** If your tests genuinely are not methods — scenarios in a `.feature` file, architecture rules over a class graph, property-based generators — modelling them as methods forces awkward compromises in naming, reporting and failure attribution. ## When the answer is *not* an engine Most in-house needs are better served one level up: - **Cross-cutting behaviour** (timing, retry-ish reporting, resource injection, conditional disabling) → a Jupiter `Extension`. - **Many cases from data** (a CSV, a directory of JSON fixtures, a database of examples) → `@ParameterizedTest` with a custom `ArgumentsProvider`. - **Cases known only at runtime** (walk a directory, build a test per file) → `@TestFactory` returning `DynamicTest`s. This covers a surprising share of "I need my own engine" cases while inheriting Jupiter's reporting. - **Selecting or slicing runs** (change-based selection, quarantine) → the Launcher API with selectors and filters, not a new engine. - **Custom reporting across everything** → a `TestExecutionListener`, which already sees all engines. A good rule: if the answer to "who else would use this engine?" is "only us, in one repo", it is almost certainly the wrong level. Engines are a *library-shaped* investment; extensions and dynamic tests are an *application-shaped* one. ## Legitimate examples - **Cucumber** — tests are scenarios in feature files, discovered from resources, not classes. - **ArchUnit** — tests are rules evaluated over a class graph; results are violations, not method outcomes. - **jqwik / property frameworks** — the unit is a property with generated inputs and shrinking, a distinct execution model. - **Spock, Kotest** — whole alternative specification languages. - **The suite engine** — even JUnit's own `@Suite` support is an engine, because a suite is a container of other engines' tests. Notice the pattern: each is a reusable product with a distinct notion of what a test *is*. ## How I'd decide in a review 1. Write the tests as `@TestFactory` dynamic tests first. If reporting, navigation and rerun are acceptable, stop. 2. If the friction is only cross-cutting behaviour, write an extension. 3. Only if the model itself is wrong — identity, sources, containers, or lifecycle cannot be expressed — consider an engine, and then budget for unique-id stability, test sources, parallelism semantics and version maintenance as first-class work, with engine-level tests of its own. The failure mode I actively guard against is an engine written to avoid learning the extension model. It ships quickly, then quietly degrades everyone's IDE experience and rerun tooling for years.
- Why is UniqueId stability the hardest part of writing an engine?Every downstream capability keys off it: rerun-a-single-test, historical result tracking, flaky-test quarantine, and IDE selection all use `selectUniqueId`. If ids shift when cases are regenerated, reordered, or discovered on a different machine, those features silently break and users blame the tooling. Deterministic, content-derived naming is essential, and it constrains how you generate cases.
- Your team wants a test per JSON fixture in a directory, discovered at runtime. Engine or not?Not an engine. `@TestFactory` returning `DynamicTest`s covers it: cases are created at execution time, each gets its own display name and reporting, and you inherit Jupiter's extension model, tags and IDE integration for free. The main limitation is that dynamic tests are registered during execution, so they don't appear in a pre-run discovery listing — acceptable for most fixture-driven suites.
- What would make you reject an engine proposal in review?If the tests are still ordinary Java methods and the motivation is cross-cutting behaviour or setup, that's an extension. If the engine would live in exactly one repository with one consumer, the maintenance cost — id stability, test sources, parallelism, Platform upgrades — outweighs the benefit. I'd also reject it if no plan exists for testing the engine itself.
Writing an engine is publishing a new file format plus its reader; writing an extension is adding a feature to an existing editor. Only do the former when your content genuinely isn't the existing format.
saying these in an interview costs you the question
- Proposing a custom engine to avoid learning the Jupiter extension model
- Assuming an engine automatically gets IDE navigation without providing TestSource
- Ignoring UniqueId stability, breaking rerun and history
- Thinking an engine is needed for runtime-generated cases (@TestFactory covers it)
- Believing a custom engine replaces or disables the Jupiter engine