You are designing a reusable JUnit 5 extension that gives every integration test a clean database, a stubbed HTTP dependency and diagnostics on failure. How do you decide which callback interfaces to implement, and what failure modes do you design against?
answer
- scope first: JVM-wide → BeforeAll; per-test → BeforeEach; body-only → TestExecution
- singleton container via root context store + closeable resource
- diagnostics = capture + rethrow; add LifecycleMethodExecutionExceptionHandler
- teardown must survive failed setup
- PER_CLASS sharing, parallel execution, per-test cost
basics
~20 sMap each concern to lifecycle scope: expensive shared resources in BeforeAll/AfterAll callbacks, per-test isolation in BeforeEach/AfterEach callbacks, diagnostics via the exception handler plus AfterTestExecution. Design for teardown that always runs, deterministic ordering, and honest reporting.
solid answer
~50 sDecide by **scope and position**, then by failure behaviour. - **Expensive, shareable** (container, stub server, connection pool) → `BeforeAllCallback`/`AfterAllCallback`, ideally started once per JVM with a store-based guard rather than per class. - **Per-test isolation** (transaction open/rollback, truncate, reset stub mappings) → `BeforeEachCallback`/`AfterEachCallback`. These wrap the test class's own `@BeforeEach`, so its fixtures land inside your transaction. - **Bracketing only the body** (timings, state snapshots) → `BeforeTestExecutionCallback`/`AfterTestExecutionCallback`. - **Diagnostics on failure** → `TestExecutionExceptionHandler` that captures and **rethrows**, plus `LifecycleMethodExecutionExceptionHandler` so setup failures are covered too; `TestWatcher` for pure reporting. - **Injection** of handles into the test → `TestInstancePostProcessor` for fields and `ParameterResolver` for parameters. Failure modes to design against: teardown that assumes setup succeeded; state leaking under `@TestInstance(PER_CLASS)`; nondeterministic ordering when composed with other extensions; swallowing failures; and per-test cost that makes the suite unusable.
go deeper
Match each need to a scope: once-per-class work in BeforeAll callbacks, per-test cleanup in BeforeEach/AfterEach callbacks, diagnostics when a test throws.
Explain why the extension's beforeEach wrapping the class's own @BeforeEach matters for transactional fixtures, and cover injection via post-processor and parameter resolver.
Add the singleton-container pattern with the root context store, exception handling for lifecycle methods, defensive teardown, and per-test cost budgeting.
Own the whole contract: composition via a meta-annotation, ordering guarantees, behaviour under per-class lifecycle and parallel execution, honest reporting, and keeping the extension small enough to reason about.
## Start from scope, not from the interface list Every concern in the brief has a natural lifetime. Write the lifetimes down first, then the interface follows mechanically. **JVM-wide / expensive.** A database container, a stub HTTP server, a Kafka broker. Starting these per class multiplies the suite runtime by the number of classes. Implement `BeforeAllCallback`, but guard it so the resource starts once: put the handle in the root `ExtensionContext`'s store under a namespace, and register a closeable resource so the framework closes it when the root context is closed. This is the standard "singleton container" pattern and it is worth more suite time than any other decision here. **Per test.** Isolation. Transaction begin/rollback, or truncate-and-reseed, plus resetting stub mappings and recorded requests so one test's stubbing cannot satisfy the next test's call. Implement `BeforeEachCallback` and `AfterEachCallback`. The wrapping position matters: your `beforeEach` runs *before* the test class's own `@BeforeEach`, so fixtures the class inserts land inside your transaction and are rolled back with it; your `afterEach` runs *after* theirs, so their cleanup still sees a live connection. **Around the body only.** If you want the duration of the test itself, or a snapshot of the database diff caused strictly by the body, use `BeforeTestExecutionCallback`/`AfterTestExecutionCallback` — the innermost ring, excluding user setup and teardown. **On failure.** Capture container logs, the stub server's recorded requests, and a dump of relevant tables. Implement `TestExecutionExceptionHandler` that captures, publishes a report entry (or writes an artefact path the CI picks up), and **rethrows the original throwable**. Add `LifecycleMethodExecutionExceptionHandler` so failures in `@BeforeEach`/`@AfterEach` — often the most confusing ones — are covered as well. If you only need to *observe* outcomes for metrics, `TestWatcher` is better because it cannot accidentally change the verdict. **Handing things to the test.** Tests will need the JDBC `DataSource`, the stub server's base URL, maybe a data builder. Offer both `TestInstancePostProcessor` (annotated fields) and `ParameterResolver` (test-method parameters) so users pick their style. Keep the contract narrow and annotation-driven, and fail fast with a clear `ExtensionConfigurationException` when the annotation is used somewhere you cannot support. ## Composition Expose one meta-annotation (`@IntegrationTest`) that composes the extensions rather than making every test class list mechanisms. Remember that before-callbacks fire in registration order and after-callbacks in reverse, so if the transaction extension must wrap the stub-server extension, register it first. Do not build hidden dependencies between extensions; if one truly needs another's state, pass it through the extension store under an explicit, documented namespace key rather than relying on incidental ordering. ## Failure modes to design against **Teardown that assumes setup succeeded.** `After*` callbacks still run when setup threw. Your `afterEach` may find a null transaction or a server that never started. Null-check, guard, and never let a teardown NPE mask the real failure — that turns a clear setup error into a confusing secondary one. **Exception masking.** Multiple callback failures are aggregated by JUnit, but a teardown that throws its own exception can still dominate the report. Log-and-continue inside teardown where a partial failure is tolerable, and let the original propagate. **State leaking under the per-class lifecycle.** If a user adds `@TestInstance(Lifecycle.PER_CLASS)`, your `TestInstancePostProcessor` runs once and injected handles are shared for the whole class. Reset mutable state in `BeforeEachCallback` rather than relying on fresh injection, or detect the lifecycle and fail with a clear message if you cannot support it. **Parallel execution.** If the suite enables parallel test execution, a single shared container and shared stub server become contended. Decide explicitly: serialise with resource locks, give each thread its own schema, or namespace stub mappings per test. Silently sharing produces the worst kind of flake. **Dishonest reporting.** Never swallow a failure to keep the build green, and never mark something passed that was tolerated. If a test could not run because infrastructure was unavailable, use a condition to skip it with a reason, so the report says "skipped: database unavailable" rather than "passed". **Cost.** Per-test callbacks run thousands of times. Truncating fifty tables per test can dominate the suite; a transaction rollback is usually orders of magnitude cheaper. Measure the overhead your extension adds per test and treat it as a budget. **Opacity.** Every hidden action makes a test harder to read. Publish what you did (report entries), name the extension after its effect, and document the state a test starts in. A newcomer should be able to answer "what is in the database when my test begins?" from the annotation name and one paragraph of docs. ## What good looks like One annotation on the class. A container started once for the whole build. Each test starting in a known empty state, ending rolled back. Failures accompanied by logs and recorded HTTP traffic, with the original assertion error still on top. No test able to affect another regardless of order. And an extension small enough that its behaviour fits on one page.
- Why prefer transaction rollback over truncating tables for per-test database isolation, and when does rollback fail you?Rollback is dramatically cheaper per test and leaves nothing behind, so it scales to thousands of tests where truncation would dominate the runtime. It fails when the code under test manages its own transactions or commits (new-transaction propagation, background threads, or anything committing outside the test's connection), and when the test must exercise real commit semantics — in those cases you need truncation, a per-thread schema, or a container reset.
- The suite turns on parallel test execution. What breaks in an extension that starts one container and one stub server for the whole JVM?Both become shared mutable state: one test's stub mappings or database rows are visible to another running concurrently, and per-test resets wipe state a parallel test is relying on. You must either serialise the affected tests with resource locks, partition the resources (a schema or stub namespace per thread), or make matching strict enough that concurrent tests cannot collide — and decide this deliberately rather than discovering it as flake.
saying these in an interview costs you the question
- Starting the container in BeforeAllCallback per test class instead of once per JVM
- Writing teardown that assumes setup completed successfully
- Swallowing failures in the exception handler instead of capturing and rethrowing
- Ignoring @TestInstance(PER_CLASS) and parallel execution when designing shared state