A colleague writes a JUnit 5 TestWatcher that re-runs failed tests and throws an exception when it cannot write its report file. What documented limits of TestWatcher make both of those a mistake, and what does JUnit do with an exception thrown from a watcher callback?
answer
- watcher observes, never influences — all callbacks void
- runs after the test AND after after-each callbacks
- exceptions logged at warning, otherwise ignored → silent data loss
- no callbacks for containers; @BeforeAll failure = nothing observed
- retry belongs to @TestTemplate, not to a watcher
basics
~20 sA TestWatcher may not influence execution: its callbacks return void and run after the test has already finished and after its after-each callbacks, so it cannot retry or change an outcome. Exceptions thrown from a watcher callback are logged and otherwise ignored — the build stays green while data is silently lost.
solid answer
~50 sTwo constraints kill both ideas. **It cannot influence execution.** `TestWatcher` is a result *observer*. All four callbacks return `void` and are invoked after the test node has completed — after its after-each callbacks — with the verdict already recorded. There is no hook to re-run, no way to convert a failure into a pass. Retry belongs to a test template or a custom invocation-based extension, not to a watcher. **Exceptions from a watcher are logged and otherwise ignored.** JUnit deliberately does not let a broken observer break the run, so a watcher that throws on an I/O error produces a warning in the log and nothing else: no failure, no red build, and the report is silently incomplete. If missing report data must fail the build, that check has to live somewhere that can fail — an after-each callback, a launcher-level listener, or a post-run build step. Also: watchers see only test methods, never containers, so a class whose `@BeforeAll` fails produces no callbacks at all.
code
java · 26 linespublic class SafeOutcomeReporter implements TestWatcher {
private static final Queue<String> ROWS = new ConcurrentLinkedQueue<>();
@Override
public void testFailed(ExtensionContext context, Throwable cause) {
append(context.getUniqueId() + "|FAILED|" + cause);
}
@Override
public void testSuccessful(ExtensionContext context) {
append(context.getUniqueId() + "|PASSED|");
}
private static void append(String row) {
try {
ROWS.add(row);
Files.writeString(Path.of("build", "outcomes.psv"), row + System.lineSeparator(),
StandardOpenOption.CREATE, StandardOpenOption.APPEND);
}
catch (IOException | RuntimeException ex) {
// never propagate: JUnit would log and ignore it, hiding the problem
System.err.println("[SafeOutcomeReporter] degraded, kept in memory only: " + ex);
}
}
}go deeper
State the core rule: a watcher only observes results, cannot change them, and its exceptions are ignored.
Add the timing (after the test and its after-each callbacks), the void signatures, and that containers produce no callbacks.
Reason about the failure modes: silent report loss, @BeforeAll failures producing no data, teardown having already destroyed the artefacts you wanted, and thread safety under parallel execution.
Draw the boundary for the team — observation in extensions, enforcement in the build or a launcher listener, retry in the template mechanism — and decide whether flaky-test data belongs in-process at all versus in a report consumer downstream.
## TestWatcher is an observer, by design `TestWatcher` exists so that extensions can *react to* test results — write a report, take a screenshot on failure, publish a metric, tag a run in an external system. The JUnit team drew a hard line around it: a watcher may observe, but it may not participate. Two design facts encode that line. **All four callbacks return `void`.** There is no return value the engine consults, no `ConditionEvaluationResult`-style verdict object, no "replace the outcome" API. **They are invoked after the fact.** By the time `testFailed` runs, the test's execution is over, its after-each lifecycle callbacks have already run, and the result has been determined. There is nothing left to change and no live test to re-enter. ## Why the retry idea fails Retrying means executing the test node again, which means creating another test invocation. Only the engine creates invocations, driven by the `@TestTemplate` mechanism — that is how repeated and parameterized tests produce N executions from one method. A watcher sits outside that loop entirely. Concretely, a retry implementation has to decide *before or during* execution how many invocations exist and what each one's outcome means; a watcher learns about an invocation only once it is finished and cannot add another. The correct shapes for retry are: a `@TestTemplate` with a custom `TestTemplateInvocationContextProvider` that yields further invocations while previous ones failed; a third-party retry extension built on that mechanism; or the retry support the build tool offers around whole test executions. A watcher can *feed* such a scheme — recording which tests are flaky — but it cannot implement it. ## Why throwing from a watcher fails JUnit treats extensions that observe results as non-critical infrastructure: an exception escaping a `TestWatcher` callback is **logged (at warning level) and otherwise ignored**. It does not fail the test that was just observed, it does not fail the containing class, and it does not fail the build. That is defensible — a reporting bug should not invalidate a thousand test results — but it has a sharp practical consequence: **failures in a watcher are effectively invisible**. A watcher that cannot open its output file will log a warning that no one reads, and the run finishes green with an empty report. Teams discover this months later when they try to use the report. The engineering answer is to make the watcher itself total: catch its own I/O errors, degrade gracefully (buffer in memory, write to `System.err`, publish through the reporting API), and never rely on a thrown exception to signal a problem. If "the report exists and is complete" is a real requirement, assert it somewhere that can fail — a final verification step in the build, or a launcher-level listener whose summary the build checks. ## Other limits worth knowing **Only test methods, never containers.** `TestWatcher` callbacks are produced for `@Test` methods and for each `@TestTemplate` invocation (each repetition, each parameterized argument set). Test classes and `@Nested` classes are containers and produce no callbacks. Corollary: if a class-level `@BeforeAll` throws, its tests are never started and the watcher hears nothing at all — the very case where you most want a report entry. Suite-complete reporting therefore belongs at the launcher level, where a `TestExecutionListener` sees containers, skips and the whole plan across engines. **Registration should be class level or higher.** In practice a watcher is registered with class-level `@ExtendWith`, on a shared test interface, via a static `@RegisterExtension` field, or through automatic `ServiceLoader` registration for suite-wide coverage. That is also how you avoid the surprise of a watcher that only observes a fraction of the run. **Concurrency.** With parallel execution enabled, callbacks arrive on multiple threads. Any accumulation inside the watcher needs a thread-safe structure; a plain `ArrayList` field will lose entries or throw. **Ordering with teardown.** Because watcher callbacks run after after-each callbacks, anything teardown destroys — a closed driver, a deleted temporary directory, a rolled-back transaction — is gone by the time the watcher runs. If you need to capture live state on failure (a screenshot, a database dump), do it in a handler that runs while the fixture is still alive, and let the watcher record only the outcome and a pointer to the artefact. ## How to answer the colleague Split the work by capability. Use a test-execution exception handler or an after-each callback for anything that must run while the test's world still exists and that may legitimately fail the build. Use a test template for retry. Use `TestWatcher` for what it is good at: a cheap, non-throwing, thread-safe record of each test method's verdict. And if the report must be complete across containers and engines, lift it out of the extension model into a launcher listener.
- Where would you implement a retry-on-failure mechanism in JUnit 5 instead?In the test-template machinery: annotate the method as a `@TestTemplate` and supply a `TestTemplateInvocationContextProvider` that keeps yielding invocations while the previous one failed, or use an existing retry extension built that way. That mechanism owns invocation creation, which is what retry needs and what a watcher can never do.
- Your watcher must capture a browser screenshot when a UI test fails, but by callback time the driver is already closed. How do you restructure it?Move the capture to a point where the fixture is still alive — a `TestExecutionExceptionHandler` that grabs the screenshot and rethrows, or an after-each callback that inspects `ExtensionContext.getExecutionException()` before teardown closes the driver. Leave the watcher to record the verdict and the path to the saved artefact.
- If exceptions from a watcher are swallowed, how do you detect that your reporting is broken?Do not rely on exceptions at all: have the watcher catch its own errors and record a degraded-mode marker, then verify the report as a separate build step that can fail — for example a launcher-level listener comparing observed test counts against report rows, or a post-run check that the file exists and has the expected number of entries.
saying these in an interview costs you the question
- Believing a TestWatcher can convert a failure into a pass or trigger a re-run
- Assuming a throw from a watcher fails the test or the build
- Expecting watcher callbacks for test classes or for tests never started because a class-level fixture failed
- Capturing live fixture state in a watcher, after teardown has already run
- Accumulating results in a non-thread-safe field while parallel execution is enabled