Your JUnit 5 suite occasionally wedges in CI and a job has to be cancelled manually. How would you design a time-limit policy across the suite, and what are the tradeoffs of the options you would consider?
answer
- liveness net vs latency assertion
- generous global default, tighter for fixtures
- same-thread default protects transactions
- interrupt-based = advisory; job watchdog is the real cap
- thread dump before the kill
basics
~20 sSet a generous suite-wide default via junit.jupiter.execution.timeout.default as a deadlock net, tighten only where a hang is plausible, keep the default same-thread mode, disable timeouts on debug, and add an out-of-JVM job watchdog because interrupt-based limits cannot stop uninterruptible code.
solid answer
~50 sI would layer it. **Layer 1 — a suite-wide net.** `junit.jupiter.execution.timeout.default` set well above any legitimate runtime (minutes, not seconds). It never fires on a slow agent, but no single method can wedge the pipeline indefinitely. A tighter default for `@BeforeEach`/`@BeforeAll`, since hanging fixtures are the usual culprit. **Layer 2 — targeted `@Timeout`.** Class-level on integration classes touching networks or containers; method-level where a hang is a plausible failure mode. These are budgets for liveness, not latency assertions — latency belongs in a benchmark. **Layer 3 — outside the JVM.** Because same-thread enforcement is interrupt-based, uninterruptible code still hangs. A hard job-level or CI-level cap is the only real guarantee, and it should dump thread stacks before killing so the hang is diagnosable. I would keep `SAME_THREAD` as the default mode to protect transaction and `ThreadLocal` state, use `disabled_on_debug`, and treat every firing timeout as a bug to fix rather than a number to raise.
go deeper
Recognise that a global default exists and that a hanging test should fail rather than stall the build; a specific number matters less than the idea.
Distinguish liveness from latency, mention the configuration default plus targeted annotations, and note that debugging needs an escape hatch.
Own the diagnosis path: thread dumps, which layer catches what, why the JUnit-level limit is advisory in same-thread mode, and how to keep budgets from creeping.
Present it as layered policy with explicit tradeoffs — flakiness cost versus hang cost, thread-mode safety versus enforcement, and governance so firing timeouts drive fixes rather than number inflation.
## Framing the decision The question is not "what number" but "what is a timeout for here". Two different goals get confused: **liveness** (no run may hang forever) and **latency** (this operation must be fast). Test-level time limits are good at the first and bad at the second, because CI machines are shared, throttled, and highly variable. A policy that mixes the goals produces flaky tests, which erodes trust in the suite far more than the original hang did. ## Layer 1: the deadlock net The cheapest, highest-value move is one line of configuration: a suite-wide default applied to every method through `junit.jupiter.execution.timeout.default`. Choose it an order of magnitude above the slowest legitimate method — if the slowest honest test is twenty seconds, five minutes is a fine net. It will effectively never fire during normal operation, so it adds no flakiness, but it converts an infinite hang into a reported failure with a stack trace. Within the net it is worth differentiating by category. Fixtures deserve stricter budgets than tests: `junit.jupiter.execution.timeout.beforeeach.method.default` of thirty seconds is defensible when no honest per-test setup takes that long, and setup hangs are the most common wedge (a container that never becomes ready, a port already bound, a pool waiting for a connection that never returns). ## Layer 2: targeted annotations Above the net, annotate deliberately. Class-level `@Timeout` on integration test classes gives a tighter budget where hangs are plausible without touching the fast unit suite. Method-level annotations are for two cases: a method whose whole point is that it must not block (a circuit-breaker test, a client with a configured read timeout), and the reverse — a genuinely slow test that needs to opt out of a tighter class budget. The discipline to insist on: the number must be justified by an argument, not by measurement plus a fudge factor. "Our HTTP client is configured with a two-second read timeout, so this test cannot honestly exceed five seconds" is a maintainable rule. "It took 800 ms on my laptop so I set one second" is a future flake. ## Layer 3: the hard stop outside the JVM The uncomfortable truth is that in the default same-thread mode Jupiter's limits are advisory. The watchdog interrupts the test thread; code that ignores interruption — a busy loop, a classic blocking socket read, a native call, a `catch (InterruptedException) { /* retry */ }` — carries on, and the failure is only recorded when the method returns. So the only genuine guarantee is above the JVM: a per-job wall-clock cap in CI, ideally preceded by a thread dump so the hang is diagnosable rather than merely killed. If you skip the dump you get a red job and no evidence, and the hang recurs. ## The thread-mode call Switching the suite-wide thread mode to `SEPARATE_THREAD` looks tempting because it makes limits preemptive. I would not do it globally. It abandons the runaway thread — which may hold a lock, an open transaction, or a pooled connection — so later tests can deadlock or see dirty state, turning one hang into a cascade. And it removes thread affinity, breaking Spring transactional rollback, `SecurityContextHolder`, and MDC. The right shape is same-thread by default with a narrow, commented opt-in for the handful of tests exercising known-uninterruptible code. ## Developer experience A strict policy must not punish debugging. `junit.jupiter.execution.timeout.mode = disabled_on_debug` keeps a single committed configuration that enforces on CI and stands down under a debug agent, so nobody edits the file locally and forgets to restore it. Time limits should also produce good failures: the message names the method and duration, but the useful artefact is a thread dump at the moment of the timeout, which an extension can produce. ## Governance Finally, treat firing timeouts as a signal to manage. If a limit fires, the default response is investigation, not a bigger number; a repository where budgets creep upward over time has quietly turned its net into decoration. Track which tests time out, and if the same one fires repeatedly, either the code under test lacks its own bounded I/O timeouts — the real fix, since a client with a proper read timeout cannot hang — or the test is doing something it should not. ## The summary answer Generous global net, category-specific tightening for fixtures, targeted annotations with justified numbers, same-thread default for safety, debugger-aware mode, hard cap outside the JVM with a thread dump, and a policy that a firing timeout is a bug report.
- A single test times out intermittently on CI but never locally. What do you do?First determine whether it is a real hang or contention: capture a thread dump at the timeout and compare against a healthy run. If the thread is blocked on a lock or an unbounded network call, that is a defect in the code or fixture and gets fixed. If it is simply slow under load, the budget was set as a latency assertion rather than a liveness net, and it should be widened substantially or removed in favour of the global default. Quarantining without diagnosis is the wrong answer, because intermittent CI hangs usually indicate a genuine concurrency or resource-leak bug.
- Why not just set an aggressive global default so hangs surface fast?Because an aggressive default fires under ordinary CI variance — a noisy neighbour, a cold JIT, a slow container pull — and produces failures unrelated to the change under test. Once a suite fails for reasons developers cannot act on, people start rerunning builds by reflex and stop reading failures, which costs far more than the occasional hang. The net should be set where firing is strong evidence of a defect.
It is layered like fire safety: sprinklers everywhere set to a temperature normal cooking never reaches, extra sensors in the kitchen, and a fire brigade outside the building for the case where the sprinklers cannot reach.
saying these in an interview costs you the question
- Proposing one tight global limit as the whole policy, without acknowledging CI variance and flakiness.
- Claiming a JUnit-level timeout guarantees the build cannot hang, ignoring that same-thread enforcement is interrupt-based.
- Switching the suite-wide thread mode to SEPARATE_THREAD as a hardening measure, ignoring thread leaks and lost transaction context.
- Using test time limits as performance regression detection.
- Responding to a firing timeout by raising the number as a matter of course.