You own a large JUnit 5 suite where occasional hangs block the build for an hour and a scattering of hand-written time-budget assertions flap. What overall strategy would you set for time limits across that suite, and what would you deliberately not do?
answer
- hang protection = policy; performance = measurement
- junit.jupiter.execution.timeout.default + @Timeout overrides
- disabled_on_debug so breakpoints don't fail tests
- thread mode: separate aborts but loses tx/security context
- never retry, auto-widen, or skip-on-CI
basics
~20 sSeparate two concerns: hang protection is policy — one global JUnit timeout plus per-test overrides, enforced declaratively — while performance verification belongs in benchmarks and production SLOs. Keep only a few coarse timeout assertions as regression guards, and never retry or auto-widen them.
solid answer
~50 sTreat time limits as two distinct problems. **Hang protection is policy.** Set a suite-wide default via the JUnit configuration parameter `junit.jupiter.execution.timeout.default` (with finer-grained variants for test methods and lifecycle methods), override per test with `@Timeout` where a case is legitimately long, and use `junit.jupiter.execution.timeout.mode=disabled_on_debug` so breakpoints do not fail tests. Choose the thread mode consciously: separate-thread aborts hangs but loses thread-bound transaction and security context. Back it with a build-level wall so a wedged JVM cannot burn an hour. **Performance verification is measurement.** Wall-clock assertions on shared runners measure the machine. Keep a handful only as order-of-magnitude regression guards — "this must not become a network call" — with budgets several times steady state, timing only the operation. Real latency questions go to JMH benchmarks on stable hardware and to production SLOs. Deliberately avoid: retries or auto-widened budgets, per-environment budget switches that disable the check on CI, and preemptive timeout assertions sprinkled through Spring integration tests.
go deeper
Say that a global default timeout protects against hangs and that individual timing assertions need generous budgets.
Distinguish the declarative policy (configuration parameter plus @Timeout overrides) from per-test assertions, and mention debugger-friendly timeout mode.
Add thread-mode consequences for transactional tests, the CI job-level backstop, and triage of flapping budgets by measuring duration trends.
Own the framing: two problems with different owners, policy sized from the duration distribution, measurement relocated to benchmarks and production SLOs, and an explicit list of anti-patterns you will not accept.
## Two problems that look like one Every conversation about "timeouts in the test suite" conflates two goals with different owners, different mechanisms and different failure modes: 1. **Hang protection** — no single test may wedge the build. This is an operational property of the suite. 2. **Performance regression detection** — this operation must not get dramatically slower. This is a property of the code. Mixing them produces the symptom in the question: hand-written budgets that were meant as guards flap because they are calibrated like measurements, while genuine hangs still escape because no policy covers the tests nobody annotated. ## Designing hang protection **Make it a default, not a per-test decision.** JUnit Jupiter reads configuration parameters (from `junit-platform.properties`, the build tool, or system properties). The relevant knobs are the default timeout for all testable and lifecycle methods, plus narrower variants for test methods, `@BeforeAll`/`@BeforeEach` and their `After` counterparts. One line gives the whole suite a ceiling, including tests written next year by someone who never read this policy. That is the property you want: coverage by default. **Pick the ceiling from data, not intuition.** Look at the distribution of test durations. Set the default well above the slowest legitimate test — if the p100 legitimate test is 8 seconds, a 60-second default catches wedges without touching anything healthy. The default exists to stop hangs, not to police slowness. **Allow explicit, reviewed exceptions.** A container-starting integration test may legitimately need minutes; annotate it with `@Timeout` and a comment. An explicit override is auditable; silently disabling the policy is not. **Choose the thread mode with eyes open.** Jupiter can enforce `@Timeout` on the same thread (the test simply fails afterwards — no rescue from a true hang) or on a separate thread (it aborts at the deadline, but the test body then runs on a different thread and loses thread-bound transaction, security and MDC context, and cancellation is only an interrupt). Same-thread is the safe default for Spring integration tests; separate-thread is appropriate for context-free, hang-prone code. The distinction mirrors `assertTimeout` versus `assertTimeoutPreemptively` exactly. **Do not let timeouts fight the debugger.** Configure the timeout mode to disable while a debugger is attached, otherwise every breakpoint session ends in a spurious failure and developers start turning the policy off globally. **Add a hard outer wall.** JUnit cannot save you from a JVM wedged before or outside test execution (a stuck static initialiser, a container that never becomes healthy, a shutdown hook that blocks). A job-level timeout in the CI configuration is the backstop, sized in minutes, whose only job is to fail fast rather than diagnose. **Fix causes, not just symptoms.** A suite that hangs usually has production code missing timeouts: HTTP clients without connect/read timeouts, unbounded `lock()` instead of `tryLock(timeout)`, `queue.take()` instead of `poll(timeout)`, JDBC statements with no query timeout. Those hangs will also happen in production, where no JUnit policy is watching. A timeout policy in tests buys time to fix them; it is not the fix. ## Designing performance checks **Accept that CI wall clock is noisy.** Shared runners, CPU quotas, cold JIT, GC pauses, parallel execution and cold caches all move the number. Any assertion whose budget sits within a small multiple of the mean will flap forever. **Keep a small number of coarse guards.** The legitimate use of a timeout assertion is an order-of-magnitude tripwire with a nameable failure it exists to catch — "this lookup must stay in memory and not become a database round trip". Budget it 5–10x steady state on the slowest environment, time only the operation (no fixture construction inside the block), and comment the intent. **Put real measurement where measurement works.** Microbenchmarks belong in JMH, run on stable hardware, ideally nightly rather than per-commit. System latency belongs to load tests and, above all, to production SLOs with alerting — production is the only environment whose latency actually matters to users. **Delete guards that never fire.** An assertion that has produced only flake triage for a year has negative expected value. Removing it is a legitimate outcome of this review, not a lowering of standards. ## The explicit "do not" list - **No retries** around timing assertions, and no test-retry plugin applied to them: with three attempts the check only fails when every attempt is slow, so it stops detecting moderate regressions while still costing runtime. - **No auto-widening budgets** derived from previous runs — a ratchet that only loosens converges on "always passes". - **No environment-conditional budgets or skips** that disable the check on CI, which is precisely where it would fire. - **No preemptive timeout assertions inside transactional or security-aware integration tests**, because the thread hop breaks rollback, lazy loading and authentication in ways that surface as unrelated failures later in the run. - **No global default so tight** that it becomes the suite's dominant source of failures; a policy people disable is worse than no policy. ## How you would know it worked Build wall-clock has a hard ceiling; hangs fail in a minute rather than an hour and name the offending test. Timing-related flake drops to near zero because few assertions remain and each has real headroom. Performance regressions are caught by a nightly benchmark trend and by production latency alerts, both of which produce numbers you can act on instead of a single boolean from a noisy machine.
- Why not simply set the global JUnit default timeout aggressively low, say two seconds, and let teams override where needed?It inverts the policy's purpose: a low ceiling turns ordinary slowness into build failures, so the dominant signal becomes annotation churn and people eventually disable the policy wholesale. The default should be far above the slowest legitimate test so it only ever catches wedges, with reviewed @Timeout overrides for genuinely long cases.
- Your global policy uses separate-thread enforcement so hangs are actually aborted. What must you check before rolling it out across Spring integration tests?Whether the tests depend on thread-bound state — the transactional test's bound EntityManager, SecurityContextHolder, MDC, request-scoped beans — because running the test body on another thread loses all of it and produces lost rollbacks and detached-entity errors. In practice you keep same-thread enforcement for those suites and rely on the CI job timeout as the outer wall, reserving separate-thread mode for context-free code.
saying these in an interview costs you the question
- Treating hang protection and performance assertions as the same mechanism
- Setting a global timeout tight enough that normal tests fail, then letting teams switch it off
- Relying on JUnit alone with no CI job-level timeout for JVM-level wedges
- Adding retries to timing assertions and calling the suite stable