A JUnit 5 test that asserts an operation completes within a fixed time budget passes locally but fails intermittently on the CI machine. How do you decide whether the test or the code is at fault, and how would you make the check trustworthy?
answer
- wall clock = code + machine + moment
- cold JIT, GC pause, CPU quota, parallel tests, cold caches
- overrun size + trend distinguishes regression from noise
- budget 5–10x steady state, time only the operation
- hang protection = @Timeout policy; latency = JMH/SLO
basics
~20 sWall-clock budgets measure the machine as much as the code. Check whether CI is slower, noisier or colder (JIT, GC, shared runners, cold caches). Keep timeout assertions only as order-of-magnitude regression guards with generous budgets; move real latency checks to a benchmark or production SLO.
solid answer
~60 sFirst separate signal from noise. Collect the actual durations across runs — JUnit's `assertTimeout` failure message reports the overrun, which tells you whether you missed by 3 ms or 3 seconds. A small, variable overrun points at the environment: shared or throttled CI runners, cold JIT on a short-lived JVM, GC pauses, parallel test execution competing for CPU, cold DB/filesystem caches, container CPU quotas. A large, consistent overrun points at a real regression such as an added network call or an N+1 query. Then fix the design of the check. A timeout assertion should be a **gross-regression guard**, not a benchmark: budget several times the observed steady-state cost so only order-of-magnitude changes trip it. Remove fixture setup from inside the timed block. Eliminate real I/O with stubs so you are timing your code, not someone's network. If you actually need latency numbers, use a proper harness (JMH) or measure in production against an SLO. And if the test's real purpose is "never hang the build", express that as a `@Timeout` policy rather than a per-assertion budget.
go deeper
Say that timings vary by machine and load, that CI is slower and colder, and that budgets need generous headroom.
Add concrete sources of variance (JIT, GC, CPU quota, parallel tests, cold caches) and the practice of timing only the operation under test.
Lead with triage: measure durations across runs, distinguish a step change from a wide band, name a mechanism before declaring a regression, then right-size or relocate the check.
Argue about where performance verification belongs at all — regression guards in the suite, benchmarks in a nightly job on stable hardware, latency objectives in production — and about the cost of flake to team trust in the suite.
## Why wall-clock assertions are structurally flaky A timeout assertion compares elapsed wall-clock time against a constant. Elapsed time is a property of *code plus machine plus moment*. On CI the machine and moment vary far more than developers expect: - **Shared and throttled runners.** Hosted runners are usually virtualised, oversubscribed, and CPU-quota-limited. A container with a 1-CPU quota that briefly needs two threads is throttled, not merely slow. - **Cold JIT.** A CI JVM lives for one build. Code that is C2-compiled after thousands of local iterations may still be interpreted on CI, which is routinely 10–50x slower for hot loops. - **GC pauses.** A young-gen collection landing inside a 200 ms budget is enough to blow it. - **Parallel execution.** If the suite runs tests concurrently, your timed block competes with everything else for CPU and connection-pool slots. - **Cold caches.** First-touch page cache, empty database buffer pool, empty connection pool (paying handshake cost), lazily initialised singletons. - **Noisy neighbours and I/O contention** on shared infrastructure. So an intermittent failure is the expected behaviour of a tight budget, not an anomaly. ## Triage: is it the code or the environment? A practical sequence: 1. **Read the failure message.** `assertTimeout` reports `execution exceeded timeout of 200 ms by 17 ms`. An overrun that is a small fraction of the budget and varies run to run is noise. An overrun of many multiples, appearing on every run since a specific commit, is a regression. 2. **Instrument rather than guess.** Log the measured duration on every run (pass or fail) and chart it over builds. A step change at a commit is a regression; a wide band with occasional excursions is noise. Percentiles beat single observations. 3. **Bisect against the budget, not the pass/fail.** Run the suspect commit range with the assertion relaxed and durations logged; the step will be visible even where the test passed. 4. **Look for a mechanism.** Real regressions have causes you can name: an added HTTP call, a lost cache, a query that became N+1, a lock introduced on a hot path, a new synchronous serialization step. If you cannot name one, treat it as noise. 5. **Reproduce the constraint locally.** Pin the JVM to one CPU (`taskset`, or a container with a CPU quota) and run with a cold JVM. If it reproduces, the budget was calibrated on a machine you do not deploy or build on. ## Making the check trustworthy **Right-size the budget.** Decide what the assertion is *for*. If it is "catch a 10x regression", set the budget at roughly 5–10x the observed steady-state duration on the *slowest* environment that runs it. A budget within 2x of the mean will flap forever. **Time only the operation.** Move object construction, container startup, warm-up calls and data seeding outside the timed lambda. It is astonishingly common to time the fixture. **Remove real I/O.** Stub HTTP with an in-process server, use an in-memory or containerised database that is already warm, pre-open the connection pool. Every real dependency inside the block imports its variance into your test. **Warm up if the point is steady-state cost.** Call the operation a few times before the timed call, or accept that you are measuring cold-start and budget accordingly. **Make it a policy, not an assertion, when that is what you mean.** "No test may run longer than 30 seconds" is hang protection: express it with the `@Timeout` annotation or a global timeout configuration parameter, not with per-call `assertTimeout`. That keeps a single knob you can tune, and it does not pretend to be a statement about the code's performance. **Consider deleting it.** Many timeout assertions were added after one incident and never tuned. If the assertion has never caught a real regression and has cost hours of flake triage, its expected value is negative. Replace it with a benchmark in a nightly job (JMH, or a load test) and a production latency SLO with alerting — those measure on stable hardware or on the only hardware that matters. **Never make it self-adjusting.** Auto-widening the budget on failure, retrying until it passes, or skipping the assertion on CI all convert a flaky test into a test that never fails — worse than deleting it, because it still costs runtime and implies coverage that does not exist. ## What good looks like A small number of timeout assertions on operations with hard, well-understood budgets ("this in-memory index lookup must not become a database call"), each with a budget an order of magnitude above the norm, each timing only the operation, each with a comment naming the regression it exists to catch. Everything else about performance lives in benchmarks and production monitoring.
- A developer proposes fixing the flake by retrying the timed assertion up to three times and passing if any attempt is fast enough. What is your response?That turns a noisy signal into essentially no signal: with three attempts the test only fails when every run is slow, so genuine moderate regressions stop being detected while the suite still pays the runtime. If the budget is too tight, widen it deliberately and document why; if the check has no value, delete it and move performance verification to a benchmark or a production SLO.
- How would you keep a timeout assertion useful when the same suite runs on both a fast developer laptop and a throttled CI container?Calibrate the budget against the slowest environment that runs the test and size it as an order-of-magnitude guard, so the fast machine simply passes with a wide margin. Avoid per-environment budgets driven by system properties, which quietly disable the check where it matters most; if the environments really differ that much, log durations everywhere and assert only on the coarse limit.
saying these in an interview costs you the question
- Tightening the budget to "make the test more valuable" instead of accepting environment variance
- Retrying or conditionally skipping the assertion on CI to hide flake
- Timing fixture setup inside the lambda and blaming the code
- Treating a single wall-clock sample on shared hardware as a performance measurement