skip to content

You switched a large JUnit 5 suite to concurrent execution and a handful of tests started failing intermittently. How do you diagnose the failures and roll the change out safely?

level: seniorimportance: should knowfreq 34%

answer

  1. disable the switch to attribute the failure
  2. classes concurrent / methods same_thread as bisect
  3. fixed parallelism = 2 to reproduce
  4. statics, system properties, ports, temp files, DB rows
  5. SAME_THREAD is debt, record it

basics

~20 s

First prove concurrency is the cause: rerun with the parallel switch off. Then narrow it — set the default execution mode back to same_thread and opt classes in with @Execution(CONCURRENT), or keep concurrent globally and mark suspects @Execution(SAME_THREAD). Fix the real cause: static fields, system properties, fixed ports, shared files or database rows.

solid answer

~50 s

Treat it as a bisect, not a mystery. 1. **Confirm the cause.** Rerun with `junit.jupiter.execution.parallel.enabled=false`. Green sequentially and red concurrently means shared state, not a broken test. 2. **Narrow the blast radius.** Drop to the safest shape — `mode.classes.default=concurrent`, `mode.default=same_thread` — so the unit of concurrency is a whole class. If failures vanish, the shared state was inside classes; if they persist, it is JVM- or environment-global. 3. **Reproduce deterministically.** Pin a low fixed parallelism (2) so the interleaving is simple, and rerun the failing pair repeatedly. 4. **Find the state.** Look for `static` mutable fields, `System.setProperty`, default locale/time zone, singletons and caches, hardcoded ports, a shared temp directory, and shared database rows. 5. **Fix, don't mute.** Make the test own its state — fresh instances, per-test temp dirs, random ports, unique identifiers. Serialising with `@Execution(SAME_THREAD)` is a stop-gap you record as debt. 6. **Ratchet.** Land parallelism class-by-class behind annotations, with a repeat-run job to smoke out remaining races.

code

java · 12 lines
java
// Racy: mutates a JVM-wide setting other tests may read mid-run
@Test
void formatsInBerlinTime_racy() {
    TimeZone.setDefault(TimeZone.getTimeZone("Europe/Berlin"));
    assertEquals("12:00", format(instant));
}

// Fixed: pass the zone in instead of mutating the JVM default
@Test
void formatsInBerlinTime() {
    assertEquals("12:00", format(instant, ZoneId.of("Europe/Berlin")));
}

go deeper

for a junior

Say that concurrency exposes shared state, name the obvious kinds (static fields, system properties, files) and that running with the switch off tells you whether concurrency is the cause.

for a middle

Add the bisect procedure using the two default-mode parameters and per-class @Execution, and describe reproducing with a low fixed parallelism.

for a senior

Own the rollout: ratcheted enablement, repeat-run jobs, fixing isolation instead of serialising, and tracking the SAME_THREAD annotations as debt.

for a principal

Judge whether parallelism pays for itself at all: compare wall-clock gain against flake rate and engineering time, and set the policy for how new tests are expected to be written.

## Why concurrency exposes bugs that were always there A sequential suite grants every test an accidental guarantee: nothing else is running. Under `@Execution(CONCURRENT)` or a `concurrent` default mode that guarantee disappears, and any state shared between two tests becomes a race. The tests were latently broken; parallelism made them observable. Framing it that way matters in an interview, because the wrong instinct — "parallel execution is flaky, turn it off" — is exactly the answer that fails the question. ## Step 1 — attribute the failure Rerun the same commit with the master switch off (`junit.jupiter.execution.parallel.enabled=false`, easiest as a JVM system property so you change nothing in the repository). If it is green sequentially and intermittently red concurrently, concurrency is the trigger. If it is flaky both ways, you have an ordinary flaky test (time, randomness, external service) and parallelism is a red herring. ## Step 2 — narrow with execution mode, not with luck Execution mode is your bisect tool, and it works at two granularities: - Set `junit.jupiter.execution.parallel.mode.classes.default=concurrent` with `junit.jupiter.execution.parallel.mode.default=same_thread`. Now the class is the unit of concurrency and everything a class privately owns is single-threaded again. Failures that disappear here were caused by two methods of the same class colliding — usually a `static` field, a `@TestInstance(PER_CLASS)` instance field, or a `ThreadLocal` populated in `@BeforeEach` and read on another thread. - Failures that survive are cross-class: JVM-global or environment-global state. Then bisect further by annotating individual classes `@Execution(SAME_THREAD)` and seeing which annotation makes the failure disappear. That identifies the culprit far faster than reading code. ## Step 3 — make the race reproducible Random interleavings are hard to debug. Pin the pool with `config.strategy=fixed` and `config.fixed.parallelism=2`: with two threads there are few possible interleavings and the failure often becomes near-deterministic. Then run just the two suspect classes together in a loop. Logging `Thread.currentThread().getName()` in the setup and teardown of the suspects gives you a cheap interleaving trace when the failure is rare. ## Step 4 — the usual culprits In order of how often they turn up: 1. **`static` mutable fields** — caches, counters, a lazily-built singleton, a shared `SimpleDateFormat`. 2. **JVM-global settings** — `System.setProperty`, `Locale.setDefault`, `TimeZone.setDefault`, redirected `System.out`. One test's setup silently reconfigures another test that is already running. 3. **Fixed external coordinates** — a hardcoded port, a fixed file path in the system temp directory, a fixed record id in a shared database. 4. **Shared containers or servers** started once for the suite and mutated per test. 5. **Order dependence** — test B only passed because test A ran first and populated something. Parallelism destroys the ordering assumption even before it creates a data race. 6. **Frameworks with global registries** — anything that installs a global default and restores it in teardown is a landmine when two tests overlap. ## Step 5 — fix at the right level The durable fix is to remove the sharing: give each test its own instance, its own temporary directory, an ephemeral port, a unique identifier prefix, its own database schema or a rolled-back transaction. Second best is to make the shared thing thread-safe or thread-confined. Only when neither is affordable do you serialise the test — pinning it to one thread with `@Execution(SAME_THREAD)` if the collision is between methods of one class, or declaring the shared thing as an exclusive resource when the collision is with other classes (that is a separate JUnit mechanism, and the right one when the state is genuinely global). Whatever you choose, record it: an `@Execution(SAME_THREAD)` with no comment becomes permanent, and permanent serialisation is why suites stop getting faster. ## Step 6 — roll out as a ratchet A safe rollout looks like this: - Enable the switch with defaults still `same_thread`, so nothing changes behaviourally and the configuration lands on its own. - Turn on class-level concurrency and fix the fallout. - Add a scheduled job that runs the suite several times in a row at a low fixed parallelism; races that appear once in twenty runs will not survive that. - Only then consider method-level concurrency, and expect a second, smaller wave of failures. Quarantining rather than fixing is acceptable for a day, not a quarter. Track the serialised classes as a list that must shrink, otherwise the suite converges back to sequential with extra annotations. ## Signals you should watch After the rollout, watch three numbers: wall-clock time (the reason you did this), the flake rate per run, and the count of `SAME_THREAD` annotations. If the second and third rise while the first barely moves, the parallelism is not paying for itself and you should either invest in test isolation or revert.

  • A test only fails when the whole suite runs, never on its own, even sequentially. Is that a concurrency bug?
    Not necessarily — that is the signature of an order dependency: some earlier test left state behind that this one relies on or is broken by. Confirm by running the suite sequentially in a shuffled or reversed order; if it still fails, it is leakage, not a race. Concurrency will also expose it, but the fix is the same: make the test set up and tear down everything it needs.
  • Would you ever ship parallel execution with several classes pinned to SAME_THREAD?
    Yes, as an explicit interim state. The speed-up from the rest of the suite is real and the pinned classes are no worse than before. The condition is that each pin carries a comment naming the shared state and there is a tracked list that shrinks, otherwise the annotations become invisible permanent debt and the suite quietly reverts to sequential.

saying these in an interview costs you the question

  • Concluding that parallel execution is inherently unreliable and reverting
  • Adding retries or sleeps instead of finding the shared state
  • Assuming JUnit will detect and serialise unsafe tests for you
  • Blaming the test that failed rather than the one that mutated global state
  • Leaving @Execution(SAME_THREAD) annotations with no comment or follow-up

context