skip to content

How do you decide the degree of parallelism for a JUnit 5 test suite, and how do you judge whether running tests concurrently is paying off at all?

level: principalimportance: nice to knowfreq 24%

answer

  1. CPU-bound plateaus at core count
  2. I/O-bound: factor 2–4 until the far side saturates
  3. fixed on CI — containers misreport cores
  4. measure wall-clock AND flake rate AND heap
  5. Amdahl: one slow class caps the gain

basics

~20 s

Measure, don't guess. Start with the dynamic strategy at factor 1.0, then sweep the factor or a fixed parallelism and plot wall-clock time and flake rate. CPU-bound suites plateau near the core count; I/O-bound ones keep improving until an external system saturates. Pin a fixed value on CI for reproducibility.

solid answer

~60 s

I treat it as capacity planning with two axes: wall-clock time and reliability. **Sizing.** Classify the workload. CPU-bound unit tests plateau around one thread per core, so `dynamic` with factor 1.0. I/O-bound tests idle most of the time, so factor 2–4 until an external component — a stub server, a connection pool, a database — becomes the bottleneck rather than the CPU. On CI I switch to `fixed` parallelism, because the JVM's view of `availableProcessors` on a shared or containerised executor is often the host's, and over-subscription makes runs slower and flakier than sequential. **Judging payoff.** Sweep the value, record total wall-clock time, the p95 of the slowest class, the flake rate over repeated runs, and peak memory. Parallelism pays when wall-clock drops materially while flake rate stays flat. It does not pay when the suite is dominated by one slow class (parallelism can't split it), when everything contends on one database, or when the cost is a growing pile of serialisation annotations. I also budget the human cost: making tests isolated is real engineering time.

code

properties · 7 lines
properties
junit.jupiter.execution.parallel.enabled = true
junit.jupiter.execution.parallel.mode.default = same_thread
junit.jupiter.execution.parallel.mode.classes.default = concurrent
junit.jupiter.execution.parallel.config.strategy = dynamic
junit.jupiter.execution.parallel.config.dynamic.factor = 1.0
# CI overrides with -Djunit.jupiter.execution.parallel.config.strategy=fixed
#                  -Djunit.jupiter.execution.parallel.config.fixed.parallelism=4

go deeper

for a junior

Say that the number depends on cores and on whether tests are CPU- or I/O-bound, and that you would measure rather than guess.

for a middle

Give the concrete sweep procedure and the dynamic-versus-fixed choice, and mention external systems as the real ceiling for I/O-bound suites.

for a senior

Add reliability and memory as first-class metrics, the reproducibility argument for fixed parallelism on CI, and the tail-dominated case where threads cannot help.

for a principal

Present it as a cost-benefit decision across the whole engineering org: isolation effort, debuggability, serialisation debt, and alternatives such as sharding across machines or deleting tests.

## Framing: this is capacity planning, not a config value The question has no single right answer, which is the point. Choosing parallelism means trading three things against each other: wall-clock feedback time, the reliability of the signal, and the engineering effort required to make tests independent enough to overlap. A good answer talks about all three and about how you would *measure*, rather than naming a number. ## Step 1 — characterise the workload Different suites have different ceilings: - **CPU-bound pure unit tests.** Every thread is genuinely computing. Throughput peaks near one runnable thread per core; beyond that you pay context switching and cache pressure for nothing. `dynamic` with factor 1.0. - **I/O-bound tests** — HTTP calls to a stub, database round-trips, container start-up, deliberate waits. Threads are idle most of their life, so parallelism well above the core count helps: factors of 2–4 are common. The ceiling is no longer the CPU but the far side: the stub server's own thread pool, a JDBC connection pool, the database's max connections. - **Memory-heavy tests** — each concurrent test may hold a Spring context, a container client or a large fixture. Here the binding constraint is heap, and the honest limit is whatever keeps you clear of GC thrash and out-of-memory failures. This is the case people forget: doubling threads can double peak heap. - **Suites dominated by one long test or one long class.** Amdahl's law applies bluntly — if one class takes 8 minutes of a 10-minute suite, no thread count gets you below 8 minutes. The fix is splitting that class, not more threads. ## Step 2 — pick the mechanism to match the environment `dynamic` derives parallelism from `availableProcessors()`. That is right on a developer machine, where you may even want a factor below 1.0 so the IDE stays responsive. On CI it is often wrong: shared executors and containers can report far more processors than your job is actually allowed to use, and the resulting over-subscription produces runs that are slower *and* flakier than sequential. `fixed` with an explicit number, chosen per executor class, is the reproducible choice — and reproducibility matters independently, because a race that only appears at parallelism 8 should appear identically on every machine. When the number must be derived at runtime — a shard count, a provisioned schema count, an environment variable set by the CI system — a `custom` strategy class keeps that logic in one place instead of scattering it through build configuration. ## Step 3 — measure the right things Sweep the parallelism value (1, 2, 4, 8, 16) and record, for each: 1. **Total wall-clock time** — the reason for the exercise. 2. **Flake rate over repeated runs** — run the suite five or ten times per setting. A configuration that is 40% faster and 5% flaky is worse than one that is 30% faster and stable, because a flaky suite is a suite people stop trusting and start rerunning. 3. **Peak memory / GC time** — the silent limit. 4. **The tail** — the slowest class or method. If the tail dominates, further threads are wasted. 5. **Contention on external systems** — connection-pool wait time, container CPU throttling, database lock waits. The shape you expect is a curve that improves steeply, then flattens, then degrades. Choose a point on the flat part, slightly conservative, not the exact minimum — the minimum is the most fragile setting. ## Step 4 — count the costs people forget - **Isolation work.** Every test that shares global state must be fixed or serialised. That is engineering time and it competes with product work. If the suite is 4 minutes, spending three weeks to make it 2 minutes is a bad trade; if it is 40 minutes and blocks every merge, it is an excellent one. - **Debuggability.** Interleaved output and stack traces from several tests at once make failures harder to read. Teams often keep a sequential mode available for local debugging. - **Serialisation debt.** Each class pinned to a single thread or guarded by an exclusive resource is a permanent brake. A suite where half the classes are serialised has the complexity of parallelism and the speed of sequential — that is the state to detect and either finish or abandon. - **Alternative levers.** Parallelism inside one JVM is not the only tool: splitting the suite across CI machines, running slow and fast tiers separately, and simply deleting redundant tests can beat it. A principal-level answer names these rather than assuming thread count is the only dial. ## Step 5 — decide and revisit Commit a default in the shared configuration, allow per-environment overrides through system properties, and record the reasoning where the next person will read it. Revisit when the suite's shape changes — a new integration tier, a move to different CI hardware, a jump in memory per test. The number is not a constant; it is a function of the suite and the machine, and it should be re-measured rather than inherited.

  • When would you conclude that parallel execution is not worth it for a given suite?
    When the measured wall-clock gain is small relative to the cost. Typical cases: a suite already fast enough that minutes are not the bottleneck; one dominating slow class that no thread count can split; heavy contention on a single database that turns concurrency into queueing; or a suite so entangled with global state that most classes end up serialised anyway. In those cases splitting the suite across CI machines or deleting redundant tests usually beats threads.
  • How do you keep a parallel suite's flake rate honest over time?
    Measure it deliberately rather than waiting for complaints: a scheduled job that runs the suite several times in a row at the production parallelism, and a record of test failures per run. New races arrive with new tests, so the check has to be continuous. I would also track the count of classes pinned to same-thread execution, because a rising count means the team is quarantining rather than fixing.

Sizing parallelism is like adding checkout lanes: the first few cut the queue dramatically, but once the packing area behind them jams — the database, the heap, one enormous trolley of a test — extra lanes just move the queue somewhere less visible.

saying these in an interview costs you the question

  • Naming a fixed number like 'always 4 threads' with no reference to the workload
  • Ignoring memory as a limiting factor when each test holds a heavyweight context
  • Optimising only wall-clock time and never measuring the flake rate
  • Assuming more threads always helps, even when one slow class dominates the suite
  • Overlooking that dynamic sizing over-subscribes on containerised CI executors

context