skip to content

A JUnit 5 suite runs concurrently, but wall-clock time barely improved: many classes declare exclusive resource locks and several are marked to run in isolation. How do you reason about the tradeoff between locking shared resources and redesigning the tests?

level: principalimportance: nice to knowfreq 18%

answer

  1. measure key × held time before tuning
  2. ladder: delete sharing → narrow key → narrow scope → downgrade mode
  3. isolation = barrier, audit every one
  4. some sharing is irreducible (one JVM locale, one licence)
  5. sharding CI may beat isolating tests

basics

~20 s

Measure the contention first: which keys are hot, how long each lock is held, how many classes are isolated. Then attack in order — remove the sharing where a test can own its state, narrow keys and lock scope where it cannot, and keep whole-suite isolation only for genuinely unnameable interference like timing measurements.

solid answer

~60 s

I treat the locks as a map of the suite's shared-state debt and work it in priority order. **Measure.** Rank keys by total held time and by how many tests declare them, and count the isolated classes and their duration. Usually a couple of keys and two or three isolated classes explain most of the lost parallelism. **Then escalate through four moves.** (1) *Remove the sharing*: per-test temp dirs, ephemeral ports, injected clock/locale/config, a schema or transaction per test. This deletes the lock rather than tuning it. (2) *Narrow the key*: one lock per partitionable resource instead of one coarse key everything queues on. (3) *Narrow the scope*: annotate the methods that touch the resource rather than the class, and shorten what happens inside the lock. (4) *Downgrade the mode*: tests that only observe declare shared read access, so readers overlap. **Isolation last.** Every isolated class is a suite-wide barrier; most turn out to touch one nameable resource. Keep it for measurement-style tests, and consider moving those out of the unit suite entirely. And I compare the engineering cost against simply sharding the suite across CI machines.

go deeper

for a junior

Say that broad locks and isolation serialise the suite, and that locking fewer methods and using shared read access where the test only observes helps.

for a middle

Present the escalation ladder — remove the sharing, narrow the key, narrow the scope, downgrade the mode — with a concrete example of each.

for a senior

Lead with measurement: rank keys by held time, quantify the barrier cost of isolation, and name the irreducible cases worth keeping.

for a principal

Make it an economic and organisational decision: effort versus recovered minutes times run frequency, alternatives outside the JVM such as sharding and tiering, and the review guardrails that stop the debt returning.

## Recognising the situation A suite that is nominally parallel but effectively sequential is a common end state. It happens because locking is the cheap local fix: a test flakes, someone adds an exclusive lock or an isolation annotation, the build goes green, and nobody measures the aggregate. After a year the suite has the complexity of concurrency and the throughput of sequential execution — the worst of both. The principal-level skill here is not knowing the annotations; it is knowing how to decide where to spend effort, and when to stop. ## Step 1 — quantify before touching anything Without data you will optimise the wrong lock. The things worth counting: - **Per key: number of declaring tests × total time held.** A key declared by 200 fast tests can cost less than one declared by 4 slow ones. Time held includes setup and teardown, since the lock spans the whole node. - **Exclusive versus shared declarations per key.** A key where every declaration is exclusive is a pure queue; if most of those tests only observe the resource, the mode is simply wrong and the fix is one-line. - **Isolated classes: count and duration.** Each is a barrier that drains and refills the pool, so its real cost is roughly its duration multiplied by the parallelism, plus the ramp on either side. - **Achievable floor.** The critical path is the longest chain of work that cannot overlap. If that floor is close to your current wall-clock, no amount of lock tuning helps and the answer is elsewhere. ## Step 2 — the escalation ladder **1. Delete the sharing.** The best lock is the one that does not exist. Most shared state in test suites is incidental: a fixed port becomes an ephemeral one, a shared temp file becomes a per-test directory, a mutated JVM default becomes an injected parameter, a shared database becomes a schema or a rolled-back transaction per test. This costs the most engineering time and returns the most, because it removes the constraint permanently instead of scheduling around it. **2. Narrow the key.** One coarse key that every integration test declares is a global bottleneck wearing a scoped disguise. If the underlying resource is partitionable — per table, per tenant, per port, per property — give each partition its own key so unrelated tests stop queueing. The cost is a convention only your codebase knows, which is a real downside for JVM-wide resources where interoperability with other people's tests matters more than throughput. **3. Narrow the scope.** A class-level lock is held for every test in the class, including the many that never touch the resource, and it also forces the class's methods onto one thread. Moving the annotation to the two methods that need it often recovers most of a class's parallelism for the cost of a two-line diff. Likewise, move slow setup that does not need the resource outside the locked node. **4. Downgrade the mode.** Tests that only observe should declare shared read access so they overlap each other, pausing only around genuine writers. This is frequently the highest ratio of benefit to effort, because the exclusive default gets copy-pasted onto read-only tests without thought. **5. Question the isolation.** Audit each isolated class and ask what exactly it interferes with. Most answers name a specific resource, which means it can be downgraded to a scoped lock. The ones that survive are usually measurements — throughput, latency, memory — where the interference is CPU and memory contention with no key to name. Those may not belong in the parallel unit suite at all; a separate, sequential performance suite gives a better signal and stops penalising everyone else. ## Step 3 — know when locking is the right answer Not every lock is debt to repay. Some shared state is imposed and irreducible: the JVM has one default locale and one set of system properties; a licensed external system may allow one session; a hardware device or an emulator may be genuinely single. Locking these is correct engineering, and the review question is only whether the key is right, the scope is minimal and the mode is honest. The judgement call is economic. Making a test independent might take an hour or a week. Compare that against the wall-clock minutes it returns, multiplied by how often the suite runs and how many engineers wait on it. A ten-minute suite run 200 times a day justifies far more investment than a two-minute suite run twice. ## Step 4 — consider the levers outside the suite Intra-JVM parallelism is only one dial. Sharding the suite across several CI machines sidesteps in-process contention entirely and often costs less engineering time than isolating tests, at the price of CI spend. Splitting fast unit tests from slow integration tests lets the fast tier stay clean and highly parallel while the contended tier runs on its own schedule. And deleting redundant tests is the underrated option: suites accumulate coverage that no longer earns its runtime. ## Step 5 — keep it from regressing Whatever you fix will come back unless the ratchet holds: track the number of exclusive locks and isolated classes as a visible metric, require a reason string on every isolation, and make new occurrences a review conversation rather than a silent commit. The failure mode is cultural, not technical — locking is the path of least resistance for an engineer under deadline, so the guardrail has to sit in review, not in the framework.

  • How do you decide whether making a test independent is worth the engineering time?
    Compare the wall-clock minutes recovered, multiplied by how often the suite runs and how many people wait on it, against the estimated effort. A ten-minute suite on every pull request for a large team justifies days of work; a two-minute suite run occasionally does not. I would also weigh the secondary benefit — independent tests are easier to run individually and to debug — which often tips a marginal case.
  • Which locks would you keep permanently, without treating them as debt?
    Those guarding genuinely irreducible singletons: the JVM's default locale and time zone, system properties, the standard output streams, a licensed external system that permits one session, or a physical device or emulator. For those the review question is only whether the key is canonical, the scope minimal and the access mode honest — not whether the lock should exist.
  • What would make you conclude the whole parallel-execution effort should be abandoned?
    If measurement shows the critical path is dominated by a small number of irreducibly serial tests, or that most classes end up locked or isolated anyway, then concurrency is buying little while adding non-determinism and debugging cost. At that point sharding across CI machines, splitting fast and slow tiers, or deleting redundant tests will return more per hour invested, and reverting to a simpler sequential run is a legitimate outcome.

The locks are a road network's traffic lights. Retiming them helps a little; the real gains come from finding the two junctions everything funnels through and building a bypass — or accepting the junction and routing half the traffic to another road entirely.

saying these in an interview costs you the question

  • Tuning individual locks without measuring which keys actually dominate
  • Treating every isolated class as untouchable rather than auditing what it conflicts with
  • Assuming all shared state can be designed away, including JVM-wide singletons
  • Optimising wall-clock time while ignoring the flake rate the change introduces
  • Ignoring cheaper levers such as sharding across CI machines or splitting fast and slow tiers

context