You inherit a Kotest suite of several thousand tests that takes 40 minutes and runs entirely serially. How would you decide how much of it to run concurrently, and how would you keep the result trustworthy rather than intermittently red?
answer
- measure per-spec timings before touching knobs
- concurrentSpecs before concurrentTests, module by module
- N consecutive green runs + randomised order as the bar
- @Isolate contains, tracked as debt; fix shared state
- budget connections, containers, ports; expect log interleaving
basics
~20 sMeasure first to find whether the time is IO-bound or dominated by a few slow specs, then raise Kotest's concurrentSpecs before concurrentTests, module by module. Run each step many times, @Isolate the specs that fail while tracking them as debt, fix shared state rather than serialising around it, and pair the rollout with randomised ordering so latent coupling surfaces deliberately.
solid answer
~60 sI would treat it as a migration, not a config change. **Measure.** Per-spec timings tell you whether 40 minutes is thousands of small IO-bound tests (concurrency helps a lot) or three specs taking 12 minutes each (concurrency across specs buys almost nothing until those are split). **Pick the axis.** IO-bound suites gain from coroutine concurrency (`concurrentSpecs`) with modest threads; CPU-bound bodies need `parallelism`. Raise cross-spec concurrency first — specs are usually better isolated from each other than tests are within a spec — and only then consider `concurrentTests`. **Prove it.** Concurrency bugs are probabilistic, so a single green run means nothing. Run the module repeatedly, ideally with randomised ordering on, so coupling shows up as a deliberate finding rather than a surprise months later. **Contain, then fix.** `@Isolate` the specs that fail so the build moves, but log each as isolation debt with an owner. The durable fixes are per-test transaction rollback, a schema or container per worker, removing global mutable state, and unique ports and temp names. **Watch the cost.** More concurrency means more connections and containers, and interleaved logs; budget the infrastructure and improve per-test output capture before widening.
go deeper
Know the suite runs serially by default and that concurrency is configured centrally, and defer the rollout judgment.
Describe raising concurrentSpecs first, running repeatedly to prove it, and using @Isolate for offenders.
Own the module-by-module migration, the acceptance bar, and the structural fixes for shared database and global state.
Frame it as measurement, sequencing and cost: match the lever to the bottleneck, budget infrastructure, keep PR runs deterministic with a randomised nightly, and stop when trust outweighs the remaining minutes.
## Start with data, not settings The first mistake is turning knobs before knowing the shape of the 40 minutes. Collect per-spec durations. Three shapes are common and they lead to different answers. - **Long tail of small tests** — thousands of specs at a second or two, mostly waiting on IO. Coroutine concurrency across specs is close to a free win. - **A few whales** — three specs consuming most of the wall clock. Concurrency across specs cannot beat the longest single spec, so the first fix is splitting those specs, or enabling within-spec concurrency for them specifically. - **CPU-bound bodies** — heavy serialisation, crypto, big in-memory fixtures. Here threads (`parallelism`) matter more than coroutine concurrency, and you are bounded by the agent's cores. Stating this distinction is what separates a principal answer from a mid-level one: the target is not "make it concurrent", it is "identify the bottleneck and apply the matching lever". ## Order of adoption Raise `concurrentSpecs` first. Specs are usually reasonably independent of one another, whereas tests inside a spec frequently share fields on the spec instance and are the likeliest to race. `concurrentTests` is the second step, applied where a whale spec forces it, and it interacts with how much state the spec instance holds. Roll out module by module rather than repo-wide. A module is a unit someone owns and can triage; a repo-wide flip produces a wall of failures that nobody claims and that ends with the flag being reverted. ## Proving it, not hoping Concurrency defects are probabilistic. Establish an acceptance bar before widening: the module runs N times consecutively green, with randomised spec and test ordering enabled. Randomised ordering belongs in the same rollout because it finds the same class of defect — hidden coupling — in a form that is easier to debug than a race. A nightly job that runs the suite several times with concurrency and randomisation, separate from the PR job, is the standard structure. PR runs stay deterministic and fast; the nightly hunts for coupling. That keeps developer trust intact while still surfacing problems. ## Containment versus cure `@Isolate` is the pressure valve: a spec marked with it never runs alongside others, so one offender does not block the whole rollout. The discipline is that every `@Isolate` is recorded as debt with an owner and a reason. Left untracked, the annotation spreads until the suite is serial again, but now with the illusion of concurrency configured. The real fixes are structural: per-test transactional rollback or a schema per worker so the database stops being shared; deleting global mutable state; injecting a clock rather than freezing a global one; unique ports and generated temp paths; per-test rather than static mocking. `.config(blockingTest = true)` handles the narrower case of a body that blocks its thread and would otherwise starve the dispatcher. ## The costs nobody budgets Concurrency multiplies resource demand. Eight concurrent specs may need eight database connections, eight containers, eight ports — on an agent sized for one. The suite can get *slower* from contention, or start failing on pool exhaustion, and the team blames the framework. Size the agents and the pools as part of the change. Debuggability degrades too: interleaved log output means a failure's context is mixed with unrelated work. Per-test output capture and structured logging are worth doing before widening, not after the first confusing failure. ## Knowing when to stop There is a point where more concurrency buys minutes and costs trust. If the suite is at eight minutes and the remaining gain is two minutes at the price of intermittent redness and doubled infrastructure, stop — and put the next increment of effort into deleting redundant tests or moving coverage down the pyramid instead. The goal is a fast suite people believe, not a maximally parallel one.
- The suite gets slower after enabling concurrency. What would you look at?Resource contention is the usual cause: a connection pool or container set sized for serial execution, CPU oversubscription on the agent, or blocking test bodies starving the dispatcher so work queues instead of overlapping. Check pool sizes and agent cores first, look for tests that block rather than suspend, and confirm the suite is not dominated by one long spec that concurrency across specs cannot shorten.
- How do you decide whether a failing spec gets @Isolate or a real fix?@Isolate is a containment step so the rollout continues, never the destination. Give it a reason and an owner. Whether the fix follows immediately depends on blast radius: a spec sharing a database is worth fixing because the same pattern will exist elsewhere, while a genuinely process-global concern such as a system property may justify permanent isolation. The test is whether the annotation is a considered decision or a way of avoiding one.
saying these in an interview costs you the question
- Flipping concurrency on repo-wide and treating the resulting failures as flakiness
- Accepting one green run as proof the suite is concurrency-safe
- Leaving @Isolate scattered untracked until the suite is effectively serial again
- Ignoring infrastructure limits so concurrency causes connection-pool or port exhaustion
- Chasing maximum parallelism past the point where the team still trusts the result