How would you tune `maxParallelForks` to make best use of a CI machine, and what trade-offs do you weigh?
answer
- start ~half cores, then profile
- memory: forks × heap must fit RAM
- CPU-bound ~cores; I/O-bound can oversubscribe
- bounded by --max-workers
- availableProcessors may ignore cgroup limits
- find the knee of wall-clock curve
basics
~20 sStart from available processors (often half of them), then profile. Weigh CPU throughput against per-fork memory (heap × forks must fit RAM), GC overhead, and the --max-workers cap. Match the value to the actual CI machine, not your laptop.
solid answer
~50 sThere's no magic number — `maxParallelForks` must be tuned to the *target machine*. Begin with a CPU-derived value: ```kotlin maxParallelForks = (Runtime.getRuntime().availableProcessors() / 2).coerceAtLeast(1) ``` Then weigh: - **Memory**: each fork is a full JVM with its own heap. `forks × (heap + metaspace + overhead)` must fit container RAM, or you'll OOM or swap. On constrained CI containers this, not CPU, is usually the binding constraint. - **CPU vs I/O profile**: CPU-bound suites scale roughly with cores; I/O-bound suites (DB, network) may benefit from *more* forks than cores since each blocks. - **GC and context-switching**: too many forks cause GC thrash and scheduler churn that can make the suite slower. - **Global caps**: `maxParallelForks` is bounded by `--max-workers` / `org.gradle.workers.max`. - **Container CPU detection**: `availableProcessors()` may report host cores, not the container's cgroup limit — on older JVMs or misconfigured containers this over-forks. Pin an explicit value in CI. Measure wall-clock at a few settings and pick the knee of the curve.
code
kotlin · 6 linesval ciForks = (project.findProperty("testForks") as String?)?.toInt()
tasks.withType<Test>().configureEach {
maxParallelForks = ciForks
?: (Runtime.getRuntime().availableProcessors() / 2).coerceAtLeast(1)
maxHeapSize = "768m" // keep forks × heap within container RAM
}go deeper
Mention starting from core count / half-cores; not expected to reason deeply about memory or container caps.
Cover the CPU-vs-memory trade-off and the --max-workers cap; know the half-cores starting heuristic.
Reason about workload profile (CPU vs I/O), the throughput knee, per-fork heap budgeting, and container CPU detection pitfalls; advocate empirical tuning on the real agent.
Standardise a tunable, machine-aware policy across modules; tie fork counts to CI agent specs and capacity planning, ensuring reproducibility and avoiding OOM at scale.
## The goal Maximise throughput of a single `Test` task without exhausting the machine. `maxParallelForks` is the lever; the right value depends entirely on the hardware and the suite's resource profile. ## Starting point ```kotlin tasks.withType<Test>().configureEach { maxParallelForks = (Runtime.getRuntime().availableProcessors() / 2).coerceAtLeast(1) } ``` Half the cores is conservative because each fork uses CPU *and* memory and may spin its own threads. ## Constraint 1 — memory Each fork is an independent JVM. Total footprint ≈ `forks × (maxHeap + metaspace + thread stacks + native)`. On a CI container with, say, 4 GB, four forks each allowed `-Xmx1g` already saturates RAM before overhead. When memory is the binding constraint (common in CI), more forks just causes OOM kills or swap that destroys throughput. Set per-fork heap deliberately: ```kotlin tasks.test { maxParallelForks = 4 maxHeapSize = "768m" } ``` ## Constraint 2 — workload profile - **CPU-bound** (parsing, crypto, pure compute): speedup tracks core count; don't exceed cores by much. - **I/O-bound** (testcontainers, HTTP, DB): each test spends time blocked, so *oversubscribing* (more forks than cores) can help — but watch the downstream resource (DB connection pool, container limits). ## Constraint 3 — overhead curve Throughput rises with forks, plateaus, then *declines*: GC pressure across many heaps, OS context-switching, and contention on shared resources (disk, the build cache, a single DB) dominate. The optimum is the **knee** of the wall-clock-vs-forks curve, found empirically. ## Constraint 4 — global caps & container awareness - `maxParallelForks` can never exceed `--max-workers` (a.k.a. `org.gradle.workers.max`), the daemon-wide worker cap shared with other parallel tasks. - `Runtime.availableProcessors()` historically reported *host* cores inside containers, ignoring cgroup CPU quotas. Modern JVMs (with `UseContainerSupport`) respect cgroups, but to be safe in CI, pin an explicit number via a property: ```kotlin val ciForks = (project.findProperty("testForks") as String?)?.toInt() tasks.withType<Test>().configureEach { maxParallelForks = ciForks ?: (Runtime.getRuntime().availableProcessors() / 2).coerceAtLeast(1) } ``` ## Method 1. Benchmark the suite at forks = 1, 2, 4, 8 on the *actual CI agent*. 2. Watch wall-clock **and** peak memory / GC logs. 3. Pick the highest value before memory pressure or the plateau. 4. Re-tune when the machine size or suite changes. ## Don't forget shared state Raising concurrency exposes class coupling (fixed ports, shared schemas). Combine tuning with isolation fixes (random ports, per-fork DB) or a separate sequential Test task for the offenders.
- Why might more forks than CPU cores still help?For I/O-bound tests (DB, network, containers) each fork spends much of its time blocked, so oversubscribing keeps cores busy — provided the downstream resource (connection pool, DB) can take the load.
- Why pin an explicit fork count in CI instead of relying on availableProcessors()?In containers, availableProcessors() can report host cores rather than the cgroup CPU quota (especially on older JVMs), causing over-forking and OOM. An explicit property matched to the agent is safer and reproducible.
- What caps maxParallelForks regardless of what you set?The daemon-wide worker limit, --max-workers / org.gradle.workers.max, which is shared across all parallel tasks.
saying these in an interview costs you the question
- Recommending a fixed number (e.g. 8) without reference to the target machine's cores/RAM.
- Ignoring per-fork memory — assuming only CPU matters.
- Assuming availableProcessors() always reflects container CPU limits.