skip to content

A self-hosted runner pool reuses agents across many Gradle jobs. How do you decide whether to disable or keep (but constrain) the daemon?

level: seniorimportance: should knowfreq 35%

answer

  1. speed vs predictability axes
  2. reuse real → keep + constrain
  3. tight limits / single-shot → disable
  4. incompatible jvm args spawn new daemons
  5. A/B benchmark a sequence

basics

~20 s

Measure reuse. If agents run many back-to-back Gradle builds, keeping the daemon speeds them up — but cap its memory and prevent accumulation. If reuse is rare or limits are tight, disable it for predictability.

solid answer

~40 s

It's a measured trade-off between **warm-build speed** and **predictability/leak risk**. Keep the daemon when reuse is real: a **reused agent** running many Gradle builds in sequence amortises start-up and benefits from JIT warm-up, so a resident daemon genuinely cuts wall-clock time. The cost is that you must **constrain** it — bound heap via `org.gradle.jvmargs` policy, prevent daemons accumulating, and clean up between jobs — or you risk OOM and orphans. Disable it (`org.gradle.daemon=false`) when: agents are effectively single-build, container memory limits are tight, or reproducibility/clean-room matters more than a few seconds. A single-use daemon is predictable and leak-free at the cost of warm-up. Decide with data: benchmark identical pipelines daemon-on vs daemon-off, watch peak memory against the cgroup limit and the orphaned-process count, and pick per agent class — not one global rule.

code

bash · 4 lines
bash
# benchmark a SEQUENCE of builds both ways, then compare
for i in 1 2 3; do time ./gradlew build; done                # daemon on
for i in 1 2 3; do time ./gradlew build --no-daemon; done    # daemon off
# also track: peak RSS vs cgroup limit, leftover java procs

go deeper

for a junior

Know the headline rule: reuse → daemon helps; throwaway → disable.

for a middle

List the conditions favouring each side (reuse, memory headroom, reproducibility).

for a senior

Drive the decision with A/B benchmarks and per-agent-class policy, and constrain a kept daemon.

for a principal

Set org policy mapping runner classes to daemon strategy, encode it in shared templates, and require evidence before exceptions.

## Framing the decision This is not dogma; it's an engineering trade-off with two axes: - **Speed:** daemon-on can be faster on reused agents (no repeated JVM start-up, accumulated JIT, warm caches). - **Predictability & safety:** daemon-off is cleaner — bounded resource use, no orphans, stronger reproducibility. The right answer depends on the **agent lifecycle** and **resource envelope**, so you decide per agent class. ## When KEEPING (constrained) wins - Agents are **long-lived and reused** for many Gradle jobs back-to-back. - The machine has enough RAM headroom that a warm daemon won't threaten the limit. - Build start-up is a meaningful share of pipeline time (many short builds). If you keep it, you must constrain: - bound the daemon/worker heap so it can't grow into the cgroup limit, - ensure a single compatible daemon is reused rather than many spawning (incompatible JVM args spawn *new* daemons), - add teardown so nothing accumulates across jobs. ## When DISABLING wins - Agents are **destroyed or effectively single-build** — no reuse to amortise. - **Tight container memory limits** where a growing daemon risks OOM kills and flaky failures. - **Reproducibility/clean-room** requirements outweigh a small speed gain. - You want operational simplicity: `org.gradle.daemon=false` and forget it. ## Decide with measurement ```bash # A/B the same pipeline time ./gradlew build # daemon on (default) time ./gradlew build --no-daemon # daemon off ``` Compare: wall-clock across a *sequence* of builds (not just one), **peak RSS** versus the agent's memory limit, and the count of lingering Java processes after N jobs. Let those numbers, plus your agent lifecycle, drive the choice. ## Common anti-patterns - Globally disabling everywhere including local dev, throwing away the daemon's real local benefit. - Globally *enabling* on tight CI containers and then chasing flaky OOMs. - Treating it as a moral stance instead of measuring. ## Bottom line Match daemon policy to **agent reuse** and **memory envelope**. Reused + roomy → keep and constrain. Single-shot or constrained → disable. Always validate with a benchmark rather than assuming.

  • Why might keeping the daemon on a reused runner still spawn many daemons?
    Gradle only reuses a daemon whose JVM args/version match the request. If jobs use varying jvmargs or Java versions, each spawns an incompatible new daemon, defeating reuse and bloating memory.
  • What metric tells you the daemon is dangerous on a given runner?
    Peak resident memory approaching or exceeding the container's cgroup limit, plus OOM kills or rising counts of leftover Java processes across successive jobs.
  • Why benchmark a sequence rather than a single build?
    The daemon's benefit is amortised reuse; a single cold build understates it. Timing several back-to-back builds reveals the warm-reuse speed-up the daemon actually provides.

saying these in an interview costs you the question

  • Insisting the daemon must always be off on CI with no regard for agent reuse.
  • Keeping the daemon on tight-memory containers without bounding heap or cleaning up.
  • Deciding without any measurement of time or peak memory.

context