skip to content

A CI agent intermittently OOMs and a stale daemon is suspected. How would you use Gradle's daemon-management commands to diagnose and harden the setup?

level: seniorimportance: should knowfreq 30%

answer

  1. --status: spot piled-up IDLE daemons
  2. --stop then re-check
  3. ephemeral → org.gradle.daemon=false
  4. shared → lower idletimeout + cleanup --stop
  5. version bump → orphaned daemons need OS kill

basics

~10 s

Use --status to see lingering daemons and their state, --stop to clear them between jobs, and harden by disabling the daemon (org.gradle.daemon=false) or shortening the idle timeout on long-lived agents.

solid answer

~40 s

First diagnose: run `gradle --status` on the agent to confirm whether old daemons are accumulating (multiple IDLE entries, growing uptime) — a classic cause of cross-build memory pressure on **reused** CI hosts. Clear them with `gradle --stop` and confirm with another `--status`. Then harden based on the agent model. For **ephemeral** agents (fresh container per job) the daemon buys nothing and can leak heap across builds, so set `org.gradle.daemon=false` (or pass `--no-daemon`). For **long-lived shared** agents where warmth is valuable, keep the daemon but lower `org.gradle.daemon.idletimeout` so abandoned daemons reclaim memory promptly, and run `--stop` in a job-cleanup step. Remember `--stop`/`--status` are per Gradle version, so a version bump can leave orphaned old-version daemons that need an OS kill or directory cleanup of `GRADLE_USER_HOME/daemon`.

code

bash · 10 lines
bash
# diagnose
gradle --status

# reset between jobs
gradle --stop && gradle --status

# ephemeral CI hardening (gradle.properties):
#   org.gradle.daemon=false
# shared-agent hardening:
#   org.gradle.daemon.idletimeout=300000   # 5 min

go deeper

for a junior

Know to run --status and --stop to inspect/clear daemons.

for a middle

Add the ephemeral-vs-shared decision and the relevant properties.

for a senior

Reason about per-version registry orphans, idletimeout tuning, cleanup steps, and that the daemon may be a symptom not the root cause.

for a principal

Define an org-wide CI daemon policy (image topology, version pinning, memory budgets) and codify it in shared properties/templates.

## Framing the failure Intermittent OOMs on CI that *correlate with agent reuse* strongly suggest **daemon state surviving across jobs**: a daemon kept alive between builds retains heap, and a misconfigured `-Xmx` or a leak means each successive build starts closer to the ceiling until one tips over. The daemon-management commands are your first-line diagnostics. ## Step 1 — observe ``` gradle --status ``` Look for: several daemons of the same version, long uptimes, multiple IDLE entries. That pattern = daemons piling up rather than being reused/cleaned. Because `--status` is **per version**, also check whether a recent Gradle upgrade left orphaned old-version daemons (they won't show under the new version). ## Step 2 — clear and confirm ``` gradle --stop gradle --status # expect an empty/clean list for this version ``` For orphaned **other-version** daemons that `--stop` won't touch, drop to the OS (`kill <pid>`) or, as a blunt reset, remove `GRADLE_USER_HOME/daemon/<oldVersion>`. ## Step 3 — harden, by agent topology **Ephemeral agents (container discarded per job):** ```properties # gradle.properties org.gradle.daemon=false ``` The daemon's warmth can't be reused across discarded containers, and disabling it removes the cross-build leak surface entirely. Equivalent ad-hoc: `--no-daemon`. **Long-lived shared agents (warmth is worth keeping):** ```properties org.gradle.daemon=true org.gradle.daemon.idletimeout=300000 # 5 min: reclaim memory fast ``` plus a pipeline cleanup step running `gradle --stop`. Right-size `org.gradle.jvmargs` (`-Xmx`) so a single daemon can't grow unbounded. ## Step 4 — make it durable - Pin a known Gradle version (wrapper) so you don't silently accumulate multi-version daemons. - Add `--status` output to CI logs on failure for forensics. - Document the policy so it survives team turnover. ## The key mental model `--status` = observe, `--stop` = reset, `idletimeout` / `org.gradle.daemon=false` = prevent recurrence. The right hardening depends entirely on whether the host is throwaway or shared — there is no single correct answer, and saying so is what separates a senior response.

  • After a Gradle version upgrade, `--stop` doesn't clear an old daemon. Why?
    `--stop` and `--status` operate per Gradle version via the per-version registry. An old-version daemon is invisible to the new version; you OS-kill it or clean its `daemon/<oldVersion>` directory.
  • For a container that's destroyed after each build, is tuning idletimeout worthwhile?
    No. The container takes the daemon with it; better to disable the daemon (`org.gradle.daemon=false`) so there's no cross-build state at all.

saying these in an interview costs you the question

  • Prescribing one fix for all CI without distinguishing ephemeral vs. shared agents.
  • Assuming `gradle --stop` clears daemons of every Gradle version on the host.
  • Treating the daemon as the OOM root cause rather than investigating heap sizing/leaks too.

context