What breaks if you set TESTCONTAINERS_RYUK_DISABLED=true on CI, and when is that acceptable?
answer
- A sidecar that watches the JVM connection
- Normal shutdown was never the problem
- Cancelled and killed builds are
- Safe only when the host is discarded
- Try the privileged setting before disabling
basics
~20 sDisabling Ryuk removes the sidecar that deletes a session's containers, networks and volumes when the JVM dies, so anything a crashed or killed build leaves behind stays. It is acceptable only where the whole runner is discarded after the job.
solid answer
~40 sRyuk is the small sidecar container Testcontainers starts once per JVM session; it holds a connection back to the JVM and, when that connection drops, removes everything labelled with the session id. Setting `TESTCONTAINERS_RYUK_DISABLED=true` (or `ryuk.disabled` in `~/.testcontainers.properties`) skips it, so cleanup relies entirely on the test framework's own shutdown path. A normally finishing build still stops its containers; a build that is cancelled, times out, or has its JVM killed does not, and the leftovers accumulate. That is fine on a single-use runner VM that is destroyed after the job, and on managed environments that forbid the sidecar. On a long-lived self-hosted runner it is a slow leak, and you need an explicit per-job cleanup step instead. Restricted daemons sometimes need `ryuk.container.privileged=true` rather than disabling it outright.
code
bash · 6 lines# Opt out (only where the runner host itself is ephemeral)
export TESTCONTAINERS_RYUK_DISABLED=true
# Compensating cleanup for a long-lived runner: label-filtered, always runs
docker ps -aq --filter label=org.testcontainers=true | xargs -r docker rm -f
docker network ls -q --filter label=org.testcontainers=true | xargs -r docker network rmgo deeper
Recall that Testcontainers starts a small helper container to clean up after itself, and that turning it off means leftovers survive when a build is killed.
Explain the mechanism — a session-labelled sweep triggered when the connection to the test JVM drops — and why ordinary passing or failing builds never needed it.
Show the operational judgment: tie the decision to whether the runner host is ephemeral, and specify the label-filtered, always-run cleanup you would put in place where it is not.
Own the fleet policy on container hygiene — who guarantees reclamation, how disk and network exhaustion are monitored, and whether restricted runtimes get a privilege exception instead of a blanket opt-out.
## What you are turning off When a Testcontainers JVM starts its first container it also starts a small companion container from the `testcontainers/ryuk` image. Ryuk talks to the Docker daemon and holds an open connection back to the test JVM. Every resource Testcontainers creates — containers, networks, volumes — is tagged with labels identifying the library and that JVM's session. When the connection to the JVM drops, for any reason, Ryuk deletes everything carrying that session's labels. The key word is *any reason*. Orderly shutdown is already handled by the library's own hooks: a suite that finishes, even with failures, stops what it started. Ryuk exists for the disorderly cases — the build cancelled from the CI UI, the job killed by a step timeout, the JVM OOM-killed, the runner process terminated. Those are exactly the cases where no shutdown hook runs. ## What disabling changes `TESTCONTAINERS_RYUK_DISABLED=true`, or `ryuk.disabled=true` in `~/.testcontainers.properties`, means no sidecar starts. You save one container start per JVM and one image to pull, and you lose the crash-safety net. Concretely, on a runner where jobs are cancelled with any regularity you will accumulate stopped and running containers, orphaned bridge networks, and anonymous volumes holding database data. The consequences show up as disk exhaustion on the runner, exhausted network address space (a bridge-network limit is hit long before disk is), and port or name contention that presents as "flaky" tests in unrelated jobs. ## When it is genuinely safe One case: the environment that would hold the leak is destroyed anyway. A managed CI runner that boots a fresh VM per job and discards it afterwards has nothing to leak into — Ryuk's guarantee is subsumed by the VM lifecycle, and skipping it removes a container start from every job. Teams with large fanned-out pipelines do this deliberately. A second case is environmental rather than chosen: some restricted environments will not run the reaper. It needs access to the Docker socket, and container runtimes with different security models may require it to run with extra privileges. The setting `ryuk.container.privileged=true` in `~/.testcontainers.properties` exists for exactly that, and trying it is the right first move — disabling the reaper should be the fallback, not the reflex, because you are trading a configuration problem for an operational one. ## When it is not safe Any long-lived runner. Self-hosted machines, a fixed pool of build agents, a developer's laptop — all of these keep whatever the reaper would have removed. If policy or the environment forces the reaper off there, replace it with something explicit: a job-level cleanup step that removes resources carrying the Testcontainers labels, run in a stage that executes even when the job fails or is cancelled, plus periodic host maintenance that prunes stopped containers, unused networks and dangling volumes. Removing *all* containers unconditionally is the wrong hammer on a shared host, because other jobs may be mid-run — filter on the labels the library applies. ## How to talk about it in an interview The strong answer is not "disable it, it causes trouble". It is: name what the reaper protects against (abnormal termination, not normal shutdown), state the exact condition that makes disabling safe (the host itself is ephemeral), name the compensating control when it is not (label-filtered cleanup that runs on cancellation), and mention that a restricted runtime usually needs a privilege setting rather than removal. The weak answer treats it as a performance tweak with no downside, which is how runners fill up with a hundred stopped Postgres containers.
- If Ryuk must stay disabled on a long-lived runner, what do you put in its place?An explicit cleanup stage that always runs, including on failure and cancellation, removing containers, networks and volumes that carry the Testcontainers labels — not a blanket prune, which would destroy resources belonging to concurrent jobs. Back it with scheduled host maintenance so anything the job-level step misses is still reclaimed.
- A restricted CI environment refuses to start the reaper container. What do you try before disabling it?Setting `ryuk.container.privileged=true` in `~/.testcontainers.properties`, since the usual cause is a runtime whose security model denies the reaper the socket access it needs. Disabling is the fallback if that is not permitted, and then only alongside an explicit cleanup step.
- Does disabling the reaper affect a build that fails its assertions but completes normally?No. Normal termination — passing or failing — runs the library's own shutdown path and stops the containers it started. The reaper only matters when the JVM dies without running that path: cancellation, step timeouts, OOM kills.
saying these in an interview costs you the question
- Says disabling it is a free speed-up with no downside
- Thinks it cleans up after ordinary test failures
- Replaces it with an unfiltered prune of all containers on a shared host
- Disables it on a long-lived self-hosted runner
- Believes leaked containers only waste disk, not networks or ports