How does Testcontainers make sure its containers are removed if the test JVM crashes?
answer
- Crashes skip shutdown hooks entirely
- One extra container per session
- A held-open socket is the signal
- Session labels tell it what to delete
- The sidecar is named Ryuk
basics
~20 sTestcontainers starts a sidecar container called Ryuk that holds an open socket to the test JVM and labels every resource it creates. When that socket closes — including after a crash or kill -9 — Ryuk deletes everything carrying the session's labels.
solid answer
~50 sCleanup cannot rely on the test process behaving well, because a `kill -9`, an IDE stop button or an OOM kill skips shutdown hooks and JUnit callbacks entirely. So Testcontainers starts one extra container per session, `testcontainers/ryuk`, with the Docker socket mounted into it. The JVM opens a TCP connection to Ryuk's mapped port and sends it a filter describing the labels Testcontainers stamps on every container, network and volume it creates for this session. Ryuk holds that connection open as a dead-man's switch: the moment it breaks, it asks Docker to remove every resource matching those labels. Normal shutdown still goes through the extension's own `stop()` call — Ryuk is the safety net for abnormal exits, which is why you see a second container in `docker ps` next to your Postgres or Kafka one.
code
console · 4 lines$ docker ps --format 'table {{.Image}}\t{{.Ports}}\t{{.Names}}'
IMAGE PORTS NAMES
postgres:16-alpine 0.0.0.0:32770->5432/tcp nostalgic_bohr
testcontainers/ryuk 0.0.0.0:32769->8080/tcp testcontainers-ryuk-8f2c1ago deeper
Recall that Testcontainers starts one extra container you did not ask for, that it is called Ryuk, and that its job is cleaning up leftovers. Do not be alarmed by it in docker ps.
Explain the mechanism: session labels on every created resource plus a socket held open from the JVM, with the socket closing as the trigger. Be able to say why a shutdown hook is not enough.
Show you have debugged this in the wild — orphaned containers filling a build agent, and the diagnosis that the reaper could not reach the Docker socket. Distinguish insurance from lifecycle discipline.
Own the policy question: whether build agents may run a privileged reaper at all, what replaces it when they may not, and who owns disk and port exhaustion on shared CI hardware.
## The problem Ryuk exists to solve A container started by a test outlives the test process by default. If the JVM exits normally, Testcontainers can stop what it started — the JUnit integration calls `stop()`, and there is also a JVM shutdown hook. But test runs die in ways that skip all of that: you hit the stop button in the IDE, CI cancels the job, the kernel OOM-kills the JVM, or the machine loses power. None of those run Java code. Without a second mechanism, a developer laptop or a CI agent slowly fills with orphaned Postgres, Kafka and Redis containers, along with the networks and volumes they were attached to, until Docker runs out of disk or ports. ## What Ryuk is Ryuk (image `testcontainers/ryuk`, project name moby-ryuk) is a tiny Go daemon that Testcontainers starts once per session, before your first container. It runs as a container itself, with the host's Docker socket mounted into it so it can talk to the same daemon your test containers live on, and it publishes its own port to a random host port. When Testcontainers creates any resource — a container, a user-defined network, a volume, or an image built on the fly — it stamps session-scoped labels on it. It then opens a TCP connection to Ryuk and sends a filter expressing "everything with these labels". Ryuk acknowledges the filter and then does nothing except hold the connection open. ## The connection is the heartbeat This is the key design idea, and the thing interviewers want you to articulate: **the open socket is the liveness signal**. Ryuk does not poll, does not run on a timer keyed to your test duration, and does not need the JVM to announce that it is finishing. The operating system closes the socket whenever the JVM process disappears, for any reason at all — clean exit, uncaught error, SIGKILL, container-of-the-build-being-torn-down. Ryuk sees the connection drop and issues the delete calls against the Docker API for everything matching the filter it was given. Because the signal is the absence of a process rather than an action by a process, it survives exactly the failures that shutdown hooks do not. A shutdown hook is code you have to be alive to run; a closed socket is something you cannot avoid emitting once you are dead. ## Ryuk versus the normal path On a healthy run both paths fire, harmlessly. The JUnit integration stops containers at the end of the class or the method, so the resources are already gone by the time the JVM exits and Ryuk finds nothing left to reap. Ryuk removes itself afterwards. Treat Ryuk as insurance, not as your primary lifecycle: a test that leaks containers for minutes inside a long run is still a leak, because Ryuk only acts when the whole process ends. Containers explicitly marked as reusable are deliberately kept out of session-scoped reaping — the entire point of a reusable container is to survive the run that created it, so it is not labelled for the session reaper. ## Requirements and failure modes Ryuk needs to talk to the Docker daemon, which means the socket must be reachable from inside a container. In hardened or rootless environments that is sometimes not permitted, and Testcontainers offers a switch to run without the reaper; the trade-off is that abnormal exits then leak, and you take on cleanup as an operational concern. The other thing worth knowing is that Ryuk is per session, not per container — one reaper watches everything a single JVM created, which is why you see exactly one extra container regardless of how many services your suite starts. ## What to say in an interview A strong answer names the sidecar, explains that cleanup is driven by a dropped socket rather than by cooperative shutdown, notes that it therefore covers `kill -9` and cancelled CI jobs, and adds that labelling is what lets one reaper find every container, network and volume belonging to that session. A weak answer says "Testcontainers uses a shutdown hook" and stops there — which is true of the happy path and useless in the case the question is actually about.
- If Ryuk is the safety net, why does the JUnit integration still call stop() at all?Because the reaper only acts when the whole JVM exits. Inside a long run you want each container released as soon as its scope ends, so ports, memory and disk are freed for the rest of the suite. Relying on Ryuk alone would keep every container of the run alive until the very end.
- How does one Ryuk instance know which containers belong to it?Testcontainers stamps session-scoped labels on every container, network, volume and locally built image it creates, and sends Ryuk a filter matching those labels. Ryuk deletes only resources matching the filter, so containers you started by hand, or from another parallel JVM, are untouched.
- Why is a container marked for reuse not deleted by the reaper?Reuse exists precisely so a container survives the run that started it and is picked up by the next one. Such containers are kept outside the session's reaping scope; the trade-off is that you become responsible for removing them yourself.
Ryuk is a dead-man's switch: the train keeps running only while a hand is on the lever, and the hand coming off — for any reason, including the driver collapsing — is what triggers the stop.
saying these in an interview costs you the question
- Says a JVM shutdown hook covers kill -9
- Thinks Docker --rm removes the containers
- Believes Ryuk polls Docker on a timer
- Claims one Ryuk container per test container
- Assumes JUnit callbacks run when CI cancels a job