What is a zombie process, why do they pile up inside containers in particular, and what do Docker's `--init` flag and the tini init process do about it?
answer
- Zombie = exited, status not yet wait()ed
- Orphans re-parented to PID 1, init must reap
- Apps as PID 1 do not reap → PID table fills
- --init inserts tini: reap + forward signals
- ps STAT Z / <defunct> to diagnose
basics
~20 sA zombie is a dead process whose exit status no parent has collected with wait(). Orphans are re-parented to PID 1, which must reap them. Application processes running as container PID 1 usually never call wait(), so zombies accumulate and can exhaust the PID table. --init inserts tini as PID 1, which reaps them.
solid answer
~60 sWhen a Linux process exits, the kernel keeps a small entry holding its exit status until the parent calls `wait()`/`waitpid()`. Until then it is a **zombie** — no memory, no CPU, just a PID table slot. If the parent dies first, the orphan is re-parented to PID 1, whose job is to reap it in a loop. In a container, PID 1 is your application, and applications almost never implement a reaping loop. So any process orphaned inside the container — a shell one-liner that backgrounded something, a subprocess spawned by a language runtime, a `docker exec` session whose command forks — becomes a permanent zombie. They consume PIDs, and with a `--pids-limit` or a full kernel PID table you eventually cannot fork at all. `docker run --init` inserts Docker's bundled **tini** as PID 1. tini calls `waitpid(-1, ...)` in a loop to reap any orphan, and forwards signals to your process, which is now PID 2 and gets normal default signal behavior. You can also bake `ENTRYPOINT ["/sbin/tini","--"]` into the image or use dumb-init.
code
bash · 3 linesdocker exec app ps -eo pid,ppid,stat,comm | awk '$3 ~ /Z/ {print}'
docker exec app sh -c 'ls -d /proc/[0-9]* | wc -l'
docker run -d --pids-limit 200 --name app myimagego deeper
Know the definition — a finished process whose exit status has not been collected — and that docker run --init is the standard fix in containers.
Explain re-parenting to PID 1, why application processes do not reap, and what --init/tini does (reap plus forward signals).
Connect it to the real failure: PID exhaustion under --pids-limit presenting as fork failures, and show the ps STAT Z diagnosis path.
Decide it as policy — init in base images or platform defaults, per-container PID limits with alerting — and weigh it against the deeper question of why a container runs multiple processes at all.
## What a zombie actually is On Linux, when a process terminates the kernel does not immediately discard everything about it. It retains a minimal task entry holding the exit status, resource usage and PID, so the **parent** can retrieve it with `wait()`, `waitpid()` or `waitid()`. Between termination and that call, the process is in state `Z` — a *zombie* (or *defunct*). It holds no memory and consumes no CPU; it occupies a PID and a task-struct slot. Once the parent reaps it, the entry disappears. If the parent never calls wait, the zombie stays forever — until the parent itself exits, at which point the zombie is re-parented to PID 1 and reaped there. ## Orphans and the init contract If a parent dies while children are still running, those children are **orphans** and the kernel re-parents them to the init process of their PID namespace. When they later exit, their zombies land on init's doorstep. That is why every real init — sysvinit, systemd, launchd — runs a reaping loop: ```c while ((pid = waitpid(-1, &status, 0)) > 0) { /* discard */ } ``` This is a hard requirement of being PID 1, not an optional nicety. ## Why containers hit this A container has its own PID namespace, and PID 1 is whatever you set as ENTRYPOINT — a JVM, a Node process, an nginx, a Python script. None of those were written to be init. They will reap *their own* direct children if they use a subprocess API that waits, but they will not reap arbitrary orphans re-parented to them. Orphans arise more often than people expect: - An entrypoint shell script backgrounds something (`some-daemon &`) and then execs, or exits. - A language runtime shells out and the intermediate `sh` exits while its grandchild lives on. - `docker exec` runs a command that forks and returns before the child finishes; the child is re-parented to container PID 1. - Sidecar-ish helper scripts, cron-like loops, and supervisor-less multi-process containers. Each unreaped exit leaves a zombie. In a long-lived container that starts subprocesses per request or per minute, the count climbs monotonically. ## What breaks Zombies do not leak memory, which is why people dismiss them. What they leak is **PIDs**. The kernel's PID space is bounded (`/proc/sys/kernel/pid_max`), and Docker lets you bound it per container with `--pids-limit` (containers under Kubernetes often have a pod PID limit too). When the limit is reached, `fork()`/`clone()` fails with EAGAIN and everything falls apart in confusing ways: the app cannot spawn threads, `docker exec` fails, health checks that shell out start failing. Zombie exhaustion presents as "resource temporarily unavailable" and is easy to misdiagnose as memory pressure. ## Diagnosis ``` docker exec c ps -eo pid,ppid,stat,comm | awk '$3 ~ /Z/' docker exec c sh -c 'ls /proc | grep -c "^[0-9]"' ``` Processes in state `Z` with `<defunct>` in the command column are zombies; a growing count with PPID 1 is the signature. `docker stats` will show the container's PID count if you have pids cgroup accounting. ## tini and --init `docker run --init` runs a tiny init binary (Docker ships **tini** as `docker-init`) as PID 1. Your ENTRYPOINT becomes its child at PID 2. tini does exactly two jobs: 1. **Reap**: loop on `waitpid(-1, ...)` and discard statuses of any re-parented orphan. 2. **Forward signals**: relay SIGTERM and friends to the child (optionally to the whole process group with `-g`), so `docker stop` still works — and because the child is no longer PID 1, the kernel's PID 1 exemption no longer applies, so even a program with no SIGTERM handler now terminates on the default action. tini also propagates the child's exit code, so `docker inspect` still shows a meaningful status. Alternatives: bake it in with `ENTRYPOINT ["/sbin/tini","--","/app/start"]`, use `dumb-init`, or set `init: true` in a Compose service. Full supervisors (s6-overlay, supervisord, systemd-in-a-container) also reap, but they bring a lot more machinery and encourage multi-process containers. ## When you do and do not need it You do **not** need an init if your PID 1 never spawns subprocesses that can be orphaned — a static Go binary serving HTTP, for example. You **do** want it when the container shells out, when PID 1 is a third-party binary you cannot inspect, or when the image is a base others will build on. Many teams simply set `--init` (or `init: true`) everywhere: the cost is one tiny extra process, and the failure mode it prevents is obscure and slow-burning. What `--init` does **not** do is give you graceful shutdown semantics. tini delivers the signal; only your application can drain connections and flush state.
- Do zombie processes leak memory, and if not, what resource actually runs out?They do not leak meaningful memory — a zombie has released its address space and holds only a small kernel task entry. What they exhaust is the PID space: the kernel's global `pid_max` and, more commonly, a per-container `--pids-limit`. Once exhausted, `fork()` fails with EAGAIN and the container cannot start threads, run health checks, or accept `docker exec`.
- Besides reaping, what else does tini do that matters for container lifecycle?It forwards signals to its child, so `docker stop` still reaches the application. Crucially, because the app is now PID 2 rather than PID 1, the kernel's PID 1 exemption no longer applies and default signal actions work again — so even a binary with no SIGTERM handler terminates promptly instead of waiting for SIGKILL. tini also propagates the child's exit code.
- When is `--init` unnecessary?When PID 1 never spawns processes that can be orphaned — a single static binary that serves requests without shelling out, for example — or when the entrypoint is already a proper init or supervisor that reaps. Even then many teams enable it by default, since the overhead is one tiny process and the failure it prevents is slow and hard to diagnose.
A zombie is a death certificate nobody has filed. It takes no space in the morgue — but the filing cabinet of case numbers is finite, and when it fills, no new births can be registered.
saying these in an interview costs you the question
- Saying zombies consume CPU or large amounts of memory
- Thinking `kill -9` on a zombie removes it (the process is already dead; only the parent's wait() clears it)
- Believing every container needs a full supervisor like supervisord to avoid zombies
- Assuming --init gives graceful shutdown, not just signal forwarding and reaping
- Not knowing orphans are re-parented to PID 1 of the PID namespace