skip to content

On a Linux host, a program has finished running but still shows up in the process list as `<defunct>` in state Z. What is that entry, and what makes it go away?

level: juniorimportance: must knowfreq 70%

answer

  1. dead already, but still listed
  2. the parent owes a collection
  3. exit status has nowhere to go yet
  4. wait() frees the entry, not SIGKILL
  5. fix the parent, not the child

basics

~20 s

A zombie is a process that has already terminated but whose parent has not yet collected its exit status. The kernel keeps only its process-table entry, PID and exit code; it disappears as soon as the parent waits on it.

solid answer

~50 s

When a process terminates, the kernel tears down almost everything about it — address space, open file descriptors, its place in the run queue. What it deliberately keeps is a small task entry holding the PID, the exit status and accounting figures, because that information is owed to the parent. The parent claims it by calling `wait()`, `waitpid()` or `waitid()`, an act usually called reaping, and the entry is freed at that moment. Until then the process shows as state `Z` and `<defunct>`. A zombie consumes no memory and no CPU, so a handful is harmless; `kill -9` does nothing to it because there is no longer a running process to signal. Piles of them mean the parent is buggy — it never reaps — and the cure is to fix or restart the parent, whose orphaned zombies are then inherited and reaped by PID 1.

code

c · 14 lines
c
#include <stdio.h>
#include <unistd.h>

int main(void) {
    pid_t pid = fork();

    if (pid == 0) {
        _exit(0);              /* child finishes immediately */
    }

    printf("pid %d is now a zombie\n", (int)pid);
    sleep(60);                 /* parent never calls wait(): entry stays */
    return 0;
}

go deeper

for a junior

Be able to say plainly that a zombie is an already-finished process whose parent has not collected its exit status, that it shows as Z or <defunct>, and that signalling it does nothing.

for a middle

Explain the mechanics: what the kernel keeps versus releases at termination, the SIGCHLD notification, and how wait/waitpid/waitid frees the entry. Know that ignoring SIGCHLD makes the kernel reap automatically.

for a senior

Show you would treat a zombie pile as a symptom: identify the negligent parent, know the non-queued-signal trap in a SIGCHLD handler, and connect the leak to PID exhaustion under kernel.pid_max or RLIMIT_NPROC before it takes the host down.

for a principal

Own the design rule: any long-lived process that spawns children must have an explicit reaping strategy — a draining wait loop, SA_NOCLDWAIT, or delegation to a real init — and that requirement belongs in the platform's service checklist, not in each team's debugging folklore.

## What terminating actually does When a Linux process calls `exit()` (or is killed), the kernel releases nearly all of its resources immediately: the address space and its page tables, open file descriptors, memory locks, and its position in the scheduler. From that instant the process cannot run again and holds no memory. What survives is the small kernel structure describing the task, now marked `EXIT_ZOMBIE`, carrying three things the system still owes to somebody: the PID, the termination status (a normal exit code, or the signal that killed it), and resource-usage accounting such as CPU time consumed. `ps` renders this state as `Z` and prints the command as `<defunct>`. The word "zombie" is exact: the process is dead, but its record has not yet been filed away. ## Why the record has to stay Unix defines a contract between parent and child: a child's fate is reported to the parent that created it. If the kernel discarded the exit status the moment the child died, a parent that asked a second later would have no way to learn whether its child succeeded. So the kernel holds the status until it is collected, and simultaneously sends `SIGCHLD` to the parent to say "one of your children has ended". The parent collects it — reaps it — with one of the wait family: ```c int status; pid_t pid = waitpid(-1, &status, WNOHANG); /* -1 = any child, WNOHANG = do not block */ ``` As soon as the wait returns for that child, the kernel frees the entry and the PID becomes reusable. ## Why `kill -9` does not help Signals are delivered to a running process. A zombie has no code, no stack and no thread of control; there is nothing there to receive `SIGKILL`. Sending signals to a zombie's PID is a no-op, and this is the single most common wrong answer in interviews. The only actor who can clear it is the parent (or, indirectly, whoever removes the parent). ## The three ways a zombie is cleared 1. **The parent waits.** Normal, correct behaviour. 2. **The parent opts out of ever knowing.** If a process explicitly sets the disposition of `SIGCHLD` to `SIG_IGN`, or installs a handler with the `SA_NOCLDWAIT` flag, the kernel reaps children automatically and they never become zombies. The trade-off is that the parent can no longer learn any child's exit status. 3. **The parent dies.** Orphaned children — zombies included — are reparented to PID 1 (or to the nearest ancestor that marked itself a subreaper). An init process waits in a loop precisely so that inherited zombies are reaped immediately. This is why "restart the parent service" makes a wall of zombies vanish. ## When zombies actually matter Each zombie occupies a PID. PIDs are a finite resource bounded by `kernel.pid_max`, and per-user process counts are bounded by `RLIMIT_NPROC` (`ulimit -u`). A supervisor that spawns a worker per minute and never reaps will, over weeks, exhaust one of those limits, and then *nothing* on the machine can fork — a failure that looks nothing like the original bug. So a growing zombie count is a leak indicator, not a problem in itself. ## The bug that produces them Almost always: a long-running parent that forks children and never waits. A close second is a `SIGCHLD` handler that calls `wait()` exactly once per signal. Standard signals are not queued — if three children die while one `SIGCHLD` is pending, the parent receives one signal, reaps one child and leaves two zombies forever. The correct shape is to loop: ```c while (waitpid(-1, &status, WNOHANG) > 0) ; /* drain every child that has already finished */ ``` ## What this is not A zombie is not an orphan. An orphan is a *live* process whose parent has exited; it keeps running quite happily under a new parent. A zombie is a *dead* process whose parent is still alive but negligent. Candidates who blur those two usually also believe zombies burn CPU — they do not.

  • If a zombie holds no memory and burns no CPU, why would anyone treat a rising zombie count as an incident?
    Because each one pins a PID. PIDs are capped by `kernel.pid_max` and per-user process counts by `RLIMIT_NPROC`, so a parent that leaks zombies steadily will eventually make fork fail — for every process on the box, not just the guilty one. The count is a leak gauge for a broken parent, which is the real defect.
  • A parent installs a SIGCHLD handler that calls wait() once, yet zombies still accumulate. Why?
    Standard signals are not queued. If several children exit while a `SIGCHLD` is already pending, the parent still gets one signal, reaps one child, and the rest stay zombies indefinitely. The handler must loop on `waitpid(-1, &status, WNOHANG)` until it returns zero or -1, draining every child that has already finished.
  • How can a parent that genuinely does not care about exit statuses avoid creating zombies at all?
    Set the disposition of `SIGCHLD` to `SIG_IGN`, or install a handler with `SA_NOCLDWAIT`. The kernel then discards each child's status at termination instead of holding it, so children are reaped automatically. The cost is exactly what it says: the parent permanently gives up the ability to learn how any child ended.

A zombie is a death certificate sitting in the registry: the person is gone, but the paperwork stays on file until the next of kin comes to collect it.

saying these in an interview costs you the question

  • Says kill -9 clears a zombie process
  • Claims zombies consume memory or CPU time
  • Confuses a zombie with an orphaned background process
  • Thinks the kernel reaps zombies after a timeout
  • Reboots the machine instead of fixing the parent

context