Inside a container, your application runs as PID 1. What does the Linux kernel treat differently about PID 1, and what problems does that cause?
answer
- PID 1 ignores default-disposition signals
- docker stop = SIGTERM, wait 10s, SIGKILL
- orphans reparent to PID 1 → zombies if unreaped
- shell form makes /bin/sh PID 1; use exec form
- --init / tini reaps and forwards
basics
~20 sPID 1 is special: default signal handlers are ignored, so it will not die from SIGTERM unless it handles it, and it inherits orphaned children and must reap them. An app that does neither ignores graceful shutdown and accumulates zombie processes.
solid answer
~60 sA new PID namespace makes the first process PID 1, and the kernel gives PID 1 two special behaviours: 1. **Signals with default dispositions are not delivered.** The kernel drops SIGTERM, SIGINT and similar to PID 1 unless that process has installed a handler. So `docker stop`, which sends SIGTERM then waits, does nothing — after the timeout (10s default) the container is SIGKILLed. Symptoms: every stop takes ten seconds and in-flight work is lost. 2. **It inherits orphans and must reap them.** When any process in the namespace dies leaving children, they are reparented to PID 1, which must `wait()` on them. If it doesn't, terminated children stay as **zombies** consuming PID-table entries. A container spawning subprocesses can eventually exhaust PIDs. There is a third trap: if PID 1 is a shell (`CMD app` in shell form runs `/bin/sh -c app`), the shell may not forward signals to the app at all. Fixes: handle SIGTERM in the app, use exec form so the app *is* PID 1, and add a tiny init (`--init`, tini) when spawning children.
code
dockerfile · 14 lines# Wrong: /bin/sh becomes PID 1 and swallows SIGTERM
CMD node server.js
# Right: node is PID 1 and its own SIGTERM handler runs
CMD ["node", "server.js"]
# If a wrapper is required, end it with exec
# entrypoint.sh:
# #!/bin/sh
# set -e
# /prepare-config.sh
# exec "$@"
ENTRYPOINT ["/entrypoint.sh"]
CMD ["node", "server.js"]go deeper
Know that the main process is PID 1, that it must handle SIGTERM to stop cleanly, and that exec form is the right CMD style.
Explain both kernel rules — dropped default-disposition signals and orphan reparenting — and connect them to the ten-second stop and zombie symptoms.
Discuss the full shutdown path including grace-period tuning, when tini/--init is warranted, exec "$@" in entrypoint scripts, and diagnosing PID exhaustion.
Set a platform-wide convention: exec-form entrypoints, mandatory SIGTERM handling for graceful drain, injected init where subprocess spawning is expected, and grace periods aligned with request timeouts.
## Why PID 1 exists inside a container Creating a PID namespace gives its first process the number 1 within that namespace. The kernel then applies to it the same special rules it applies to the system's real init — because PID 1 is, by definition, the root of that process tree and its death tears the namespace down. ## Special rule 1 — signal delivery For an ordinary process, a signal with a *default* disposition (no handler installed) is acted on by the kernel: SIGTERM terminates it. For PID 1, the kernel **does not deliver** signals whose disposition is default, for SIGKILL and SIGSTOP excepted (those cannot be caught anywhere). The rationale is to prevent accidentally killing init and panicking the system. Inside a container this inverts the usual expectation. `docker stop` sends SIGTERM to PID 1, waits `--time` seconds (10 by default), then sends SIGKILL. If your application never installed a SIGTERM handler, the SIGTERM is silently dropped and the container is always killed hard. Practical consequences: shutdown takes the full grace period every time; connections are cut mid-request; buffered data, in-flight transactions, and clean deregistration from a load balancer are lost. Orchestrators behave the same way — the same grace-period-then-kill sequence applies when a Kubernetes Pod terminates. Note the asymmetry: run the same binary as a child of an init inside the container and default SIGTERM works normally, because the rule attaches to PID 1 specifically, not to the program. ## Special rule 2 — orphan reaping When a process exits, its parent must call `wait()`/`waitpid()` to collect the exit status; until then the kernel keeps a **zombie** entry (state `Z`) holding just that status and the PID. If a process dies while it still has children, those children are **reparented to PID 1** of their namespace. The real system init loops on `wait()` forever, so on a normal host zombies vanish immediately. A typical application is not written to do that. If your container spawns subprocesses — a shell script, an image converter, a `git` invocation, a language runtime forking workers — and any intermediate parent dies, the grandchildren land on your app as PID 1 and are never reaped. Zombies do not consume CPU or memory, but each holds a PID-table slot; with a PID limit (or the namespace's `pids.max`) applied, you eventually cannot fork at all and the container fails with resource-unavailable errors while looking idle. Diagnosis: `ps` inside the container shows many entries marked `defunct`/`Z` parented to PID 1. ## The shell-form trap `CMD npm start` (shell form) is executed as `/bin/sh -c "npm start"`, making `sh` PID 1 and your app its child. Now two things break: the shell is PID 1 and ignores default-disposition SIGTERM, and even if it handled it, it typically does not forward signals to children. `docker stop` becomes a guaranteed SIGKILL after the grace period. Using the **exec form** — `CMD ["npm", "start"]` — runs the program directly as PID 1, so its own signal handling applies. If a wrapper script is genuinely needed, end it with `exec "$@"` so the final program *replaces* the shell and inherits PID 1. ## Fixes, in order of preference 1. **Handle SIGTERM in the application.** Stop accepting new work, finish in-flight requests, close resources, exit. Most frameworks and runtimes offer a hook; this is the only fix that gives genuinely graceful shutdown. 2. **Use exec form** in ENTRYPOINT/CMD, or `exec` at the end of an entrypoint script, so the real process is PID 1. 3. **Add a minimal init** when the container spawns children: `docker run --init` injects one (tini), or bake `tini`/`dumb-init` in as the entrypoint. It reaps orphans and forwards signals to your process, which then still needs its own SIGTERM handling to shut down cleanly. 4. **Tune the grace period** (`docker stop -t`, or the equivalent termination grace setting in an orchestrator) so shutdown has enough time — but only after the signal is actually being received. ## Other PID-namespace properties worth knowing - **PID 1's death ends the namespace.** When it exits, the kernel SIGKILLs every remaining process in that namespace and tears it down. That is why the container stops when your main process stops, regardless of background children. - **`/proc` must be remounted** inside the namespace for `ps` to show the namespace's own view; `unshare --pid --fork --mount-proc` does exactly that, which is why the flag exists. - **Two PIDs per process.** The host sees a container process under a different, host-namespace PID; `docker top` and `/proc/<hostpid>/status` (field `NSpid`) show both. Killing from the host uses the host PID. - **Nesting is one-directional.** A parent namespace sees all descendants' processes; a child sees nothing above it.
- A container always takes exactly ten seconds to stop. What is happening?SIGTERM is reaching PID 1 and being discarded, because PID 1 has no handler for it and the kernel does not deliver default-disposition signals to PID 1. Docker then waits out its ten-second grace period and sends SIGKILL, which cannot be ignored. The fix is to install a SIGTERM handler in the application, and to make sure the application really is PID 1 by using exec form rather than a shell wrapper.
- When is `docker run --init` actually necessary?When the container's main process spawns child processes it does not itself reap — build tools, shell pipelines, forking runtimes. The injected init sits at PID 1, reaps orphaned children so zombies do not accumulate against the PID limit, and forwards signals to your process. For a single static server that spawns nothing and handles SIGTERM, it adds little.
- Does a zombie process consume memory or CPU?Essentially none — a zombie is just a kernel task entry holding an exit status and its PID. The real cost is PID-table slots, so under a PID limit enough zombies make fork() fail while the container appears idle. That is why the symptom presents as mysterious resource-unavailable errors rather than as high resource usage.
PID 1 is the building superintendent: the building's rules say you cannot simply evict them, and every abandoned box in the hallway becomes their problem to clear.
saying these in an interview costs you the question
- Claiming SIGKILL can be caught or handled by PID 1 — it cannot be, anywhere.
- Saying zombies consume large amounts of memory; they consume PID entries.
- Using shell-form CMD and expecting the application to receive SIGTERM.
- Thinking `--init` alone gives graceful shutdown without the app handling SIGTERM.
- Believing PID 1 in the container is the host's init process.