skip to content

questions

5

Walk through exactly what happens, step by step, when you run `docker stop` on a running container, and how that sequence differs from `docker kill`.

level: juniorimportance: must knowfreq 70%

answer

  1. SIGTERM → wait 10s → SIGKILL
  2. Signal goes to PID 1 only
  3. -t / --timeout changes the grace window
  4. docker kill = SIGKILL now, no grace
  5. Exit 137 = 128 + 9 = killed

basics

~20 s

docker stop sends SIGTERM to the container's PID 1, waits a grace period (10 seconds by default, changed with -t), then sends SIGKILL if the process is still alive. docker kill skips the wait and sends SIGKILL immediately.

solid answer

~50 s

`docker stop` is a two-phase shutdown. The daemon sends the container's stop signal — SIGTERM unless the image's Dockerfile `STOPSIGNAL` or `docker stop --signal` says otherwise — to **PID 1 inside the container only**, not to every process. It then waits the grace period, 10 seconds by default, overridable per call with `-t/--timeout`. If PID 1 has exited by then the container goes to `exited` with that process's exit code. If not, the daemon sends SIGKILL, which the kernel delivers uncatchably, and the container is recorded as exiting with 137 (128 + 9). `docker kill` sends SIGKILL right away (or a signal you choose with `-s`) with no grace window. Use `stop` in normal operation so the app can drain connections, flush buffers and close files; use `kill` only when the process is wedged and you have accepted losing in-flight work.

code

bash · 5 lines
bash
docker stop -t 30 api
docker kill api
docker kill -s HUP nginx
time docker stop api
docker inspect -f '{{.State.ExitCode}} {{.State.Status}}' api

go deeper

for a junior

Know the sequence cold: SIGTERM, default 10-second wait, then SIGKILL; docker kill skips straight to SIGKILL. Mention that the signal goes to PID 1.

for a middle

Add the configurability (-t, --signal, Dockerfile STOPSIGNAL, daemon --shutdown-timeout) and explain why an app that ignores SIGTERM always takes the full window and exits 137.

for a senior

Frame it as the shutdown contract: what the app must do inside the window (stop accepting work, drain, flush, exit) and what data loss looks like when SIGKILL wins.

for a principal

Tie the grace period to fleet-level behavior — deploy velocity, in-flight request loss, queue redelivery — and argue for a single consistent timeout policy across local Docker and the orchestrator.

## The shape of container shutdown A container is not a virtual machine you power off — it is a process tree, and stopping it means asking the process at the root of that tree to end. Docker models that as a two-phase handshake: a polite request first, a fatal one after a deadline. ## Phase 1 — the stop signal When you run `docker stop <container>`, the Docker daemon (through containerd and the OCI runtime) sends a signal to **PID 1 in the container's PID namespace** — the single process Docker started from your `ENTRYPOINT`/`CMD`. It does not fan the signal out to child processes; anything else running in the container is that process's responsibility. The signal sent is, in priority order: whatever `docker stop --signal` names on this invocation, otherwise the image's Dockerfile `STOPSIGNAL` instruction, otherwise SIGTERM. SIGTERM (signal 15) is the conventional "please terminate" request: a process can install a handler for it and use the moment to stop accepting new work, finish or abandon in-flight work, flush buffers, close database connections and then exit. ## Phase 2 — the grace period and SIGKILL After sending the stop signal the daemon waits. The default wait is **10 seconds**. You change it per invocation with `docker stop -t 30 <container>`, and you can change the daemon-wide default with `--shutdown-timeout` (which also governs how long containers get when the daemon itself shuts down). Two outcomes: - **PID 1 exits inside the window.** The container transitions to `exited` and records PID 1's own exit code — 0 for a clean exit, or whatever code the app chose. - **PID 1 is still alive when the window closes.** The daemon sends **SIGKILL** (signal 9). SIGKILL cannot be caught, blocked or ignored; the kernel destroys the process. The container is recorded as exiting with **137**, which is the shell convention 128 + signal number (128 + 9). A SIGTERM-killed process would similarly show 143. When PID 1 dies, the kernel tears down the PID namespace and kills every remaining process in it, so orphaned children do not survive the container. ## Why the difference matters in production SIGKILL gives the process zero opportunity to clean up. Anything buffered in userspace is lost — an unflushed log batch, a partially written file, an HTTP response half-sent. A database engine killed with 9 typically has to run crash recovery on next start. A queue consumer killed with 9 leaves messages in-flight until the broker's visibility timeout expires and redelivers them. That is exactly why the graceful phase exists, and why an app that ignores SIGTERM is a real defect: it converts every deploy and every restart into a crash. ## docker kill `docker kill <container>` sends SIGKILL immediately with no grace period. It also doubles as a general signal-sending tool: `docker kill -s HUP nginx` sends SIGHUP to PID 1, which many servers use for config reload. Note the asymmetry — `docker kill` defaults to SIGKILL and ignores `STOPSIGNAL`, while `docker stop` honors it. `docker restart` is `stop` followed by `start` and takes the same `-t` flag, so it inherits the same grace semantics. `docker rm -f` on a running container also goes through a kill path rather than a graceful stop. ## Common traps 1. **The signal only reaches PID 1.** If your entrypoint is a shell script that launches the real server as a child, SIGTERM hits the shell, not the server. Use `exec` in the script, or the exec form of `ENTRYPOINT`, so the real process *becomes* PID 1. 2. **PID 1 has no default SIGTERM behavior.** The Linux kernel does not apply the usual default action to signals for PID 1, so a program with no explicit SIGTERM handler simply ignores it and always eats the full 10 seconds before SIGKILL. 3. **Ten seconds is a policy, not a law.** A service that legitimately needs to drain 30-second requests must be given a longer timeout everywhere it is stopped, and its shutdown code must be bounded so it actually finishes inside that window. 4. **A slow `docker stop` is a symptom.** If stopping consistently takes exactly the timeout, the container is being SIGKILLed every time; treat that as a bug to fix, not background noise. ## Verifying behavior Stop a container and time it: `time docker stop app`. Sub-second means the app handled the signal. Bang on 10 seconds means it did not. Then confirm with `docker inspect -f '{{.State.ExitCode}}' app` — 137 confirms SIGKILL.

  • What exit code does a container show after Docker SIGKILLs it, and why that number?
    137. The shell convention for a process terminated by a signal is 128 plus the signal number, and SIGKILL is signal 9. A container killed by SIGTERM would show 143 (128 + 15). Seeing 137 consistently after `docker stop` is strong evidence the app never handled SIGTERM and always burned the full grace period.
  • Does `docker stop` signal every process inside the container?
    No. Only PID 1 in the container's PID namespace receives the signal. Propagating it to children is the application's or an init process's job. When PID 1 finally exits, the kernel tears down the PID namespace and kills whatever is left, but those processes get no chance to clean up.

docker stop is the fire alarm: everyone gets a set number of minutes to grab their laptop and leave. docker kill is the building being demolished with people still inside.

saying these in an interview costs you the question

  • Saying docker stop sends SIGKILL, or that stop and kill are the same thing
  • Believing the signal is broadcast to all processes in the container
  • Claiming SIGKILL can be caught or handled to run cleanup code
  • Thinking the 10-second grace period is fixed and cannot be changed
  • Treating a stop that always takes exactly the timeout as normal rather than as a bug

context

open as a page

A service built with `CMD npm start` in its Dockerfile never reacts to `docker stop` — it always takes the full grace period and dies hard. Explain why the shell form of CMD/ENTRYPOINT causes this, and show how you would fix it.

level: middleimportance: must knowfreq 62%

basics

~20 s

Shell form wraps the command in /bin/sh -c, so the shell is PID 1 and your app is a child. Docker signals PID 1 — the shell — which does not forward SIGTERM. Use exec form (CMD ["npm","start"]) or exec in the entrypoint script so the app is PID 1.

open as a page

Why does a program running as PID 1 inside a container often ignore SIGTERM entirely, even though the same binary shuts down correctly when you run it on a normal Linux host?

level: seniorimportance: must knowfreq 48%

basics

~20 s

The Linux kernel treats PID 1 specially: signals with no explicitly installed handler are discarded instead of applying their default action. On a host your process is not PID 1, so SIGTERM's default action (terminate) applies. As PID 1, no handler means nothing happens.

open as a page

What is a zombie process, why do they pile up inside containers in particular, and what do Docker's `--init` flag and the tini init process do about it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

A zombie is a dead process whose exit status no parent has collected with wait(). Orphans are re-parented to PID 1, which must reap them. Application processes running as container PID 1 usually never call wait(), so zombies accumulate and can exhaust the PID table. --init inserts tini as PID 1, which reaps them.

open as a page

You own a containerized service that both serves HTTP requests taking up to 20 seconds and consumes messages from a queue. Design its shutdown path: what should the process do on SIGTERM, and how do you choose the stop timeout it runs under?

level: principalimportance: should knowfreq 34%

basics

~20 s

On SIGTERM: fail readiness first, stop accepting new HTTP connections and stop pulling from the queue, finish or abandon in-flight work under a bounded deadline shorter than the stop timeout, flush and close, exit 0. Set the stop timeout above your worst realistic drain, not above your worst theoretical one.

open as a page