Two services are stopped the same way. One records exit code 143 and the other records 0. What accounts for the difference, and which outcome do you want in production?
answer
- `docker stop` = SIGTERM → grace period → SIGKILL
- 143 = no handler, default kill; 0 = handled and exited cleanly
- 137 on stop = shutdown outran the grace period
- `exec` in entrypoint so the app is PID 1 and gets the signal
- Bound the drain; size `-t`/`stop_grace_period` to measured reality
basics
~20 s143 means the process was killed by SIGTERM's default action — it ran no shutdown code. 0 means the process caught SIGTERM, shut down on its own terms, and exited successfully. In production you want 0: it proves graceful shutdown actually executed.
solid answer
~60 s`docker stop` sends SIGTERM first. What happens next is entirely the application's choice: - **No handler** → the kernel's default action for SIGTERM kills the process, and Docker records `128 + 15 = 143`. Nothing was flushed or drained. - **A handler that shuts down and returns normally** → the process exits with its own status, normally `0`. In-flight requests were completed, connections closed, offsets committed. So the exit code is a cheap, reliable *assertion about graceful shutdown*: if your service records 143 on every deploy, it has no shutdown path, whatever the code review claimed. Why it matters: without a handler, a rolling deploy drops in-flight requests, leaves locks or leases held until they expire, and can lose buffered writes. Caveats: 143 is harmless for a genuinely stateless process with nothing to finish; and a shutdown handler that runs too long is worse than none, because `docker stop` escalates to SIGKILL after the grace period (10s by default) and you get 137 instead — cleanup cut off halfway.
code
bash · 11 lines# No handler: dies on SIGTERM's default action
docker run -d --name a alpine sleep 600
docker stop a && docker inspect --format '{{.State.ExitCode}}' a # 143
# Handler that exits cleanly
docker run -d --name b alpine sh -c 'trap "echo draining; exit 0" TERM; while :; do sleep 1; done'
docker stop b && docker inspect --format '{{.State.ExitCode}}' b # 0
# Handler that ignores the signal: escalated to SIGKILL after the grace period
docker run -d --name c alpine sh -c 'trap "" TERM; while :; do sleep 1; done'
docker stop -t 3 c && docker inspect --format '{{.State.ExitCode}}' c # 137go deeper
Say that 143 means SIGTERM killed the process outright while 0 means it shut itself down cleanly after catching the signal.
Explain the stop sequence (SIGTERM → grace period → SIGKILL) and map 0/143/137 onto its three outcomes, including the exec/PID-1 gotcha.
Tie the codes to production impact — dropped requests, unreleased locks, redelivered messages — and treat a stop-time 137 as partial cleanup, worse than no cleanup.
Make it a platform contract: measured grace periods, bounded drains, CI assertions that services exit 0 on stop, and alerting that separates intentional stops from force kills.
## The stop sequence and where the code comes from `docker stop <container>` performs a two-phase termination of PID 1 inside the container: 1. Send the stop signal — SIGTERM unless the image or run command set a different `STOPSIGNAL`. 2. Wait a grace period (10 seconds by default; `docker stop -t N`, or `stop_grace_period` in Compose). 3. If the process is still alive, send SIGKILL. The recorded exit code is a direct readout of which branch you landed in: | Recorded code | What happened | |---|---| | `0` | The process handled SIGTERM, did its shutdown work, and exited successfully within the grace period | | non-zero, below 128 | It handled SIGTERM but its shutdown path failed or chose a failure status | | `143` | It did **not** handle SIGTERM; the kernel's default disposition terminated it | | `137` | It neither exited nor finished in time; Docker escalated to SIGKILL | ## Why 143 happens SIGTERM's default action is "terminate". A program that never installs a handler is killed by the kernel the instant the signal is delivered — no destructors, no `finally` blocks, no flush. Docker reports the signal death as `128 + 15`. There is a second, sneakier cause of 143 (and of stop signals seeming to do nothing): the signal only reaches PID 1. If your entrypoint is a shell script that launches the real server as a child without `exec`, the shell is PID 1, and depending on how it handles signals the server may never be told to stop at all. Using `exec` in the entrypoint, or `ENTRYPOINT ["/app/server"]` in exec form, makes the application itself PID 1 so the signal lands where the handler lives. ## Why 0 is the target A service that exits 0 on stop is telling you its shutdown path ran to completion. Concretely, a graceful path usually does some or all of: - stop accepting new connections / deregister from load balancing; - finish in-flight requests up to a bounded deadline; - flush buffered writes, logs, and metrics; - commit consumer offsets or acknowledge messages so they aren't redelivered; - release locks, leases, and advisory database sessions rather than waiting for them to expire. Skipping that during a routine rolling deploy is exactly how you get "a handful of 502s every release", duplicate message processing, or a hot lock nobody holds for the next 60 seconds. Since deploys are frequent and voluntary, this failure mode repeats forever until fixed. ## Why 143 is sometimes fine Not every workload needs a handler. A stateless batch step that only reads input and writes to an idempotent sink, or a sidecar that holds no state, loses nothing when SIGTERM kills it outright. The judgement call is about *what is in flight at the moment of the signal*, not about ideology. Treat 143 as "no shutdown work ran — is that acceptable here?" rather than as a defect by default. ## The failure mode worse than both A handler that takes longer than the grace period is the trap. You get partial cleanup — say, listeners closed and connections drained halfway — followed by SIGKILL at the 10-second mark, and the recorded code is 137. That is worse than 143 because the state is now indeterminate rather than simply untouched. Two disciplines prevent it: - **Bound the shutdown**: give the drain an internal deadline comfortably inside the grace period, then exit. - **Size the grace period to reality**: measure how long shutdown actually takes under load and set `-t` / `stop_grace_period` above it. ## Using exit codes as a check Because the three outcomes are so distinguishable, the exit code doubles as a lightweight test: - In CI, start the container, send a stop, and assert the recorded code is 0. That catches a regression where someone removes the handler or reintroduces a non-`exec` entrypoint wrapper. - In production, alert on 137s that follow an intentional stop — they mean the grace period is being exceeded. - Watch for services that changed from 0 to 143 after a base-image or entrypoint refactor; that transition is almost always accidental. One more nuance worth stating: some runtimes deliberately re-raise SIGTERM after cleanup so the process appears to die from the signal, producing 143 *with* cleanup having run. It's uncommon, but if a service you know shuts down cleanly reports 143, check whether that is deliberate before filing a bug.
- Your service has a SIGTERM handler in code, yet the container still records 143. What would you check first?Whether the handler's process is actually PID 1. If the entrypoint is a shell script that starts the server as a child without `exec`, the shell receives the signal and the server never does. Switch to exec form (`ENTRYPOINT ["/app/server"]`) or `exec` the binary at the end of the script. Also confirm the image's `STOPSIGNAL` still matches the signal the handler installs.
- How would you choose the stop grace period for a service?Measure how long a full drain takes under representative load — longest in-flight request plus flush and deregistration time — and set the timeout above that with headroom, via `docker stop -t` or Compose's `stop_grace_period`. Pair it with an internal deadline in the shutdown code so the process always exits itself before the runtime escalates to SIGKILL. If drain time is genuinely long, that is a design signal to shorten request lifetimes rather than to keep extending the window.
saying these in an interview costs you the question
- Reading 143 as proof that graceful shutdown succeeded
- Assuming SIGTERM reaches the application even when a non-exec shell wrapper is PID 1
- Extending the grace period indefinitely instead of bounding the shutdown work
- Believing exit 0 on stop is impossible because the process was signalled — a handled signal lets the process choose its own status