skip to content

You own a containerized service that both serves HTTP requests taking up to 20 seconds and consumes messages from a queue. Design its shutdown path: what should the process do on SIGTERM, and how do you choose the stop timeout it runs under?

level: principalimportance: should knowfreq 34%

answer

  1. Fail readiness first, then stop accepting
  2. Internal deadline < stop timeout
  3. Nack in-flight messages; require idempotency
  4. Size to p99 + propagation + flush, not max
  5. 137 in dashboards = design not honored

basics

~20 s

On SIGTERM: fail readiness first, stop accepting new HTTP connections and stop pulling from the queue, finish or abandon in-flight work under a bounded deadline shorter than the stop timeout, flush and close, exit 0. Set the stop timeout above your worst realistic drain, not above your worst theoretical one.

solid answer

~50 s

Treat shutdown as a contract with a deadline. On SIGTERM the process should, in order: 1. **Signal unreadiness** so the load balancer stops sending new requests, then wait a short propagation delay — otherwise you close the listener before traffic is diverted and users see connection resets. 2. **Stop accepting new work**: close the HTTP listener to new connections, pause the queue consumer's prefetch. 3. **Drain in-flight work under its own bounded deadline** — say 15s when the stop timeout is 30s. Long-poll or streaming connections need explicit cancellation, not hope. 4. **Finish or safely abandon**: ack/nack outstanding messages so the broker redelivers rather than silently losing them. 5. **Flush and close**: log buffers, metrics, DB pools, then `exit(0)`. Sizing: stop timeout > readiness propagation + p99 request duration + flush margin. Make the app's internal deadline strictly smaller so *the app* decides how to end, not SIGKILL. Then make idempotency and redelivery-safety non-negotiable, because SIGKILL will still happen — node loss, OOM, crash.

code

bash · 3 lines
bash
docker run -d --stop-timeout 30 --name api myimage
docker stop -t 30 api
docker inspect -f '{{.State.ExitCode}} {{.State.FinishedAt}}' api

go deeper

for a junior

Focus on the basics: install a SIGTERM handler, stop accepting new work, finish what is running, then exit cleanly rather than being killed.

for a middle

Add ordering (readiness first, then listener, then drain), a bounded internal deadline, and matching the stop timeout to the app's drain time.

for a senior

Bring in the load-balancer propagation window, broker redelivery and idempotency, observability of drain duration, and diagnosing exit 137 in production.

for a principal

Argue the number from traffic data and fleet economics, define the shutdown contract every service must implement, and insist that correctness survive SIGKILL regardless of the grace period.

## Why this is a design question, not a coding question Shutdown is where availability, data safety and deploy velocity collide. Too short a grace period and every deploy sheds requests and duplicates queue work. Too long and rolling deploys crawl, a stuck process pins capacity, and incident response slows. There is no universally correct number — only a number derived from your traffic shape and your idempotency guarantees. ## The mechanism you are designing against Docker's `docker stop` sends the container's stop signal to PID 1, waits the stop timeout (default 10s, set with `-t`/`--stop-timeout`, or `stop_grace_period` in Compose), then sends SIGKILL. Orchestrators wrap the same primitive with their own grace-period setting. Everything below assumes the signal actually reaches your process — exec-form entrypoint or `exec "$@"`, and an installed SIGTERM handler, since PID 1 gets no default action. ## The shutdown sequence, in order **1. Fail readiness before you fail requests.** Load balancers and service meshes learn about your departure asynchronously. If you close the listener the instant SIGTERM lands, in-flight *routing decisions* already made send traffic to a closed socket → connection refused/reset, which retries may or may not paper over. The standard fix is a **pre-stop delay**: flip readiness to failing, keep serving normally for one or two health-check intervals, *then* stop accepting. This deliberate "lame duck" period is the single highest-value part of the design. **2. Stop taking new work.** HTTP: stop accepting new connections while continuing to serve accepted ones (Go `Server.Shutdown`, Node `server.close()`, Spring Boot graceful shutdown). Queue: stop the prefetch/poll loop first — pulling a message you cannot finish is the worst thing you can do at this moment. **3. Drain with an internal deadline.** Compute in-flight work and wait, but never unboundedly. If your stop timeout is 30s, your handler should give up at ~20-25s and proceed to step 5. The rule is that *your code* decides how the process ends, because your code can flush and log; SIGKILL cannot. Anything still running at the internal deadline gets cancelled explicitly (context cancellation, interrupt, connection close with a 503 and `Connection: close`). **4. Deal with in-flight messages honestly.** For an at-least-once broker, nack/abandon unacked messages immediately so they are redelivered fast rather than waiting out a visibility timeout. This is only safe if consumers are idempotent — dedupe key, or an upsert keyed on a message id. Long jobs should checkpoint so redelivery resumes rather than restarts. **5. Flush and exit deliberately.** Async log appenders, metrics buffers, trace exporters and connection pools all hold data. Close them in dependency order and `exit(0)`. Exiting 0 matters: it distinguishes an intentional stop from a crash in your dashboards and in restart-policy behavior. ## Sizing the timeout A workable formula: ``` stop_timeout >= readiness_propagation + p99_request_duration + flush_margin app_internal_deadline = stop_timeout - safety_margin (e.g. 5s) ``` For the stated service — 20s worst-case requests, ~5s for readiness to propagate, ~2s of flushing — a 30s stop timeout with a ~25s internal deadline is defensible. Design against **p99, not max**: sizing for the longest theoretical request means every deploy waits for the worst tail. Requests that legitimately exceed the window should be redesigned (async job + polling) rather than have the whole fleet's shutdown budget stretched around them. Also weigh the fleet cost. A 300-pod rolling deploy with a 120s grace period is a very different operation from one with 30s. And a grace period only *bounds* the wait — well-behaved services exit in a fraction of it, so the number is insurance, not latency you pay every time. ## What you cannot solve with grace periods SIGKILL is unavoidable: kernel OOM kills, node loss, hypervisor failure, a hung process that ignores its own deadline. So the design must be **crash-equivalent-safe**: idempotent consumers, transactional or checkpointed writes, no state that exists only in process memory, and no correctness that depends on cleanup code running. Graceful shutdown is an optimization that removes noise from the common case — never the mechanism that makes the system correct. ## Making it verifiable Make it observable and testable, or it will rot: log a structured line at SIGTERM receipt and at exit with the drain duration and in-flight counts; emit a `shutdown_duration_seconds` metric; alert when the p95 approaches the timeout, since that is the early warning before you start eating SIGKILLs. In CI or a staging soak, run load, stop the container, and assert zero 5xx and zero lost messages. Finally, keep the timeout consistent between local Docker, Compose and the orchestrator so behavior does not change per environment — and note that exit code 137 in production is the metric that says the design is not being honored.

  • Why add a delay between failing readiness and closing the listener, instead of draining immediately?
    Because load balancers, service meshes and DNS-based clients learn about the instance's departure asynchronously, over one or more health-check intervals. Closing the listener immediately means routing decisions already in flight hit a closed socket and clients see resets. A short lame-duck period — still serving, but advertised as unready — lets those decisions drain naturally and is usually the difference between a clean deploy and a burst of 5xx.
  • How do you handle a request that legitimately runs longer than any reasonable grace period?
    Stop treating it as a request. Convert it to an asynchronous job with a durable record and a polling or callback interface, so a shutdown mid-job is a resumable event rather than a lost response. Stretching the fleet's grace period to accommodate one long tail penalizes every deploy and still fails when the node dies.
  • If SIGKILL can happen anyway from OOM or node loss, why bother with graceful shutdown at all?
    Because it removes the common case from the failure budget. Deploys, scale-downs and restarts happen constantly; crashes are rare. Graceful shutdown makes routine operations invisible to users and keeps retry storms and duplicate processing out of your error budget. The correctness properties — idempotency, checkpointing, durable state — must still hold for the SIGKILL case, and graceful shutdown never substitutes for them.

Closing a restaurant: take the sign down first so nobody new walks in, let seated diners finish, tell anyone mid-order they will be served elsewhere, cash out the register — all before the landlord cuts the power at a fixed hour.

saying these in an interview costs you the question

  • Closing the listener the moment SIGTERM arrives, with no readiness/propagation delay
  • Draining with no internal deadline, so SIGKILL decides the ending instead of the app
  • Sizing the grace period to the maximum theoretical request instead of the p99
  • Treating graceful shutdown as a substitute for idempotency and durable state
  • Setting different stop timeouts locally and in the orchestrator, so behavior changes per environment

context