A running web service is asked to stop while requests are in flight — how should its shutdown proceed?
answer
- shutdown is ordered, not instant
- stop advertising before stop accepting
- the lame-duck window must outlast the refresh
- drain in-flight under a deadline
- release in reverse startup order
basics
~20 sFlip readiness to false first and keep serving for a moment so routers stop sending work; then stop accepting new connections; then let in-flight requests finish under a deadline; then release resources in reverse startup order and exit.
solid answer
~50 sGraceful shutdown is a sequence, and the order is the whole point. **First**, flip readiness to false while still serving — whatever routes traffic needs time to notice, so this lame-duck window must last longer than its refresh interval. **Second**, close the listener so no new connections are accepted, and mark idle persistent connections to close after their current response. **Third**, drain: let in-flight handlers finish, under a deadline shorter than the platform's kill timeout, rejecting anything new that slips in. **Fourth**, release: run shutdown hooks in reverse of startup so nothing is torn down while a dependent is still using it, flush buffers, close pools. **Then exit.** Skipping the first step is the usual bug: closing the listener immediately makes every request in the router's queue fail, which is exactly the dropped-request symptom that makes deploys visible to users.
go deeper
Know that stopping a service is not the same as killing it: requests already being handled should be allowed to finish, and new ones should go to another instance instead of failing.
Give the sequence in order and explain each step: advertise-stop, stop accepting, drain under a deadline, release in reverse startup order, exit. Know why Connection: close matters for reused connections.
Show that you have debugged this: errors clustered at deploy time, the lame-duck window versus the traffic layer's refresh interval, and a drain deadline chosen to fit inside the platform's kill timeout.
Own the time budget as a policy across services, including what long-running and streaming work is allowed to do. Decide what a shutdown that overruns its deadline must record so an incident is explainable afterwards.
Shutdown is the part of the lifecycle that gets least attention and causes the most user-visible damage. A deploy replaces every instance of a service; if each replacement drops the requests it was holding, a routine release becomes an error spike. **Graceful shutdown** is the ordered sequence that avoids that. ## The sequence 1. **Stop advertising: flip readiness false, keep serving.** The instance is still perfectly capable; it just wants to stop receiving new work. Whatever decides where traffic goes re-evaluates on an interval, so there is a lag between your flag changing and new requests actually stopping. This lame-duck window has to be at least as long as that lag — this is why the very first shutdown step is *not* to stop serving. 2. **Stop accepting.** Close the listening socket so new connections are refused (and land on another instance). Persistent connections that are idle should be closed; ones mid-request should be marked to close when their current response completes, which on HTTP means the response carries `Connection: close` so the client opens its next connection elsewhere. 3. **Drain in-flight work.** Let handlers that already started run to completion, under a deadline. Requests that somehow arrive after the decision to stop should be refused quickly with a retryable status — `503` with a `Retry-After` header is the conventional answer — rather than accepted and then abandoned. 4. **Release resources.** Run shutdown hooks in **reverse order of startup**, so a component is never closed while something that depends on it is still running. Flush buffered writes, close pooled connections, deregister from anything the service registered with at boot. 5. **Exit.** Exit deliberately once the drain completes, rather than waiting for the deadline to expire on an idle process. ## The budget | Interval | Must be | Why | |---|---|---| | Lame-duck window | longer than the traffic layer's refresh interval | otherwise requests are still being sent when you stop accepting | | Drain deadline | shorter than the platform's forced-kill timeout | otherwise the process is killed mid-request anyway | | Total shutdown | short enough not to stall a rolling replacement | a slow stop makes every deploy and scale-in slower | When the drain deadline expires with work still running, you have already lost the clean path; the honest choice is to stop waiting, record what was abandoned, and exit — an unbounded wait just moves the decision to whatever kills the process, with no logging and no flush. ## What makes requests get dropped anyway - **Closing the listener first.** Requests already in the router's queue, and connections already established, fail with a reset the caller sees as an error rather than a retryable refusal. - **A lame-duck window shorter than the refresh interval.** The flag flipped, nobody noticed yet, and traffic was still arriving when the socket closed. - **Persistent connections nobody told to close.** A client happily reuses a connection to an instance that is going away; the response for the next request on it never comes. - **Tearing down in startup order.** A pool is closed while a handler still holds a borrowed connection, turning a would-be-successful request into a failure at the last moment. - **Very long requests.** Streaming responses and slow uploads outlive any sensible deadline; they need either a cap or an application-level mechanism to end them early. ## Who decides how long you get The stop begins outside the process: a platform replacing the instance, an operator scaling in, a deploy rolling forward. The framework turns that into the sequence above, which means two things are true at once — the process chooses how gracefully it stops, and something else chooses how long it is allowed to take. A sequence that ignores the second half is not graceful, only optimistic: it gets cut off partway and calls the result a drain. Treat the externally imposed deadline as the budget the whole sequence must fit inside, and derive the lame-duck window and the drain deadline from it rather than picking round numbers. If the two cannot both fit, the honest answer is that requests here are too long for the platform's stop budget, and one of the two has to change. ## Where frameworks differ Some frameworks implement the whole sequence for you behind a single switch and only ask for a timeout; others give you the stop-accepting and drain primitives and leave the lame-duck window and the readiness flip to your own code. The drain guarantee differs too: some track in-flight requests exactly and complete when the count reaches zero, others simply wait a fixed period and hope. Check which you have, because the two behave identically on an idle service and completely differently under load. ## What an interviewer is listening for The ordering — advertise-stop before stop-accepting, drain before release, release in reverse — and the reason each step exists. A candidate who starts by closing the socket has described exactly the bug this question is about.
- Why is flipping readiness to false not enough on its own to stop receiving requests?Whatever routes traffic re-reads readiness on an interval and may hold pooled connections to your instance. Between your flag changing and its view updating, requests keep arriving, so the instance must keep serving normally for at least that lag before it closes the listener.
- What should happen to a request that arrives after the service has decided to stop?Refuse it quickly and retryably rather than accepting work that will be abandoned: a `503` with a `Retry-After` header tells the caller to try elsewhere. Accepting it and then hitting the drain deadline gives the caller a timeout, which is far more expensive than an immediate refusal.
- Why are shutdown hooks run in reverse order of startup?Startup builds dependencies before dependants, so reversing the order guarantees nothing is torn down while something still using it is alive. Closing in startup order closes a shared pool or client first, turning in-flight requests that were about to succeed into failures.
- How do streaming or long-running responses interact with the drain deadline?They can outlive any deadline you can afford, so they need their own treatment: cap the duration, send a hint that the peer should reconnect, or end the stream at a safe boundary when shutdown starts. Otherwise every stop either waits far too long or cuts a response mid-flight.
Closing a shop: first turn the sign to Closed so nobody new walks in, then lock the door, then let the customers already inside finish and leave, and only then cash out and switch off the lights. Locking the door first traps people mid-purchase.
saying these in an interview costs you the question
- Closes the listening socket as the very first shutdown step
- Thinks flipping readiness false stops traffic instantly
- Waits for in-flight work without any deadline at all
- Closes pools and clients before in-flight requests finish
- Ignores persistent connections that clients will reuse
- Treats a forced kill as an acceptable normal stop