How long should Server.Shutdown's drain window be for a Go service deployed many times a day?
answer
- two costs, one number, two owners
- the outer kill deadline is a ceiling
- read it off the latency histogram
- expiring changes nothing by itself
- count the abandoned drains
basics
~20 sDerive it from the request-duration distribution, cap it below whatever kill deadline the deploy tooling enforces, and always escalate to Server.Close when it expires. The number is a negotiated cost split between slow deploys and cut requests, not a technical constant.
solid answer
~50 sTreat the deadline you pass to `Shutdown` as a policy with two owners. From outside, the deploy tooling already enforces a hard kill deadline, so a longer Go-side window is fiction — the process dies mid-drain anyway. From inside, the honest input is the latency histogram of the endpoints that hold connections open: draining for p99 covers almost everyone, draining for the maximum lets one report download hold a deploy hostage. Then decide the escalation explicitly, because `Shutdown` returning its context error leaves those connections open and their handlers running: call `Close` after it, or you pay the whole window and cut the requests anyway. Where a class of endpoint cannot drain, make the transfer resumable rather than inflating the number. Publish a default in a shared skeleton, allow overrides with a stated reason, and revisit when deploys crawl or truncations rise.
go deeper
Know that the deadline is a real choice with consequences, and that copying a number from another service without checking how long its requests run is guesswork.
Be able to derive a candidate window from request latencies and to explain why the process's outer kill deadline caps whatever you choose.
Own the escalation and the instrumentation: the Close after the deadline, the abandoned-drain counter, and the evidence that tells you when to change the number.
Set the fleet default and the override rules, keep the inner and outer deadlines governed together, and be ready to argue for a design change instead of a bigger window.
## Why this is a decision and not a constant Every Go service that stops gracefully contains one number: the deadline on the context handed to `Server.Shutdown`. It looks like a detail in `main`, but it allocates cost between two groups who never meet. Too short, and users of the report-download service get their transfers cut on every deploy — several times a day. Too long, and a deploy that should take seconds takes minutes, rollbacks get slower exactly when they matter most, and the person on call starts to dread pushing a fix. The engineer who typed `30*time.Second` made that allocation, usually without knowing they had. ## The constraints that bound the number **The outer kill deadline is a hard ceiling.** Whatever supervises the process gives it a fixed amount of time to stop before terminating it outright. A 120-second drain under a 30-second stop timeout is not a 120-second drain; it is a 30-second drain followed by a kill, with the difference being that you no longer control what happens at the end. The Go-side deadline must sit under the outer one with headroom for everything else in the stop sequence — background workers joined after `Shutdown`, buffered writes flushed, dependencies closed. **The latency distribution is the honest input.** For each endpoint that holds a connection open, look at where requests actually finish. If the API endpoints finish under a second at p99 and the download endpoint has a p99 of two minutes, there is no single number that serves both, and the pretence that there is one is where most of these policies go wrong. **The deploy frequency multiplies everything.** A service deployed twice a month can afford a generous window nobody notices. A service deployed several times a day pays the drain cost several times a day, and any user-visible cut happens several times a day too. ## The escalation is part of the policy This is the Go-specific half that a generic "pick a grace period" answer misses. When the context expires, `Shutdown` returns `context.DeadlineExceeded` and stops waiting — and does nothing else. The connections it gave up on stay open, their handlers keep running, and the request contexts are not cancelled. So the code must say what happens next: ```go if err := srv.Shutdown(ctx); err != nil { drainAbandoned.Add(1) log.Printf("drain abandoned after %s: %v", window, err) _ = srv.Close() } ``` Without the `Close`, the process sits there holding open connections until the supervisor kills it — you have paid the full window *and* still cut the requests, which is the worst of both. With it, the behaviour is deliberate, immediate and logged. Deciding to always escalate is the easy call; the interesting one is deciding what you count and alert on when it happens. ## Making the number defensible A policy someone can be overruled on needs evidence, so instrument it: count abandoned drains, record the drain duration, and count truncated responses if the protocol lets you see them. Now the argument is empirical. When an SRE says deploys are too slow, you can show that 99% of drains complete in 900 ms and the tail is one endpoint. When a product owner says downloads are being cut, you can show how often, and at which window it stops happening. ## Structuring it across a fleet Set a default in whatever shared service skeleton the teams start from, so the common case needs no thought and no one retypes the shutdown wiring. Allow per-service overrides, but require a written reason next to the number, because that reason is what the next person needs when the constraint changes. Keep the outer kill deadline and the inner drain deadline in the same place if you can — the failure mode where somebody lowers the platform stop timeout and silently invalidates every service's drain window is entirely preventable and entirely common. And for the class of work that genuinely cannot drain — the multi-minute download — the right answer is usually not a bigger number. Make the transfer resumable so that a cut costs the client a retry rather than a restart, or move the work off the request path so the connection is not held at all. That is a design change with a cost, which is precisely the sort of trade a principal is expected to put on the table rather than absorbing it into an ever-growing timeout. ## Reviewing it The policy has two symptoms and both are observable. Deploys getting slower means the window is being consumed; a rise in abandoned drains or truncated responses means it is too small for what the service now does. Either one is the trigger to re-derive the number from the current histogram rather than to adjust it by feel.
- A team wants a ten-minute drain so their downloads are never cut. What do you say?That it only works if the supervisor's kill deadline is longer, which it usually is not, so the extra minutes are imaginary. Then push on the real question: a ten-minute drain makes every rollback ten minutes slower. Resumable transfers, or moving the work off the request path, buys the same outcome without that cost.
- What do you instrument so this number can be argued about with data?The drain duration and its outcome on every shutdown, a counter for abandoned drains, and the latency distribution of the endpoints that hold connections open. With those, "deploys are slow" and "downloads get cut" both become measurable claims instead of competing anecdotes.
- Should the escalation to Server.Close ever be omitted?In practice, no. Skipping it does not save the requests — they die when the process is killed anyway — it just removes your control over when and your chance to log it. The only variant worth considering is a short second window before Close, and even that mostly adds delay for no gain.
saying these in an interview costs you the question
- Picks a round number with no reference to request latency
- Sets a drain longer than the supervisor's kill deadline
- Leaves Shutdown's error unhandled with no Close escalation
- Applies one global window to every service and endpoint
- Grows the window instead of making long transfers resumable