skip to content

Go cannot kill a goroutine, so how do you set your service's shutdown budget when the platform owns the grace period?

level: principalimportance: nice to knowfreq 27%

answer

  1. two owners, one number
  2. the deadline abandons, it never shortens
  3. which work is actually unrecoverable?
  4. make being killed a supported path
  5. arrive with measurements, not a request

basics

~20 s

Treat the number as a contract with the platform: start from the grace period you are granted, subtract margin, and budget only for work a restart cannot redo. Everything else is made retryable, so abandoning it is safe.

solid answer

~50 s

Start from the fact that Go gives you no way to kill a goroutine: whatever is still running when the deadline arrives is abandoned, not shortened. So the budget is not how long we would like, but what must complete before we are abandoned. Split teardown work into what is unrecoverable if lost — a non-idempotent external call already half applied — and what a restart simply redoes: unacknowledged messages are redelivered, buffered metrics re-derived. Budget only for the first group, keep the internal deadline comfortably under the grace period the platform grants, and make being killed at that deadline a supported path with tested recovery rather than a bug. If the work genuinely does not fit, arrive at the platform team with measured teardown times and offer to shrink the requirement — smaller units, checkpoints, later acknowledgement — instead of asking for more seconds.

go deeper

for a junior

Know that the process is given only a limited time to stop and is then killed, and that Go cannot force a goroutine to stop within it. Say that unfinished work must be safe to lose.

for a middle

Explain how the internal deadline sits under the externally granted period, and why in-flight work has to be classified into what a restart redoes and what it cannot.

for a senior

Show that you measure teardown time rather than assume it, put deadlines on the calls cancellation cannot reach, and design recovery so an abandoned unit resumes on the next start.

for a principal

Own the negotiation: the grace period is a shared contract with a real cost in rollout speed, so bring measurements, offer to shrink the requirement instead of asking for more time, and hold the line that being killed at the deadline is a supported path.

## Two owners, one number How long a service may take to exit is owned twice. The **platform** owns the grace period: it asks the process to stop, waits some number of seconds, then kills it outright. It cares because that number multiplies across every instance in every rollout, and it sets how fast a bad release can be replaced. The **service owner** knows what work is in flight and what it costs to lose. Neither can set the number alone, and the service owner can be overruled — which is exactly why the argument has to be made in terms the platform team can act on. ## The Go fact that shapes the whole decision Go has no way to kill a goroutine. Cancellation is cooperative: a goroutine that is inside a blocking call which does not watch a context cannot be shortened at all. So the deadline is not a mechanism for stopping work — it is a mechanism for **abandoning** it. Whatever has not finished when the timer fires is left where it stands, and the process exits. That reframes the design question. Instead of "how long do we need to finish everything", the question is "what must have finished before we are allowed to be abandoned, and how do we make everything else safe to abandon". ## Classify the teardown work Walk the teardown sequence and sort each step: * **Unrecoverable if lost.** A non-idempotent external effect that is half-applied. A record written without its counterpart. These are the only steps that genuinely deserve budget, and the strongest move is usually to remove them from this category — an idempotency key, a two-phase write, a transaction that either commits or does not — rather than to buy more seconds. * **Redone on restart.** An unacknowledged message is redelivered. A batch that was not checkpointed is recomputed. This is the bulk of most services' in-flight work, and it costs a little duplicated effort, not correctness. * **Best effort.** A final flush of buffered logs or metrics. Nice to have, never worth blocking an exit for; give it a small slice at the very end and let it be cut. The budget is the sum of the first group plus margin — not the sum of everything the service happens to be doing. ## Derive the number, do not pick it Work inwards from the grant. If the platform kills at N seconds, the internal deadline is meaningfully below N, because you want your own watchdog to fire first: your exit is attributable and it produces a diagnostic, while an external kill produces nothing. Below that sit the individual step budgets, and each one is measured, not guessed — instrument how long teardown actually takes, publish the distribution, and treat the tail as the design input. That instrumentation is also what converts the conversation with the platform team from opinion into evidence. "We need 90 seconds" is a position. "Our teardown's tail is 12 seconds, 11 of which is one non-idempotent call we are removing this quarter" is a plan. ## When the answer is no Assume the grace period will not be raised for you, and design for that as the normal case. The lever you still own is the size of the requirement: * Shrink the unit of work, so less is ever in flight at once. * Checkpoint mid-work, so an abandoned unit resumes rather than restarts. * Acknowledge later, so an abandoned item is redelivered instead of lost. * Move the unrecoverable step out of the request path entirely. Every one of these makes the service cheaper to kill, which is worth more than a longer grace period: a service that is safe to abandon at any instant is also safe during a crash, a node failure or an out-of-memory kill — none of which give you a grace period at all. ## Own the kill path The last part of the judgment is cultural. If being killed at the deadline is treated as a bug, every incident produces pressure to raise the grace period, and the number ratchets up across the fleet. If it is treated as a supported path — recovery on the next start is tested, duplicate delivery is handled, the exit is logged with the step that was outstanding — the number can stay small, rollouts stay fast, and the service tolerates the failures that were never going to be graceful. ## What a strong answer sounds like It names the constraint (the grant, and Go's inability to kill a goroutine), classifies the work rather than treating it as one lump, derives the internal deadline from measurement, offers a concrete alternative to a longer grace period, and accepts the overrule without leaving the service unsafe. It does not open with a number.

  • The platform team refuses to raise the grace period. What do you change in the service?
    Shrink the requirement rather than fight the grant: smaller units of work so less is in flight, checkpoints so an abandoned unit resumes, later acknowledgement so an abandoned item is redelivered, and idempotency on the one external call that was genuinely unsafe. The service ends up cheaper to kill, which also helps in crashes that offer no grace at all.
  • Which teardown steps are you willing to skip when the deadline is close?
    Anything a restart redoes: flushing buffered metrics, draining optional caches, finishing best-effort background work. What must not be skipped is a half-applied non-idempotent effect. That ranking decides the teardown order too — put the unrecoverable steps first, so the ones that get cut are the ones designed to be cut.
  • How do you keep this number from ratcheting upward across a fleet?
    Treat being killed at the deadline as a supported path with tested recovery, not as an incident. Every service that files a slow shutdown as a bug pushes the grace period up for everyone, and it never comes back down; measured teardown times published per service keep the conversation specific instead of fleet-wide.

It is the last train, not a taxi. You do not negotiate the departure time; you decide what must be on board, and make everything else able to catch the next one.

saying these in an interview costs you the question

  • Picks a round number with no reference to what must finish
  • Assumes the process always gets to complete its teardown
  • Treats being killed at the deadline as an unrecoverable bug
  • Thinks a goroutine can be forced to stop when time runs out
  • Asks for a longer grace period without measured teardown times
  • Budgets for flushing logs as if it were correctness-critical