skip to content

How do you set a service's drain budget against the platform's stop timeout, and is a timed-out drain worth paging on?

level: principalimportance: should knowfreq 38%

answer

  1. two decisions hiding under one number
  2. the outer bound is not yours to set
  3. cost of an abandoned item decides the alert
  4. longer budgets are paid for at rollback time
  5. smaller unit of work beats a bigger budget

basics

~20 s

Fit the budget strictly inside the platform's grace period, then set the posture from what an abandoned item costs: redeliverable work means alert on the rate, a duplicate or dropped side effect means page. Write the rule down.

solid answer

~50 s

Two decisions, and conflating them is the usual mistake. The first is arithmetic: the platform kills the process after a fixed grace period, so the budget is that period minus signal delivery, the final acknowledgements and flush, and the exit. A budget at or above the grace period is theatre - the kill lands where your timeout branch would have run. The second is a posture the service owner sets and the platform team can overrule. If the queue redelivers to idempotent handlers, an incomplete drain is an accepted loss: count it, alert on the rate, do not page. If an abandoned item means a duplicate charge or a silent drop, it is a correctness event that pages. Resist simply raising the number: long budgets slow every rollback, and when the work genuinely does not fit, make the unit of work smaller rather than asking the fleet for a longer grace period.

go deeper

for a junior

Take away that the drain budget is a chosen number, not a default, and that it has to be smaller than the time the platform allows before it kills the process.

for a middle

Be able to enumerate everything that must fit inside the grace period besides the drain itself — signal delivery, deregistration, the final acknowledgements and flush, the exit — and explain why an equal budget means the timeout branch never runs.

for a senior

Argue the operational side: what the metric should alert on, why raising the budget usually masks one unbounded call, and what it costs the fleet when every instance drains for a long time during a rollback.

for a principal

Own the posture and the negotiation. Decide what an abandoned item costs for this service, whether that alerts or pages, when the answer is to reshape the work rather than ask the platform for a longer grace period, and write the rule down so it survives the next incident.

## Two decisions wearing one name "How long should we drain for" hides an arithmetic constraint and a policy choice, and they should be argued separately. ### The arithmetic The platform sends a termination signal and kills the process after a fixed grace period. Everything the program does must fit inside that, not just the drain: - signal delivery and the shutdown path starting; - stopping intake and any deregistration from routing; - the drain itself; - acknowledging the last completed items, flushing buffered telemetry, closing files; - returning from main. So the budget is the grace period minus all of that, with margin. A budget set equal to the grace period is the failure mode that looks fine in review and never works: the kill lands at approximately the instant the timeout branch would have run, so you get neither the finished work nor the report about the unfinished work. If the two numbers live in different repositories — the budget in the service, the grace period in the deployment configuration — say so explicitly in both places, because a platform-side reduction silently breaks every service that was tuned to the old value. ### The policy What should happen when the budget runs out is not a Go question, it is a question about what an item is worth: - **Redeliverable and idempotent work.** The queue redelivers anything not acknowledged, and reprocessing is harmless. An incomplete drain costs some duplicated effort. This is an accepted loss: emit the metric, alert when the rate crosses a threshold, do not wake anyone at three in the morning for one instance. - **At-most-once side effects.** The item sends money, emails a customer, or writes something that is not naturally idempotent. Now an interrupted item is either a duplicate or a silent drop, and that is a correctness event worth paging on. The better engineering answer is usually to remove the ambiguity — make the effect idempotent with a key — but until that exists, the alert has to reflect the real cost. - **Work with an external commitment.** An item holding a lease or a lock elsewhere may leave that resource unavailable until it expires. The drain policy has to account for the recovery time, not only for the item. ## Why longer is not safer The instinct after one incomplete drain is to raise the number. Push back on it with the costs that are easy to forget: - **Rollback time.** Deploys stop instance by instance. A generous budget multiplies through the whole rollout, and the number that really matters during an incident is how fast you can get back to the previous version. - **Capacity during deploys.** Instances that are draining still hold memory, connections and downstream capacity while contributing nothing. - **It hides the real bug.** Nearly every persistent drain timeout is one unbounded operation. A longer budget converts a loud, reproducible failure into a slow deploy that nobody investigates. The honest sequence is: dump what was still running, bound that operation, then reconsider the budget — usually downward. ## When the work genuinely does not fit Some items legitimately take longer than any sane grace period. The answer is structural, not numerical: - **Make the unit of work smaller.** Split a long item into steps that each fit comfortably, so an interrupted item loses one step. - **Checkpoint.** Record progress so redelivery resumes rather than restarts. - **Stop taking long items before the drain.** If the arrival of a long item is predictable, stop accepting them earlier than the ordinary ones, so the drain begins with only short work in flight. - **Split the service.** If one class of work needs a fundamentally different shutdown profile, it probably wants its own deployment, with its own grace period, rather than dragging the whole service's deploy latency up. ## Who can overrule whom The grace period is usually a fleet-wide platform setting; the drain budget is the service's own. That gives the platform team the last word on the outer bound and the service owner the last word inside it. The productive framing when they conflict is not "give us more time" but "here is the distribution of item durations, here is what an abandoned item costs us, and here is what it would take to fit". A special-cased longer grace period for one service is a real cost paid by the people who operate the fleet, and it should be justified with numbers rather than asserted. ## Write it down The artefact this decision should produce is short and explicit: the budget, the grace period it fits inside, what an abandoned item costs, and whether the metric alerts or pages. Without it, the number is re-derived by whoever last had an incident, and the alerting posture drifts to whatever felt right during the last outage.

  • A team asks for a longer platform grace period because their drain times out. What do you want to see first?
    The distribution of item durations and the goroutine dump from a timed-out drain. Almost always it shows one unbounded operation rather than ordinary work that needs more time, and bounding that is cheaper than a fleet-wide change that every other service pays for in deploy latency.
  • How would you decide whether incomplete drains alert or page?
    By what an abandoned item costs. Redeliverable, idempotent work means an occasional incomplete drain is an accepted loss: alert on the rate. Work with an at-most-once side effect means an abandoned item is a duplicate or a silent drop, which is a correctness event and belongs on a page until the effect is made idempotent.
  • What do you do when a single item legitimately takes longer than the grace period?
    Change the shape of the work rather than the number: split it into steps, checkpoint progress so redelivery resumes instead of restarting, or stop accepting long items earlier than short ones. If one class of work truly needs a different shutdown profile, give it its own deployment.
  • The drain budget and the grace period live in different repositories. What do you do about that?
    Make the dependency explicit in both places and, where possible, derive the budget from the platform value at start-up rather than hard-coding it. Otherwise a platform-side reduction silently breaks every service that was tuned to the old number, and it will be found during an incident.

saying these in an interview costs you the question

  • Sets the drain budget equal to the platform's grace period
  • Raises the budget after every incident without diagnosing
  • Pages on any incomplete drain regardless of what work is lost
  • Treats deploy and rollback latency as free
  • Assumes the grace period will never be changed by the platform team
  • Has no written statement of what an abandoned item costs