How would you set the SIGTERM drain budget for a fleet of Python workers, and decide what to abandon?
answer
- Derive it, do not guess it
- A percentile, not a mean or a maximum
- The supervisor's timeout is a ceiling
- Make abandoning work cheap instead
- Track how often the kill wins
basics
~20 sDerive the budget from measured work-item durations — a high percentile plus margin, not the mean or the maximum — keep it inside the supervisor's window, and make abandoning work cheap by keeping every unit idempotent.
solid answer
~50 sTreat the budget as a measurement, not a preference. Instrument how long one unit of work takes, then size the drain from a high percentile — a 92nd-percentile item duration with headroom, so the common case finishes and the long tail is abandoned rather than dictating the window for everyone. The supervisor's stop timeout is a hard ceiling and the process budget must sit well inside it, leaving room for flush, close and interpreter finalization. Then attack the problem from the other side: the cheaper abandonment is, the shorter the window can be. Idempotent, checkpointed, at-least-once work makes a killed process a redo rather than a loss, and that buys back deploy speed across the whole fleet. Finally, measure shutdown itself — the fraction of processes still alive when the window expires is the number that tells you the budget is wrong.
code
python · 7 linesimport statistics
durations = [0.4, 0.6, 0.5, 1.2, 0.7, 2.9, 0.8, 0.9, 1.1, 3.4, 0.6, 1.0]
p92 = statistics.quantiles(durations, n=100, method="inclusive")[91]
grace_window = 30.0
budget = min(grace_window - 5.0, p92 * 3)
print(f"p92 item {p92:.2f}s -> drain budget {budget:.1f}s")go deeper
Understand that the time a process gets to shut down is a deliberate budget, not an accident, and that anything unfinished when it expires must be safe to redo.
Be able to derive a budget from measured work-item durations and explain why it must sit inside the supervisor's window with margin for flushing and finalization.
Classify the work: drain unconfirmed external effects, abandon recomputable work, give long-lived connections their own policy, and instrument shutdown duration and kill rate.
Own the position that a short window is a forcing function for retryable design, and that a generous one quietly subsidises services which cannot survive losing a machine without warning.
## The budget is a measurement, not a preference Teams usually pick a grace window by feel, and the number then calcifies. The defensible method is to measure the distribution of *one unit of work* and size the drain from it. Use a high percentile rather than the mean or the maximum. The mean under-serves a right-skewed distribution and abandons work that would nearly have finished; the maximum lets a single pathological item hold every deploy hostage. Something like the 92nd percentile of item duration, with a multiple for the items already in flight, finishes the overwhelming majority and declares the tail abandonable on purpose: ```python import statistics durations = [0.4, 0.6, 0.5, 1.2, 0.7, 2.9, 0.8, 0.9, 1.1, 3.4, 0.6, 1.0] p92 = statistics.quantiles(durations, n=100, method="inclusive")[91] grace_window = 30.0 budget = min(grace_window - 5.0, p92 * 3) print(f"p92 item {p92:.2f}s -> drain budget {budget:.1f}s") ``` The `min` is the important part: the supervisor's window is a hard ceiling, and the process budget must sit inside it with room for flushing, closing connections and interpreter finalization — plus clock skew, since the supervisor's timer started before your handler ran. ## Attack the other side of the equation The instinct when the drain does not finish is to lengthen the window. That is the expensive fix. A long window multiplies through every rolling restart — window times batches equals rollout time — and it silently converts an operational property into a correctness dependency: the longer you drain, the more the system's correctness rests on shutdown code running, which is exactly the code that does not run when a machine loses power. The cheap fix is to make abandonment harmless: * **At-least-once with idempotency.** A work item that is claimed with a lease, executed, and only acknowledged after its external effect is confirmed can be abandoned at any instant; the lease expires and someone else redoes it. Deduplicate on a stable key where a repeat would be visible. * **Checkpointing.** Long items that cannot be made short should record progress, so a redo resumes rather than restarts. * **Small units.** Halving the unit of work halves the p92 and therefore the budget. This is usually the single highest-leverage change available. Once abandonment is free, the budget stops being a risk control and becomes a pure efficiency knob — how much duplicated work you are willing to tolerate per restart. ## Not all work is the same A single fleet-wide number is a starting point, not an answer. Classify: * **Work with an unconfirmed external side effect** — a payment submitted but not yet acknowledged, a message handed to a transport. Drain this; abandoning it creates ambiguity that costs more than the wait. * **Recomputable work** — anything whose result can be rebuilt from inputs still sitting in a queue. Abandon immediately; draining it buys nothing. * **Long-lived connections** — streaming responses, long polls, websockets. These need a policy of their own: stop advertising readiness first so new connections go elsewhere, refuse new requests on existing ones, and close after the current response rather than trying to outlast the client. The two-stage shape matters generally: mark the process unready *before* the terminate signal initiates the drain, so whatever routes work has already stopped routing by the time you stop accepting it. A process that stops accepting while still being sent work manufactures errors during every deploy. ## Escalation and the second signal Decide explicitly what a second terminate signal means. The common contract is: first one starts the graceful drain, second one abandons it immediately. Make the handler idempotent so a repeated signal cannot corrupt the shutdown state, and consider exiting non-zero when the drain was truncated so the distinction is visible downstream. ## Measure the shutdown itself The numbers worth having on a dashboard: * **Shutdown duration**, as a distribution, not an average. * **The fraction of processes killed rather than exiting on their own** — a status reported as killed by signal 9, or `-9` from `subprocess.Popen.wait()`. If this is not near zero the budget is wrong, and if it is *exactly* zero across a wide margin the budget may be needlessly generous. * **Work re-executed after a restart**, which is the actual cost of a short window and the number that makes the tradeoff concrete rather than theoretical. Without these, the window is an argument between opinions. With them it is arithmetic. ## The organizational tradeoff A generous grace window is a subsidy for non-idempotent work: it lets teams ship handlers whose correctness depends on cleanup running, and that debt only surfaces during an incident, when machines vanish without any signal at all. The position worth defending is that the window should be short enough to be uncomfortable for anyone relying on it — short enough that services are built to survive being killed at an arbitrary instant, because eventually they will be. Set a fleet default from measurement, allow a documented exception for the genuinely unconfirmable side effect, and require the exception to come with the retry story that makes it unnecessary later.
- A team asks for a much longer grace window because their drain does not finish. How do you respond?Ask what breaks when it is abandoned. If the answer is duplicated work, the window is already long enough and the fix is deduplication. If the answer is lost or ambiguous state, the defect is the missing at-least-once contract, and a longer window only hides it until a machine disappears without any signal. Grant a documented exception for a genuinely unconfirmable external effect, with the retry work scheduled.
- What signal tells you the drain budget is wrong in production?The rate of processes reported as killed rather than exiting on their own — status 137, or `-9` from `subprocess.Popen.wait()`. A non-trivial rate means the budget is under the real drain time. The inverse matters too: if shutdowns consistently finish in a fraction of the window, the window is subsidising slow rollouts and can be tightened.
- Why mark a process unready before the drain rather than at the same moment?Because whatever routes work to it needs time to observe the change. If readiness drops at the same instant the process stops accepting, in-flight dispatches arrive at a process that is already refusing them, and every rolling restart manufactures errors. Dropping readiness first, then beginning the drain a beat later, makes the deploy invisible to callers.
saying these in an interview costs you the question
- Picking a grace window by feel and never revisiting it
- Sizing the budget from the slowest observed item
- Lengthening the window instead of fixing idempotency
- Assuming shutdown code always gets to run
- Using one number for every workload shape
- Never measuring how often the process is killed