skip to content

How would you size the stop-to-kill grace period for a ledger service whose longest legitimate write takes about 90 seconds?

level: seniorimportance: should knowfreq 46%

answer

  1. measure, do not round
  2. three terms, not one
  3. pause plus work plus margin
  4. a high percentile, never the mean
  5. too short costs correctness, too long costs time

basics

~20 s

Size it from measurement: the drain pause that lets routing removal propagate, plus a high percentile of the longest legitimate unit of work, plus a margin. Ninety-second writes and a fifteen-second pause put the window near two minutes, not thirty seconds.

solid answer

~50 s

The window has to cover the whole sequence, not just the work: the drain pause while the routing removal propagates, then the longest unit of work that is still legitimately running after new work stopped, then releasing claims and exiting. So the arithmetic is `drainDelay + tail work duration + margin` - with fifteen seconds of pause and a ninety-second write, a two-minute window is honest and a thirty-second one is a decision to cut ledger writes in half. Take the work figure from a high percentile of the measured distribution, not the mean, and exclude runs that are pathological rather than legitimate - those should be bounded by their own timeout instead of by the kill. The counter-pressure is that a long window lengthens every scale-down and every replacement step, and it is only real if whatever asked for the stop actually waits that long.

code

yaml · 8 lines
yaml
shutdown:
  drainDelay: 15s        # measured: requests stop arriving ~12s after removal
  gracePeriod: 120s      # 15s pause + 90s tail write + 15s margin
  onExpiry: forceKill

work:
  longestLegitimateUnit: 90s   # 99.9th percentile of ledger write duration
  ownTimeout: 150s             # pathological runs die here, not at the kill

go deeper

for a junior

Know that the window is a number someone chose, and that it has to be at least as long as the work a replica may still be finishing. A default that nobody checked is a common source of lost work.

for a middle

Compose the budget from its parts: the drain pause, the tail of the work, and a margin. Explain why the mean duration is the wrong input and the tail is the right one.

for a senior

Demonstrate measurement and judgment: separating legitimate long work from a hang, re-measuring after the workload changes, and verifying that the declared window is actually honoured before trusting it.

for a principal

Own the asymmetry across the estate. Too short costs correctness on every shutdown, too long costs rollout time on some, and the escape hatch is restructuring the work rather than growing the number indefinitely.

## What the window has to cover The grace period is not "how long the work takes". It is the budget for the whole shutdown sequence, and the work is only one term in it: ``` gracePeriod >= drainDelay # routing removal propagates + longest legitimate unit of work # still running after new work stopped + release and exit # claims, leases, flush + margin # you measured on a good day ``` With a measured fifteen-second propagation delay and a ninety-second tail write, that lands near two minutes. Writing thirty seconds because it looked like a reasonable round number is not a compromise - it is a decision that every shutdown may cut a ninety-second write at second fifteen of its life. ## Getting the work figure honestly The number you want is **the longest unit of work that is still legitimately in progress when the replica stops accepting new work** - which is not the same as the slowest request you have ever recorded. - **Use a high percentile of the real distribution**, not the mean. The mean is dominated by the fast common case and will size the window for the requests that never needed it. - **Separate legitimate from pathological.** A write that takes nine minutes because a downstream call is hanging is not a unit of work you should size a window around; it is a unit of work that needs its own timeout. Bound it there, then size the window for what remains. - **Count what the replica accepted, not what it started.** A batch step kicked off internally on a timer is in-flight work too, and it is the one people forget. - **Re-measure after the workload changes.** A window sized before a bulk-import feature shipped is a window sized for a different service. ## What each extreme costs | Window too short | Window too long | |---|---| | Long writes are cut mid-step on every shutdown | Every replacement step waits longer for stragglers | | Multi-step effects are left half-applied | Capacity is held by replicas that are done serving | | Failures cluster on deploys, so they look like release bugs | A wedged replica lingers for the full window before it is killed | | The cost lands on callers and on the ledger | The cost lands on rollout duration and on the scale-down response time | The asymmetry matters: a window that is too long costs *time*, and only on the shutdowns where work genuinely runs long. A window that is too short costs *correctness*, on every shutdown, for exactly the most expensive requests. When you are unsure, err long. ## The window is a ceiling; the pause is a bill A generous window does not slow down the common case. A replica with nothing in flight leaves the routing set, waits out the drain pause and exits in seconds - the remaining budget is never spent. The only guaranteed cost of shutdown is the drain pause, which is paid on every replacement regardless of the window's size. That means the right instinct is to keep the *pause* tight and measured, and to let the *window* be as generous as the tail of the work genuinely requires. ## Declaring a window does not always get you one A declared window is only real if whatever requested the stop actually waits that long before killing the process. Platforms differ in how much of that they let a workload choose and whether an outer layer imposes its own cap, so a very long window is worth verifying by observation - stop a replica with work in flight and time the kill - rather than trusting the number you wrote down. ## When no number works Sometimes the honest answer is that the longest legitimate unit of work exceeds any window you would defend. A five-minute reconciliation run is not going to fit in a replacement budget for a service that deploys several times a day. At that point the window is the wrong lever, and the fix belongs in the work: 1. **Checkpoint it**, so an interrupted run resumes from its last committed point rather than restarting or half-applying. 2. **Split it** into units small enough to finish inside a defensible window. 3. **Move it off the request path**, onto a worker whose shutdown story is its own and whose claim timeouts bound the damage. That is a design decision, not a configuration one, which is why the sizing question is asked at senior level: the interviewer wants to see you reach for measurement first and recognise the point where measurement stops being enough.

  • Which statistic of work duration should size the window?
    A high percentile of the legitimate distribution - the tail is the whole reason the window exists, so the mean sizes it for requests that never needed it. Exclude runs that are pathological rather than long; those should be bounded by the work's own timeout, so the window covers real work and not a hang.
  • Does a two-minute window make every deployment two minutes slower per replica?
    No. The window is a ceiling, and a replica with nothing in flight exits in seconds. The only cost paid on every shutdown is the drain pause. The long tail of the window is spent only on the shutdowns where long work genuinely was running.
  • Is declaring a long window enough to actually get one?
    Not necessarily - it is real only if whatever issued the stop waits that long before killing the process, and platforms differ in how much of that a workload may choose. Verify it by observation: stop a replica with work in flight and time when the kill actually lands.

saying these in an interview costs you the question

  • Sizes the window from the mean request duration
  • Picks a round number without measuring anything
  • Forgets to include the drain pause in the budget
  • Thinks a long window slows down every shutdown equally
  • Sizes the window around a hanging call instead of bounding it
  • Assumes a declared window is always honoured as written