skip to content

A ledger write replica is asked to stop, then forcibly killed 30 seconds later - what is that grace period for, and what happens when it expires?

level: middleimportance: must knowfreq 68%

answer

  1. a window, not an instruction
  2. asked to stop, then killed
  3. two events with a timer between
  4. the second event cannot be handled
  5. unfinished work is cut mid-step

basics

~20 s

The grace period is a budget for finishing work already in progress. The workload is asked to stop but keeps running; when the window ends it is killed outright. Anything still running at that instant is cut where it stood.

solid answer

~50 s

Stopping a container is two events with a timer between them. First the platform asks the workload to stop - nothing is taken away, the process keeps running and keeps its connections. Then, when the grace period expires, the platform kills it, and that second event cannot be caught, delayed or negotiated. The window in between is the workload's entire budget for leaving the routing set, refusing new work, finishing what it already accepted and releasing anything it holds. It is a ceiling, not a delay: a replica that finishes in two seconds should exit in two seconds. For a ledger write service the interesting case is the write that is still open when the timer runs out - the connection to the store drops mid-transaction, and everything that had not committed is simply gone.

go deeper

for a junior

Recall the shape: a workload is asked to stop, keeps running for a fixed window, and is killed if it is still there at the end. The two events and the timer between them are the whole idea.

for a middle

Explain what the window is spent on and in what order, and why the forced kill at the end cannot be handled. Be able to say what a cut write leaves behind in a durable store and what it does not.

for a senior

Show that you have seen the consequence: the ledger gaps that line up with scale-downs, the multi-step work that is not protected by any single transaction, and why the window is a ceiling rather than a delay.

for a principal

Frame the trade-off for a fleet. Every second of window is a policy applied to every replacement across the estate, so argue about where the budget should come from - a longer window, or work restructured so it does not need one.

## Stopping is two events, not one A container is never just "stopped". Two separate things happen, and the distance between them is the whole subject. 1. **The stop request.** The platform tells the workload to end. That is all it does. The process keeps running, keeps its memory, keeps its open connections, and keeps whatever work it had in progress. Nothing has been removed yet, and the platform is now waiting. 2. **The forced kill.** When the timer expires, the platform ends the process without asking. There is no handler for this, no callback, no final flush, no chance to write one more record. The process simply stops existing. The interval between the two is the **grace period**. Platforms let you declare it per workload and ship a default in the tens of seconds, because a default that is too generous makes every replacement slow and a default of zero makes every replacement lossy. The exact number is a policy choice you own, not a property of the workload. ## What the budget is for During the window the replica is supposed to spend the time, in this order: - **Leave the routing set**, so that whatever directs traffic stops choosing it as a target. - **Keep serving briefly**, because that removal does not take effect everywhere at once. - **Stop accepting new work** once the removal has had time to propagate. - **Finish the work it already accepted** - the requests in its handlers, the writes it has opened. - **Release what it holds** - claimed messages on a work queue, an exclusive lease, a reserved slot. - **Exit**, ideally well before the timer runs out. All of that has to fit inside the budget. If it does not fit, the tail of the list is what gets cut. ## What the expiry actually costs | State at the moment of the forced kill | What happens to it | |---|---| | A request already answered | Unaffected; the response left before the kill | | A request still inside a handler | Cut; the caller sees a broken connection, never a response | | An open, uncommitted write | The store rolls it back once the connection drops | | A committed write whose follow-on step never ran | Stays committed; the follow-on never happens | | A claimed message on a work queue | Stays claimed until its own claim timeout expires | | Buffered output not yet flushed | Lost with the process | The row that surprises people is the fourth. A single transaction against a durable store is safe: the store notices the dropped connection and rolls the transaction back, and the ledger is consistent. What is *not* safe is work that was never one transaction - a balance committed here and a notification that was supposed to be published there, a first step applied and a second step that the kill pre-empted. Graceful shutdown exists mostly to protect multi-step work, because single-step work was already protected by the store. ## The window is a ceiling, not a delay A common misreading is that declaring a 30-second window means every replacement takes 30 seconds. It does not. The workload exits as soon as it is genuinely finished, and the platform proceeds immediately. The only part of the window that is paid every single time is the deliberate pause you take while your removal from the routing set propagates - typically a handful of seconds. Everything after that is spent only if there is real work to spend it on. That asymmetry is why a generous window is cheaper than it looks and why a short one is more expensive than it looks. A window that is too short does not save you time on the common case, where work finishes quickly anyway; it only costs you the rare, long, important unit of work. ## Why interviewers ask this Because the failure is invisible in a healthy system and obvious in a post-incident review. Scale-downs, rollouts and node replacements all end with some replica being asked to stop, and the ledger's mysterious gaps - a debit with no matching credit, a request the client swears it sent - line up exactly with those events. A candidate who treats the stop request as "the container stops" has no model for the gap. A candidate who can name the two events, the budget between them, and the fact that the second one cannot be handled, can reason about every one of those incidents from first principles, on any platform.

  • Does a workload have to use the whole grace period?
    No. The window is a ceiling, not a required wait. Once the replica has left the routing set, finished its in-flight work and released what it holds, exiting immediately returns capacity sooner and shortens every rollout step. Only the deliberate pause that lets the routing removal propagate is paid on every single shutdown.
  • If a ledger write is cut mid-transaction, is the data corrupt?
    A single uncommitted transaction is not: the durable store rolls it back when the connection drops. The damage is in work that was never one transaction - a committed record whose follow-on step never ran, or an effect already applied in another system. Protecting that multi-step work is most of what the grace period is for.
  • What should the replica do with a request that arrives after it decided to stop?
    During the brief drain pause it should serve it normally, because that request was dispatched before the routing removal took effect and no other instance is expecting it. After the pause, the listener is closed and nothing new arrives at all - which is the point of closing it only once the removal has propagated.

It is last call at a bar: no new orders, finish the drinks already poured, and the lights go out at a fixed time whether or not anyone is done.

saying these in an interview costs you the question

  • Thinks the stop request itself terminates the process immediately
  • Believes the platform waits as long as the work happens to need
  • Assumes the final forced kill can be caught, deferred or negotiated
  • Thinks declaring a long window makes every shutdown take that long
  • Assumes every partial effect is undone because one transaction rolled back