skip to content

Your platform team wants one shutdown contract for every service on the estate - what do you mandate, and what do you leave to each team?

level: principalimportance: should knowfreq 34%

answer

  1. standardise shape, not values
  2. the order is universal
  3. the window is a derived number
  4. mandate the formula and the evidence
  5. cut-work counter or it is unfalsifiable

basics

~20 s

Mandate the ordering and the evidence: leave the routing set first, drain, refuse, finish, release, exit - plus a measured drain pause and a counter for work cut by the forced kill. Leave the window's length and what finishing means to each service.

solid answer

~50 s

Standardise the parts that are the same everywhere and refuse to standardise the part that is not. The ordering is universal, so mandate it: leave the routing set, pause while that propagates, stop accepting, finish, release claims and leases, exit. The measurement is universal, so mandate that too: each service must publish how many units of work its replicas had cut by the forced kill, or nobody will ever know the contract is being violated. What must not be one number is the window itself, because the same value that is generous for a request-serving service starves a batch worker and inflates rollouts for a service with nothing in flight. Require each service to derive its window from its own measured tail instead. Then own the consequence: the estate pays the drain pause on every replacement, and buying it back means restructuring work, not shortening the pause.

go deeper

for a junior

Understand that the window and the pause are numbers someone chose deliberately, and that copying them from another service can be wrong in either direction for yours.

for a middle

Be able to explain why the ordering can be standardised while the window cannot: one is the same for every workload, the other is derived from how long that workload's own work runs.

for a senior

Show that you would make the contract checkable - a count of work cut by the forced kill - rather than auditing configuration, and that you know cutting the drain pause converts a time cost into a correctness cost.

for a principal

Own the estate-wide economics: the pause is paid on every replacement everywhere, the legitimate way to buy it back is restructuring long work, and the standard needs a written escape hatch or it drifts upward exception by exception.

## Why a single number is the wrong mandate The tempting version of a shutdown standard is "every service gets a sixty-second window". It fails in both directions on the same estate: - The request-serving service finishes everything in under a second and pays a drain pause it never needed sized for someone else. - The reconciliation worker has a four-minute tail and is cut on every deployment, quietly, forever. - The service with an unmeasured drain pause keeps producing transport failures during deploys, and the standard did not address the part that was actually broken. A window is derived from a property of the workload, so it is the one number that cannot be a constant. What *can* be a constant is the shape of the sequence, the requirement to measure, and the evidence that the contract is being met. ## What the contract should mandate 1. **The ordering, exactly.** Leave the routing set; pause; stop accepting new work; finish what was accepted; release claims and leases; exit. This is identical for every workload on any platform, so there is no reason for forty teams to each invent it. 2. **A measured drain pause, not a guessed one.** Require the number to come from an observation - how long requests keep arriving after removal - and require it to be re-taken when the routing layer changes. 3. **A window derived from the service's own tail.** Mandate the formula, not the value: pause plus a high percentile of the longest legitimate unit of work plus a margin. 4. **Its own timeout on long work.** Work that hangs must be bounded by something that is not the forced kill, or the window sizing is corrupted by outages. 5. **Evidence.** A count of units of work cut by the expiry, per service. Without it, a shutdown contract is unfalsifiable: every service claims compliance and the ledger gaps continue. 6. **Shutdown behaviour in the application, not only in platform configuration.** A sequence that depends on the platform running a step on the workload's behalf is not portable across the estate's environments. ## What to leave to each team | Mandated centrally | Left to the service | |---|---| | The order of the steps | What "finish in flight" means for its work | | That the pause is measured | The measured value itself | | The formula for the window | The window's length | | That cut work is counted | How the work is made resumable | | That long work has its own timeout | The timeout's value | The division is deliberate: the platform team owns what is universal and observable, and the service team owns what depends on the shape of its work, because that is the knowledge the platform team does not have and cannot acquire at scale. ## The trade-off you are actually making Every second of drain pause is paid on **every replacement of every replica across the entire estate** - each rollout step, each scale-down, each node replacement. On a large fleet that deploys often, that is a real and continuous cost, and it is the number people try to cut first. Cutting it does not remove the cost; it moves it onto callers as connection failures during deploys, which is strictly worse because it is paid in correctness rather than in seconds. The legitimate ways to buy the time back are all in the work, not in the window: - **Make long units resumable** with checkpoints, so being cut costs a partial replay rather than a lost write. - **Split long units** into pieces that fit comfortably inside a defensible window. - **Move long work off the replacement path** entirely, onto workers whose handover is explicit and whose claim timeouts bound the damage. ## The escape hatch, stated in the standard The contract should say out loud what happens when a service cannot fit: **work whose longest legitimate unit exceeds any window the platform will defend does not belong on the shutdown path.** That sentence is the point of writing a standard at all. It converts a recurring argument about one team's window request into a design conversation about that team's work, and it stops the estate's window drifting upward one exception at a time until every rollout takes an hour. ## How you know it worked Not by an audit of configuration files. By the counter: cut-work events per service per week, trending to zero, with the exceptions being services that have an open design item about resumability. A standard that cannot be checked from telemetry is a document, and a document does not survive the first incident that makes a deadline inconvenient.

  • What single measurement tells you the contract is being violated?
    A per-service count of units of work that were still running when the forced kill landed. Everything else in a shutdown standard is configuration you can claim compliance with; this is the outcome. Trending it per service per week makes the exceptions visible and turns compliance into something you can check without reading anyone's configuration.
  • Why not simply mandate a generous window for everyone and move on?
    Because a window is a ceiling for work but a floor for wedged replicas: a replica that is stuck holds its capacity for the full budget before the kill. A very generous estate-wide value lengthens the worst case of every rollout and scale-down while still being wrong for the workload whose tail exceeds it.
  • When should the answer be that the work does not belong on the shutdown path?
    When its longest legitimate unit exceeds any window the platform will defend for a service that deploys frequently. At that point the lever is the work: checkpoint it so an interrupted run resumes, split it into smaller units, or move it to a worker whose handover is explicit rather than bounded by a replacement budget.

saying these in an interview costs you the question

  • Mandates one grace-period value for every service on the estate
  • Writes a standard with no way to detect violations
  • Treats the drain pause as pure overhead to be minimised away
  • Assumes each team will derive a window without being told the formula
  • Grants longer windows case by case until rollouts take hours