skip to content

How would you set disruption-cap policy across a shared cluster where maintenance windows keep expiring with drains still blocked?

level: principalimportance: should knowfreq 38%

answer

  1. a contract, not a setting
  2. who pays, owner or operator
  3. a cap must be affordable
  4. blocked drain alerts the owner
  5. named escape hatch, always recorded

basics

~20 s

Make a cap a contract the workload owner has to fund: reject caps whose floor equals the declared copy count, prefer caps expressed against that count, default the workloads that declare none, and alert the owner - not the operator - when a drain blocks.

solid answer

~50 s

A cap is written by one team and paid for by another, so the policy's job is to say who yields. Four rules carry most of it. A cap must be affordable: a workload asking to keep N copies available must declare more than N, and a floor equal to the declared count should be rejected when it is written rather than discovered at 02:00. Prefer a cap expressed against the copy count, because an absolute one silently changes strictness as the workload scales. Give undeclared workloads a default loose enough that the fleet stays drainable. And route a blocked drain to the workload's owner with a stated deadline and a written outcome at that deadline, because the operator can neither add copies nor decide the workload may take a gap. Name the escape hatch in advance, and record every use.

go deeper

for a junior

The thing to take away is that a cap is written by the workload's owner but felt by whoever has to empty a host, so it is a shared decision rather than a private setting.

for a middle

Be able to spot the unsatisfiable cap - floor equal to the declared copy count - and explain why an absolute cap and a proportional one mean different things as a workload scales up and down.

for a senior

Argue the operational side: where the alarm goes when a drain blocks, what the waiting deadline is, and why forcing past a cap has to be recorded and treated as a signal that something needs changing.

for a principal

Own the trade explicitly. You are pricing patch latency against availability and unattended automation against per-workload attendance, and your real target is preventing a fleet that has quietly stopped being drainable.

## What the policy is arbitrating A disruption cap is a rule one team writes and another team pays for. The workload owner writes 'keep this many of my copies available', and the person who has to empty a host tonight to flash firmware or swap a failing disk is the one whose window expires. Neither side is wrong: unpatched hosts are a real risk with a clock on it, and an ingester dropping below its serving capacity at 02:00 is a real incident. A policy that does not say who yields yields to whoever waits longest - which, in the situation described, is the maintenance window. ## Four rules worth writing down 1. **A cap has to be affordable.** A workload asking to keep N copies available must declare more than N copies. A floor equal to the declared count is arithmetically unsatisfiable - no drain of any host it runs on can ever proceed - and that should be rejected when the cap is declared, not discovered during a window. 2. **Prefer a cap expressed against the copy count.** An absolute 'at most one down' means something very different at ten copies than at two, and a workload whose count moves through the day changes the strictness of its own cap with nobody editing it. Where the platform supports a proportional form, use it, and check what it does at your smallest copy count, because **platforms differ in how they round** and the small end is where drains block. 3. **Give undeclared workloads a default.** Most workloads never declare a cap, so the default decides whether the fleet is drainable at all. A loose default plus an exception process works; a strict default applied silently to everything does not, and you will find out months later. 4. **A blocked drain alerts the owner, with a deadline.** The operator cannot fix a cap, add copies, or decide that a workload may take a gap. Route the alarm to the team that can, state how long the drain will wait, and write down in advance what happens when that deadline passes. ## Absolute against proportional | | Absolute cap | Proportional cap | |---|---|---| | At ten copies | mild, roughly a tenth of headroom | tracks the count | | At two copies | strict, frequently blocking | tracks the count, subject to rounding | | When the count changes | strictness changes silently | meaning stays stable | | Easy to reason about | yes | needs a rounding check | ## The escape hatch, named in advance Some platforms let an operator force a removal past a cap and some do not. Where the capability exists, the policy question is not whether to use it but **who may**, and the answer should never be 'whoever holds the pager at the time'. Put the breach decision with the owner who set the cap or with an on-call lead, require that each use is recorded, and treat every use as evidence that a cap or a copy count needs changing. A cap that is routinely and silently breached is worse than no cap: it still blocks honest automation while protecting nothing. ## What you are actually trading - **Patch latency against availability.** Every hour a drain waits is an hour a host runs old firmware, and fleets that cannot be emptied accumulate that debt across every host at once rather than one at a time. - **Automation against attendance.** Loose caps let unattended maintenance run overnight; strict caps demand a human per workload per window, which does not scale past a few dozen services. - **Whose incident it is.** A loose cap turns a window into a small availability dip the owner did not schedule. A strict one turns it into an expired window the operator did not schedule. Pick deliberately, per workload, and say so out loud. ## The failure to design against The outcome to avoid is not a blocked drain; it is a **fleet that quietly stopped being drainable**. It arrives as a handful of workloads with unsatisfiable caps nobody reviews, an operator who has learned to force past them as routine, and a set of hosts that have not been emptied in months. Every rule above exists to make that state visible while it is still three workloads, rather than after it has become the way the cluster works.

  • Where do you draw the line and force a removal past a cap?
    At a deadline written in advance, not in the moment. A cap that has blocked maintenance past an agreed limit escalates to a decision with a named owner. Forcing is legitimate; forcing silently, by whoever happens to be holding the pager, is not, because the whole value of the cap is that breaching it was somebody's choice.
  • How do caps interact with a workload whose copy count changes through the day?
    An absolute cap drifts in meaning as the count moves: the same 'at most one down' is mild at ten copies and unsatisfiable at one. Where the platform offers a proportional cap, prefer it, and review absolute ones whenever the scaling bounds change - the night-time floor is the case that blocks.
  • How do you tell whether a cap is protecting anything at all?
    Ask when it last refused a removal that mattered, and what the workload does when a host fails instead. A cap on a workload with plenty of copies rarely binds and costs nothing; a cap on a workload that cannot survive losing one copy binds constantly and still does not protect it against failure. The second is the one to renegotiate.

saying these in an interview costs you the question

  • Just raise every cap until drains stop blocking
  • The operator should force the removal when a window is closing
  • Whoever runs the hosts should decide each workload's cap
  • A workload can promise three available while running three copies
  • Once set, a cap needs no review when the copy count changes