skip to content

A workload runs one copy and its cap allows none down - why does every drain of its host block, and what resolves it?

level: middleimportance: should knowfreq 44%

answer

  1. arithmetic with no solution
  2. one copy, floor of one
  3. refused and retried forever
  4. more copies, looser cap, or outage
  5. caps forbid only the planned outage

basics

~20 s

The arithmetic can never be satisfied: taking the only copy down leaves zero available, below any floor of one, so the removal is refused and retried forever. Resolving it means more copies, a looser cap, or a deliberate outage.

solid answer

~40 s

A cap is evaluated as arithmetic on available copies, and with a single copy no arrangement satisfies it: removing that copy takes availability to zero, below a floor of one, so the request is refused and retried indefinitely. Waiting changes nothing, which is why this blocks the host's maintenance rather than the workload. There are three honest resolutions and each is somebody's decision: run more than one copy, if the workload tolerates a second instance; relax the cap so the copy may be taken down; or accept a short planned outage and remove it deliberately inside an announced window. The insight worth saying out loud is that the cap never made the workload available - it only refused to be the thing that took it down.

code

yaml · 8 lines
yaml
workload: fleet-registry
replicas: 1
voluntaryDisruptionCap:
  maxDown: 0

# floor    = replicas - maxDown = 1 - 0 = 1
# a drain must take the only copy: available after = 1 - 1 = 0
# 0 < 1, so the removal is refused - and refused again, forever

go deeper

for a junior

Do the sum. One copy, none allowed down, means the floor is one and removing the copy leaves zero. The request fails that comparison every single time it is retried, so the drain never finishes.

for a middle

Explain that the cap is a precondition, not a queue, and list the three resolutions with what each costs. Say clearly that a second copy is only possible if the workload tolerates two instances running at once.

for a senior

Show the consequence: the cap guarantees that this workload's outages will all be unplanned ones, and the host stays unpatched in the meantime. Then say who you would alert and why the operator cannot resolve it.

for a principal

Treat it as a declaration-time check rather than a 02:00 discovery: any workload whose floor equals its copy count is undrainable by construction, and that is a policy you can enforce when the cap is written.

## The arithmetic has no solution A disruption cap is evaluated as arithmetic on one number: how many copies of the workload are currently available. The cap implies a **floor** - the declared copy count minus however many may be voluntarily down - and a removal is granted only if availability stays at or above that floor once the copy has gone. With **1 copy** and a cap of **none down**, the floor is 1 - 0 = 1. Removing the only copy takes availability to 0, which is below 1, so the request is refused. The drain retries; the retry evaluates the same two numbers and reaches the same verdict. Nothing in the fleet can change it, because there is no other copy that could ever become available to make room. This is not a hung process or a lost lease - it is a condition that is false and will stay false. ## What the cap did and did not buy The cap did not make the workload available. It made one specific thing impossible: an automation emptying a host and taking that copy down without anybody deciding to. That is worth having, and it is much less than it looks like. The host can still fail on its own tonight, and when it does the copy goes with it and the cap is silent - caps hold back **voluntary** disruption only. So the practical effect of a strict cap on a single-copy workload is that all of its outages will be **unplanned** ones: the planned outage is the one the cap forbids, and the unplanned one is the one it cannot touch. Stated that way, it is usually the opposite of what the owner wanted. ## Three honest resolutions 1. **Run more than one copy.** The floor becomes satisfiable - two copies with at most one down leaves a floor of one, which a drain can meet. This is the real fix, and it costs whatever a second instance costs in resources plus the harder requirement: the workload must tolerate two of itself running at once, and not every workload does. 2. **Relax the cap.** Say the copy may be taken down. The drain then completes at the price of a gap in service while the replacement starts elsewhere. This is the right answer whenever that gap is genuinely cheap, which for plenty of internal workloads it is. 3. **Plan the outage.** Keep the strict cap as a guard against unattended automation, and remove the copy deliberately inside an announced window. The outage happens either way; this only decides when, and who knew about it beforehand. | Resolution | Service during maintenance | Cost | Who decides | |---|---|---|---| | More copies | continuous | a second instance the workload must tolerate | the workload owner | | Looser cap | a short gap | an unannounced gap on every drain | the workload owner | | Planned outage | a known gap | coordination and a window | the owner, with the operator | Every row is the owner's decision. The operator draining hosts can make none of them, which is why a blocked drain should alert the owner rather than the person holding the maintenance ticket. ## When the strict cap is right anyway A cap that no drain can satisfy is a legitimate setting for a workload whose single copy is expensive to lose and which nobody may take down unattended - so long as somebody owns the alarm it raises. The defect is not the setting; it is a setting nobody is watching, which silently converts every maintenance window into an expired one and leaves a queue of unpatched hosts behind it. ## The review that stops it recurring The condition is detectable before the window rather than during it: any workload whose floor equals its declared copy count can never survive a drain of a host it runs on. That is a one-line check over declared workloads, it can run continuously, and it turns a blocked window at 02:00 into a conversation in daylight. The same check catches the near misses. A workload with two copies and a cap of none down has a floor of two and is equally unsatisfiable. And a workload whose copy count falls to its floor during quiet hours becomes unsatisfiable only at night - which is precisely when maintenance runs.

  • The workload genuinely cannot run two copies at once. What now?
    Then the maintenance is an outage and should be planned as one: pick a window, announce it, remove the copy deliberately and let the replacement come up on another host. A cap that forbids this does not prevent the outage - it only prevents anybody from choosing when it happens.
  • Is a cap of 'none may be down' ever the right setting?
    Yes, for a workload whose copies are few and whose loss is expensive, where you want a human in the loop rather than an automation emptying hosts unattended. It becomes a defect when nobody owns the alarm it raises, because then it silently blocks every window instead of prompting the decision it was meant to force.

A shop with one till and a rule that the till may never close. The rule does not give the shop a second till; it just means the till can only ever close by accident.

saying these in an interview costs you the question

  • A none-down cap makes a single-copy workload highly available
  • A blocked drain forces the removal once it has waited long enough
  • The fix for a blocking cap is always to relax the cap
  • Adding a second copy is free for any workload
  • The cap also protects the copy when its host fails