skip to content

A host must be drained tonight for a firmware upgrade - what does that drain actually do, and how does it differ from the host failing?

level: middleimportance: must knowfreq 62%

answer

  1. who decided, and was anyone asked
  2. two kinds of disruption
  3. closed to new placement first
  4. one removal request at a time
  5. caps govern only the chosen kind

basics

~20 s

A drain is a disruption the platform is told about first: it stops placing new work on the host, asks the copies already there to stop one at a time, and waits for replacements. A host failure announces nothing and the copies are simply gone.

solid answer

~50 s

Draining is *voluntary* disruption - an operator, not the hardware, decided it, so the platform gets to sequence it. It happens in three moves: the host is closed to new placement so nothing lands while it empties; each copy on it is then asked to stop, one request at a time, and only if a cap on how many copies of that workload may be voluntarily down still allows it; and the drain waits until the host reports empty. Because each removal is a request that can be refused, a drain can take minutes or hours, and that is the mechanism working rather than something stuck. An unplanned loss is the opposite: nothing is asked and nothing is sequenced, and the platform finds out afterwards because the host stopped reporting. The cap holds back voluntary disruption only; it cannot hold back hardware.

go deeper

for a junior

Know the two words and which is which: a drain is disruption somebody chose and announced, a host failure is disruption nobody got to announce. Being able to sort examples into those two buckets is the whole first answer.

for a middle

Explain the order - close the host to new placement, then remove copies one request at a time, then wait for the host to report empty - and why skipping the first step lets a removed copy be placed straight back onto the host you are emptying.

for a senior

Show you have watched one run. Name the step that waits, say roughly how long draining a busy host takes and why, and be clear that a slow drain is usually the cap doing its job rather than a process that has hung.

for a principal

The tradeoff is who absorbs maintenance cost. Tight caps make drains slow and windows long; loose caps make patching fast and small dips likelier. Decide it per workload and make the workload owner, not the operator, live with the choice.

## Two kinds of disruption A **disruption** is anything that takes a running copy of a workload away. Platforms sort disruptions by one property, and nearly everything about planned maintenance follows from it: **was anybody able to ask first?** - **Voluntary disruption** is disruption somebody chose - emptying a host so its firmware can be flashed, repacking a fleet onto fewer machines, retiring old hardware. It is announced before it happens, so the platform can pace it, sequence it, and refuse a step that would go too far. - **Involuntary disruption** is disruption that is already over by the time anything notices - a power cut, a kernel panic, a cut link, a host that simply stops reporting. There is nothing to pace, because there is nothing left to negotiate with. A drain is the first kind, performed deliberately, and it is the reason the platform has a vocabulary for planned maintenance at all. ## What a drain actually does Assume the workload being drained is a device-telemetry ingester running **6 copies over 12 hosts**, with **2 of them on tonight's host**. Draining that host is four moves, in order: 1. **Close the host to new placement.** The scheduler stops choosing it. Without this step, a copy removed in move 2 can be placed straight back onto the host you are trying to empty. 2. **Ask each copy on the host to stop, one request at a time.** Each request is checked against the cap on how many copies of that workload may be voluntarily down at once, and is refused if granting it would breach the cap. 3. **Wait.** The next removal is granted only once the workload is back at or above the floor the cap implies - that is, once a replacement has become available somewhere else. 4. **Report the host empty.** Only then is it safe to power it down. With a cap of at most one copy down, availability across the drain moves 6, 5, 6, 5, 6: the two copies leave in sequence and never together. That sequencing is the entire product of a drain. ## Side by side | | Planned drain | Unplanned loss | |---|---|---| | Who decided | an operator or an automation | nobody | | Announced beforehand | yes | no | | Can be paced or refused | yes | no | | Copies leave | one request at a time | together, instantly | | Held back by a disruption cap | yes | no | | How the platform learns of it | it initiated it | the host stops reporting | ## Why the platform keeps the two apart Only one side is negotiable, so only one side can be governed. A **disruption cap** - a rule saying at most so many copies of this workload may be voluntarily down at once - is a brake on automation, not a guarantee of availability. Three hosts failing in the same instant will take three copies with them and the cap will have nothing to say about it. Reading a cap as 'at most one copy of this workload is ever missing' is the most common misreading of the whole mechanism. The split also decides who is on the hook. Voluntary disruption is scheduled work whose cost can be paid deliberately - spread across the night, one failure domain at a time, with a wait between each removal. Involuntary disruption is paid whenever it arrives, which is why a workload that cannot survive losing a copy needs **more copies**, not a stricter cap. ## Reading a drain while it runs Three numbers tell you what is happening, and none of them is the drain's own status line: - **Copies still on the host.** This is what the drain is trying to get to zero. - **Copies of that workload currently available across the fleet.** This is what the cap compares against its floor before granting the next removal. - **The floor itself** - the declared copy count minus the cap. When available equals the floor, the next removal is refused by design. If the second number is not climbing back, the drain is not slow; it is waiting for something that will not happen on its own. ## What a drain does not do - It does not decide **where** the replacements land; that is the scheduler's placement problem. - It does not change what a copy does between being asked to stop and actually stopping. - It does not create capacity. If the remaining hosts have no room, the removal is refused and the drain waits. - It does not protect the copies against the host dying on its own halfway through the drain. - It is not a verdict on the workload's health: a slow drain usually means the cap is working.

  • Does a drain guarantee that no work in flight is lost while a copy is being removed?
    No. The cap controls how many copies may be missing at once, not what happens to requests already accepted by the copy that is stopping. Finishing accepted work and refusing new work is the stopping copy's own responsibility, and it is a separate mechanism from the drain - a drain can be perfectly paced and still drop work if the copy ignores the request to stop.
  • The host is empty and powered off. What is still true of the workload's disruption cap?
    Nothing changed. The cap is a standing rule on the workload, not a per-drain setting, so it is evaluated again on the very next voluntary removal - the next drain, repack or retirement. That is why a cap that is set wrong blocks every future maintenance window, not only tonight's.

Closing one lane of a bridge for resurfacing: the cones go out first so no new cars enter, the traffic already in the lane is let through, and the crew waits until it is empty. A lane collapsing does none of that.

saying these in an interview costs you the question

  • A drain and a host crash look the same to the platform
  • The disruption cap also protects against hardware failure
  • A drain just deletes the copies and the platform sorts it out
  • Once issued, a drain empties the host immediately
  • New work can still be placed on a host that is draining