A host drain has sat for an hour with three copies left and a cap on how many may be voluntarily down - what is happening?
answer
- a request, not a deletion
- checked before every removal
- floor equals copy count minus cap
- refused and retried, not failed
- an hour means availability never returned
basics
~20 sThe drain is blocked, not broken: every removal is a request checked against the cap, and while the workload is already at the floor of available copies, granting one more would breach it, so the request is refused and quietly retried.
solid answer
~40 sA drain removes copies by *asking*, not by deleting. Before each removal the platform compares the workload's currently available copies against the floor the cap implies - for six copies with at most one down, the floor is five. If removing one more would put availability below that floor, the request is refused and retried a moment later, so the drain waits instead of proceeding. A few minutes of that is normal. An hour of it means availability is not coming back, so the retry can never pass. The usual causes are no room on the remaining hosts for the replacement, a replacement that is placed but never becomes available, or another copy already down for an unrelated reason. Read the workload's available count against its floor, not the drain's status, to tell which.
code
pseudocode · 12 linesclose_host_to_new_placement(host)
for each copy on host:
while true:
available = count_available(copy.workload) # 5
floor = copy.workload.replicas - maxDown # 6 - 1 = 5
if available - 1 >= floor: # 5 - 1 = 4, not >= 5
request_stop(copy) # not reached
break
wait(retry_interval) # this branch firesgo deeper
The key idea is that a drain asks rather than deletes, and an ask can be refused. If you remember only one thing, remember that a waiting drain is usually a refused request being retried, not an error.
Be able to do the arithmetic out loud: copy count minus the cap gives a floor, and a removal is granted only if availability stays at or above it. Then say what the drain does with a refusal - retries it, indefinitely.
Diagnose it. Say which number you read first, name the three reasons availability fails to recover, and distinguish a cap that is working from capacity that ran out. Restarting the drain should not appear in your answer.
Decide in advance what happens when a drain outlives its window: who is allowed to breach a cap, how it is recorded, and whether an unpatched host or an unscheduled dip is the cheaper outcome for that workload.
## A removal is a request, not a delete A drain does not delete copies. For each copy on the host it submits a **removal request**, and the platform evaluates a precondition before granting it: would the workload still honour its disruption cap once this copy is gone? If yes, the copy is asked to stop. If no, the request is **refused**, and the drain retries it shortly afterwards. Nothing errors and nothing gives up, which is exactly why a blocked drain and a slow drain look identical from outside. ## Why the check is re-evaluated every time The cap is not a quota the drain draws down once at the start; it is a precondition re-evaluated on every request against availability at that instant. That is what makes it safe when several things disrupt the same workload at once - a second operator draining another host, an automation replacing copies under a new spec, a repack running unattended. Each of them is refused the moment the budget is spent, without any of them needing to know the others exist. ## The arithmetic State the numbers. The ingester declares **6 copies** and a cap of **at most 1 voluntarily down**, so the **floor** - the number that must remain available - is 6 - 1 = 5. Three of the six copies sit on tonight's host. | Step | Available now | Floor | Removal granted? | |---|---|---|---| | First attempt | 6 | 5 | yes - 6 - 1 = 5, still at the floor | | That copy stops | 5 | 5 | - | | Second attempt | 5 | 5 | no - 5 - 1 = 4, below the floor | | Replacement becomes available | 6 | 5 | yes, the retry now passes | That cycle is the healthy case, and it runs at the speed a replacement takes to become available: a couple of minutes per copy for a fast-starting service, far longer for one that loads state at start-up. Two remaining copies at ten minutes each is a twenty-minute drain, and it is not stuck. ## An hour is a different diagnosis An hour of refusals means availability never returned to 6, so the retry cannot pass. Three causes account for nearly all of it: 1. **The replacement has nowhere to go.** The remaining hosts have no room for another copy, so it is never placed and never becomes available. The drain is blocked on capacity; the cap is only the thing reporting it. 2. **The replacement is placed but never becomes available.** It starts, fails its check, restarts, and the available count never rises. The cap is doing precisely what it is for: refusing to take a healthy copy away while the fleet is already one short. 3. **Another copy is already down for an unrelated reason** - a host that failed on its own, an earlier drain nobody finished. The budget was spent before tonight's window even opened. ## Read the count, not the drain The drain's own status tells you only that it is waiting. The number that tells you *why* is the workload's available copy count against its floor. If available is pinned one below where it needs to be and is not moving, look at the copy that should be replacing the one you removed: is it placed at all, and if it is placed, is it reporting itself available? Those two questions separate cause 1 from cause 2, and neither is answered by restarting the drain. ## What the platform will not do for you The platform will not quietly breach the cap because a window is closing - the whole value of the cap is that it cannot be breached by accident. **Designs differ on whether it can be breached on purpose:** some platforms expose a forced removal that skips the check and records who ordered it, and others simply refuse indefinitely. Policy therefore has to say what happens when a drain outlives its window, because the platform will not decide that for you. The legitimate unblocks all work by creating headroom rather than by removing the guarantee: - raise the workload's copy count, so the floor can be met with the drained copy gone; - free room on the remaining hosts, so the replacement can actually be placed; - fix whatever stops the replacement reporting itself available. And one thing that is not an unblock: waiting. A refusal that has repeated for an hour under unchanged conditions will repeat all night.
- Can an operator override a cap that is blocking a drain?Designs differ. Some platforms expose a forced removal that skips the check and records that somebody chose to breach the guarantee; others only ever refuse. Where a force exists it is the right tool for an emergency and the wrong one for a routine window, because it silently converts the workload owner's guarantee into the operator's problem.
- The drain is blocked and the maintenance window closes in twenty minutes. What is the fastest legitimate unblock?Give the workload somewhere to go. Raise its copy count, or free room on the remaining hosts, so a replacement becomes available and the floor is satisfied with the drained copy gone. Both restore headroom without touching the cap, which is the guarantee you are trying to keep intact.
- Two operators drain two hosts at the same time. What stops them together taking the workload below its floor?Nothing coordinates the operators, but nothing has to: the cap is re-evaluated against current availability on every request. Whichever request arrives when the budget is already spent is refused, so the second drain simply waits. That is why the check is a precondition per removal rather than a quota handed out at the start.
saying these in an interview costs you the question
- The drain is stuck, so force-delete the copies and move on
- A stalled drain always means the cap is set wrong
- The cap delays removals but eventually lets them all through
- A drain that waits long enough will finish on its own
- Availability counts copies that exist, not copies reporting available