Under a strict even-spread rule across three failure domains, why does a workload's fourth replica stay pending?
answer
- a hard rule is a filter
- counts, not capacity
- four copies, three domains
- permitted difference of zero
- pending means the rule is working
basics
~20 sA strict spread rule is a filter, not a preference. With three failure domains holding one replica each and no unevenness permitted, placing a fourth anywhere would make the counts uneven, so every host is ruled out and the replica waits.
solid answer
~50 sA spread rule declares three things: the dimension to spread across, how uneven the per-domain counts may become, and what to do when the rule cannot be honoured. Declared strictly with zero permitted difference, it becomes a hard condition in the filter stage. Three domains at one copy each is even; a fourth copy makes some domain hold two, a difference of one, which the rule forbids — so every host in every domain fails the filter and the copy stays pending indefinitely, with room to spare on all of them. Permitting a difference of one legalises counts of 2/1/1 and the copy lands. The other ways out are to make the rule best-effort, to run a replica count that is a multiple of the domain count, or to add a domain. Required anti-affinity at host scope fails identically when there are fewer hosts than copies.
code
yaml · 12 linesworkload: scoring-api
replicas: 4
placement:
spread:
across: failure-domain # host label naming rack, power feed or site
maxDifference: 0 # every domain must hold the same count
ifImpossible: leave-pending # the alternative is: place-anyway
# 3 domains, 4 replicas:
# placed 1/1/1 -> difference 0, rule satisfied
# a 4th makes 2/1/1 -> difference 1 > maxDifference 0 -> every host filtered out
# maxDifference: 1 would allow 2/1/1 and the 4th replica would be placedgo deeper
Remember that placement rules can rule out every host. A copy stuck pending while hosts have free room usually means a rule was not satisfiable, not that the cluster is full.
Explain the three parts of a spread rule — the dimension, the permitted unevenness, and what happens when it cannot be met — and work the counts for four copies over three domains.
Diagnose it from the rejection reasons, then choose deliberately between a pending copy and an uneven one, and know that nothing rebalances a lopsided distribution after a domain recovers.
Set the standard: which workloads must be spread across which dimension, whether strict is ever justified, and the headroom that has to exist so losing a domain is survivable rather than merely orderly.
## What a spread rule declares A spread rule is three decisions, and confusing them is where the incident comes from: - **The dimension.** A host label that names the failure domain — a rack, a power feed, a separate data-centre zone. Spreading across hosts and spreading across failure domains are different guarantees: six copies on six hosts in one rack survive a host failure and not a rack failure. - **How uneven counts may get.** The permitted difference between the fullest and emptiest domain. Zero means the counts must be equal; one means the fullest may hold one more than the emptiest. - **What to do when it cannot be honoured.** Leave the copy pending, or place it anyway as best effort. This is the setting that decides whether an unsatisfiable rule costs you a copy or costs you the guarantee. ## Why the fourth copy cannot land Three domains, one copy each: counts are 1, 1, 1 and the difference is 0. Placing a fourth copy makes some domain hold 2, so the counts become 2, 1, 1 and the difference is 1. With a permitted difference of 0 that is a violation, so the rule rules out every host in every domain — including hosts with abundant unreserved room. The candidate list is empty and the copy stays pending. Nothing will change that: the cluster is not short of capacity, and waiting does not help, because four copies can never be spread evenly over three domains. Raise the permitted difference to 1 and the same fourth copy places immediately, because 2/1/1 differs by exactly 1 and is allowed. The whole incident is one number. The same shape appears without any spread rule at all: a required anti-affinity rule saying no two copies of this workload may share a host, with four copies and three hosts, leaves the fourth pending forever for the same reason. ## Strict against best-effort | declared as | what it guarantees | how it fails | |---|---|---| | hard requirement | no domain ever carries more than its permitted share | a copy can sit pending indefinitely while hosts stand idle | | best effort | every copy runs | the distribution can quietly become lopsided and stay that way | Best effort has a second-order trap. Placement is a one-time decision, so copies created while a domain was unavailable land wherever they could, and nothing moves them back when that domain returns. The workload then runs with the distribution of the worst moment in its history until something replaces the copies again. ## Choosing between the two 1. Ask whether an unplaced copy is worse than an uneven one. For a service with headroom, uneven almost always wins, and the rule should be best effort with alerting on the skew. 2. Where the rule must be strict, make it satisfiable: run a replica count that is a multiple of the domain count, and check it stays satisfiable at the counts you will actually run at when the copy count changes with load. 3. Be explicit about the dimension. Per-host spread is not a domain guarantee, and a domain guarantee is not a host guarantee; some workloads want both rules. 4. State which failures the rule is bought for, so the cost of a pending copy can be argued against a real risk rather than a habit. ## What spreading buys, and what it does not It buys a **bounded** loss: losing one domain costs that domain's share of the copies instead of all of them, and that is the entire point. It does not buy capacity. With three domains evenly spread, losing one removes a third of the copies, so the survivors must already be sized to take the traffic — otherwise a domain failure turns into an overload of the two domains that stayed up, and the spread rule's only contribution was to slow that down. It also protects only against the dimension you chose. A bad image, a bad configuration change or a failed shared dependency reaches every copy in every domain at once, and no placement rule has anything to say about that. Finally, note what the rule is evaluated against: the copies that exist at the moment of each placement. It is not a one-off check performed when the workload was first created, and it is not a promise about the distribution afterwards. Every new copy is filtered against the counts as they stand then, which is why a rule that was satisfiable at three copies can block silently at four, and why a workload whose copy count changes over time should be checked against the counts it will actually run at rather than the one it was designed with.
- Which change gets the fourth copy running today without giving up the intent entirely?Raise the permitted difference from 0 to 1, so counts of 2/1/1 are legal, or declare the rule best effort so it degrades instead of blocking. Both keep spread as the intent. Adding a domain or moving to six copies also restores evenness, but costs infrastructure or money.
- What does spreading across failure domains not protect you against?Anything that is not that dimension: a bad image, a configuration change, a poisoned dependency. And it does not preserve capacity — with three domains, losing one removes a third of the copies, so the survivors must already be sized to absorb the traffic.
saying these in an interview costs you the question
- Says the cluster is out of capacity when the hosts plainly have room
- Assumes a spread rule always degrades to best effort when it cannot be met
- Treats spreading across hosts as equivalent to spreading across failure domains
- Thinks the platform rebalances copies once the distribution becomes uneven
- Believes spreading keeps full serving capacity when a domain is lost