skip to content

Every partition showed a full copy set, yet one rack outage took partitions offline — how did copies end up sharing a domain?

level: seniorimportance: should knowfreq 46%

answer

  1. count is visible, placement is not
  2. the layout drifts one node at a time
  3. labels missing, defaulted, or simply wrong
  4. join copies to node domain labels
  5. audit on a schedule, not mid-outage

basics

~20 s

The layout was correct when built and drifted afterwards. Nodes replaced one at a time rejoined with a missing or wrong domain label, so the cluster placed copies as though every node were independent. Counts stayed full; the spread behind them did not.

solid answer

~50 s

A cluster reports how many copies exist, not how many domains they occupy, so co-location is invisible to every health signal you normally watch. The usual path is drift: nodes get replaced one at a time over months, and each replacement rejoins with whatever domain label provisioning gave it — often none, sometimes the wrong one, sometimes a default shared by every unlabelled node. Placement is then computed on a fiction, and it lands two or three copies of the same partition in one rack while still reporting a complete copy set with every copy caught up. The detection is an inventory join, not a metric: group each partition's copies by the domain label of the node holding each, and look for a partition whose copies are not spread. The durable fix is to set the label at provisioning and to audit the join on a schedule instead of during an outage.

go deeper

for a junior

The takeaway is that a cluster can look completely healthy and still be one rack away from losing a partition, because the numbers on the screen count copies rather than locate them.

for a middle

Explain the drift mechanics: replacement nodes, missing or wrong domain labels, and platforms that prefer spread rather than requiring it. Say why none of it shows up in a copy count or a sync check.

for a senior

Describe the check you would actually run — copies per partition grouped by the domain label of the node holding each, on a schedule and after node changes — and cover the membership holding cluster metadata in the same sweep.

for a principal

Treat it as a provisioning contract rather than an audit: the label comes from the system that knows where the machine is, an unlabelled node cannot join, and the per-domain view sits beside the copy count everyone already watches.

## What the counts do and do not tell you Operators watch two numbers about copies: how many exist, and whether they are caught up. Both can be perfect while the layout is one rack away from an outage, because neither carries any information about **where** the copies are. | Signal | What it answers | What it cannot answer | |---|---|---| | copies present per partition | are all copies there at all | which domains they sit in | | copies caught up | do the copies hold the same records | which domains they sit in | | node inventory with domain labels | which domain each node claims to be in | whether the claim is true | | copies per partition grouped by domain | how much one domain loss removes | nothing else — this is the check | The last row is the only one that answers the question, and it is almost never on a dashboard. ## The three ways a good layout drifts 1. **Replacement, one node at a time.** A cluster built deliberately across three racks loses a node, and a replacement is provisioned from whatever pool has space. Repeat over a year of hardware refreshes and the careful original arrangement is gone, one machine at a time, with no single change big enough to review. 2. **Labels missing, defaulted or wrong.** A cluster spreads copies only if it has been told which domain each node is in. A node that rejoins with no domain label is treated by some platforms as a domain of its own — which makes co-located nodes look independent — and by others as belonging to one shared default bucket, which piles every unlabelled node together. The two behaviours fail in opposite directions, and both are computed as if they were the truth. 3. **Placement that prefers rather than refuses.** Some platforms refuse to create a stream that cannot be spread; others treat spread as a preference and will place a copy anywhere rather than leave it unplaced — for instance when one domain is temporarily short of capacity. The second kind produces a complete, healthy, co-located copy set and says nothing. A fourth path is worth naming even though it is not drift: the labels describe the wrong boundary. Three racks on one breaker are three labels and one failure domain, and every arithmetic downstream is wrong in a way no cluster can detect. ## Finding it before the outage does The check is mechanical: - take the cluster's own inventory of which node holds each copy of each partition; - join it to the domain label of each node; - count copies per partition per domain, and flag any partition where one domain holds enough of the copy set to take it out; - run the same check against the membership holding cluster metadata, which is small enough that people assume it is fine; - flag any node with no domain label at all as a defect in its own right, because it makes the rest of the arithmetic meaningless. Run it on a schedule and after every membership change. A cluster that has been running for two years without this check almost certainly has some drift in it; a cluster that has had a hardware refresh certainly does. ## Making the fix stick The repair itself — moving copies until the spread is right again — is an ordinary re-placement operation with its own risks, and it is the easy half. The half that keeps it from coming back is procedural: - the domain label is set at provisioning, from the same source of truth that knows where the machine physically is, not typed in afterwards; - a node that comes up without a label is not allowed into the cluster, or is quarantined until it has one; - the per-domain copy count is audited on a schedule and after node changes, with the result visible next to the copy count that everybody already watches; - the boundary the labels claim is reviewed against reality occasionally, because power and network topology change without anybody updating a label. ## On a rented cluster Where the cluster is a managed service, node replacement happens without you and the labels are not yours to set. The drift question becomes a vendor question: what the service guarantees about spreading copies across domains, whether it is a guarantee or a best effort, and whether it reports the per-domain layout at all. A vendor that places across domains by contract removes this failure; a vendor that merely tends to do so has the same drift risk with none of the visibility.

  • Why does a caught-up-copies health signal never catch this?
    Because it measures agreement, not location. It asks whether the copies hold the same records as the copy that leads, and a set of copies sitting on three nodes in one rack can be perfectly caught up right up to the moment the rack goes. Placement is a separate fact and needs a separate check.
  • Is an unlabelled node worse than a mislabelled one?
    Both feed the placement arithmetic a fiction, and which is worse depends on the platform. Some treat a missing label as a unique domain, so co-located nodes look independent; others group all unlabelled nodes into one bucket, which over-constrains placement instead. A wrong label lies consistently, which is at least easier to spot in an inventory.
  • How do you check this on a cluster whose nodes you do not provision?
    You ask the vendor what it guarantees rather than inspecting racks. The useful questions are whether spreading copies across domains is contractual or best-effort, how many domains a cluster spans, and whether the per-domain layout is exposed at all. Without an answer, treat the spread as unverified.

saying these in an interview costs you the question

  • reads a full copy count as proof the copies are spread
  • assumes a replaced node keeps the domain label of the one it replaced
  • checks whether copies are caught up instead of where they sit
  • assumes every cluster refuses to co-locate two copies of one partition
  • waits for a domain outage to learn the real layout
  • labels racks without checking they fail independently