skip to content

Three copies of one partition sit on three record-serving nodes in the same rack — what does that placement protect against?

level: juniorimportance: must knowfreq 62%

answer

  1. count is not the same as spread
  2. what fails together, fails together
  3. one rack, one power feed, one switch
  4. three copies, one domain, one failure
  5. the metadata role is placed too

basics

~20 s

Node-level failures only: a disk, a process, one machine. The rack is a single power and network domain, so one rack event takes all three copies at once — three copies against one failure class, one copy against another.

solid answer

~50 s

A copy count and a copy spread are different facts, and only the count is reported loudly. Three copies on three record-serving nodes in one rack survive a disk, a crashed process, a re-imaged machine or one node taken out deliberately. They do not survive the rack's power feed, its switch or uplink, a cooling event, or maintenance scheduled against the rack as a unit — any of those takes every copy held there at once, and the partition with them if no copy survives elsewhere. The placement rule that follows is that copies of one partition are placed so no single failure domain holds enough of the copy set to take the partition out with it. The same reasoning applies to the small membership that holds cluster metadata: spreading the data and leaving that membership in one rack is a half-done job.

go deeper

for a junior

Remember the one-line version: copies sitting in the same rack are one copy as far as that rack's failure is concerned. Be able to name what a rack shares — power, switch, cooling — and why that makes it a single unit of failure.

for a middle

Explain the mechanics: which failures a co-located copy set survives, which it does not, and that a cluster can only spread copies if it has been told which domain each node is in. Say what the count does and does not report.

for a senior

Show that you check placement rather than trusting the copy count, and that you include the membership holding cluster metadata in the same check. Be ready to say what one domain loss would remove from a cluster you actually run.

for a principal

Frame it as a standing rule rather than a per-cluster choice: which unit counts as a domain across the estate, who enforces it at provisioning time, and what the extra hop between domains puts into the write path.

## What a failure domain is when you are placing copies A **failure domain** is the smallest unit of correlated failure: the boundary inside which equipment fails together. In most estates the first such unit is a rack — one pair of power feeds, one top-of-rack switch, one cooling path — then a building or zone, then a region. What makes something a domain is not its size but the existence of a single event that takes everything inside it at once: a breaker, an uplink, a cooling loss, or a maintenance action scheduled by somebody who does not know what is in there. A **copy** here means one stored replica of one partition (or of one queue), held by a **record-serving node**. **Placement** is the question of which node — and therefore which domain — each copy in a copy set lands in. That is a different question from how many copies exist, and the two are confused constantly, because a cluster reports the count in every console and the spread in none. ## What three copies in one rack buy They genuinely survive: - a single disk failing under one of the three copies; - a single broker process crashing, or being restarted for a patch; - one machine being rebuilt, re-imaged or replaced; - one node taken out deliberately by an operator. They do not survive: - the rack's power feed, or one of a redundant pair failing while the other is already out; - the rack's switch or its uplink to the rest of the network; - a cooling or fire suppression event in that row; - a maintenance action taken against the rack as a unit. So the honest description of that layout is: **three copies against node-level failure, one copy against rack-level failure**. The count is not wrong; it is being read against the wrong failure. And the rule that follows is short — copies of one partition are placed so that no single domain holds enough of the copy set to take the partition out with it. ## Reading a layout honestly Three questions settle whether a layout is real: 1. **Which domains exist, and which is the one that fails as a unit?** A row of racks on one breaker is one domain, not eight, whatever the inventory says. 2. **What does losing the largest domain remove?** Count the copies of one partition inside it, not the nodes. 3. **Who tells the cluster which domain a node is in?** A cluster spreads copies only if it has been told the domains; with no such information it treats every node as equivalent and may put the whole copy set in one place while reporting the full count. ## The metadata role is placed too The cluster's own metadata — which streams exist, which node holds which copy, who leads what — is held by a small membership rather than by every node. Spreading record-serving nodes across three racks while leaving that membership inside one is a common half-done job: the data survives the rack and the cluster's ability to react to the rack loss does not. The membership is spread on the same principle as the data. ## Where platforms differ What a domain loss actually removes depends on how the platform stores copies at all, and this is where a claim learned on one platform stops being true: | How the platform stores copies | What losing one domain removes | |---|---| | One copy leads a partition and others follow it | every copy of that partition held in the lost domain, the leader included if it was there | | A write commits when a majority of the copy set holds it | that domain's share of the copy set; if the share was half or more, no majority remains | | Record-serving nodes share one storage system instead of holding copies | nothing, if that storage system is itself spread — and everything if it is not; the placement question moves onto it | | A queue is mirrored to a small number of other nodes | the mirrors held in the lost domain, which is why queue-shaped brokers need the same placement rule | ## What spreading takes back Crossing a domain puts a network hop between the node that accepted a write and the node holding another copy of it. A write that must reach a second domain before it is acknowledged therefore pays an extra round trip, and the further apart the domains sit, the larger that round trip is. That is the standing constraint on how far apart you are willing to put them, and it is why "spread wider" is not automatically better. Within one building the extra hop is close to free; across a continent it is in the path of every acknowledged write. The short version a junior is expected to carry: **copies in the same domain are one copy against that domain's failure**, and the cluster will not tell you that unless you ask it where things are.

  • If a lost domain holds only a minority of a partition's copies, what happens to that partition?
    It keeps serving from the copies that survived, because enough of the copy set is still reachable to accept and serve records. It is running degraded until the lost copies are rebuilt elsewhere, so a second failure during that period is far more dangerous than the first.
  • Does any of this apply to a queue-shaped broker that has no partitions?
    Yes. Wherever the platform keeps redundant copies of a queue, those copies must not share a domain, and the reasoning is identical. Where nodes instead share one storage system, the per-node copies disappear and the question moves to whether that storage system is spread — it does not go away.
  • Why does a cluster sometimes place every copy in one domain without complaining?
    Because it places on the information it has. If nobody told it which domain each node is in, every node looks equivalent and any placement satisfies the count. Some platforms then refuse to guarantee spread, others silently place anywhere, and the console shows a full copy set in both cases.

saying these in an interview costs you the question

  • treats three copies as three survived failures regardless of where they sit
  • calls a rack too small to count as a failure domain
  • assumes a cluster spreads copies without being told which domain each node is in
  • counts nodes sharing one power feed and one switch as independent failures
  • designs only against a whole-region outage and ignores the rack
  • spreads the record-serving nodes and leaves the metadata role in one place