All three copies of a stream sit on the same host in one cluster. What does the copy count still protect against?
answer
- the count promises independence, not copies
- what do these copies share?
- a copy buys unshared domains only
- sharing the leader's host buys least
basics
~20 sVery little: a single broker process failing, and one volume failing if the copies are on separate volumes. Host, storage-array, rack and power failures take all three together, so the count bought almost no independence.
solid answer
~50 sA copy is worth exactly the failure domains it does **not** share with the others. Three copies on one host still cover one broker process crashing or being restarted, and - if each copy is on its own volume - one volume failing. They cover nothing above that line: the host, its power supply, the storage array behind those volumes, the rack and the zone are all shared, so any one of those events takes all three at once. That is why a copy count is never a durability statement on its own; it is a depth, and the placement rule that separates copies across hosts, racks or zones is the axis. The copy you most need to place away from the others is the one sharing the leader's host, because a leader failure and a host failure are then the same event.
go deeper
Remember that copies help only when they are not in the same place. If someone tells you a stream keeps three copies, the useful follow-up is where those three copies live.
Explain the mechanics: failure domains nest from process to volume to host to rack to zone, and copies separated at one level survive that level and everything below it. Read a layout and name what it does not survive.
Demonstrate the production habit - verify the actual placement rather than the configured number, check what sits underneath the nodes, and re-verify after node replacements, because a correct layout drifts.
Own the standard: define levels as a count plus a placement axis, decide which streams justify zone separation given the extra traffic each copy costs, and make the constraint enforced at creation rather than audited later.
## The rule in one line An extra copy buys you **exactly the failure domains it does not share with the copies you already had**. Everything else in this subject is that sentence applied to a specific layout. So "three copies" is not a durability claim. "Three copies on three hosts in three racks" is. The number sets how many independent failures you can absorb; the **placement constraint** sets what counts as independent. ## Failure domains nest Write them out, innermost first, and any layout becomes readable: 1. **The broker process** - one server process on one node. 2. **The volume** - the disk or logical volume the copy is written to. 3. **The storage array or host bus** the volumes hang off, if several volumes share one. 4. **The host** - the physical or virtual machine, its memory, its power supply. 5. **The rack or power/network unit** - a shared feed and a shared top-of-rack path. 6. **The availability zone** - a building or hall with its own power and cooling. 7. **The region** - which is generally out of scope for copies inside one cluster. Two copies separated at level *k* survive the loss of one unit at level *k* and everything nested below it. Three copies all inside one unit at level 4 survive only failures at levels 1 to 3. ## What three copies on one host actually cover | Where the other copies sit | Broker process crash | One volume fails | Host is lost | Rack or zone power loss | |---|---|---|---|---| | Same host, same volume | Survived | Lost | Lost | Lost | | Same host, separate volumes | Survived | Survived | Lost | Lost | | Different hosts, same rack | Survived | Survived | Survived | Lost | | Different racks or zones | Survived | Survived | Survived | Survived | The first two rows are the answer to the question, and they are also the two rows nobody chooses deliberately. ## How copies end up sharing something anyway This layout is almost never an explicit decision. It is what happens when nobody expressed a constraint: - **No domain information.** If nodes carry no label saying which host, rack or zone they are in, a placement algorithm has nothing to spread across and will use free capacity instead. Some platforms enforce a spread automatically once the labels exist; others leave the choice to whoever created the stream. - **Several broker nodes on one machine.** Two nodes scheduled onto the same physical host - common when the cluster runs on a shared scheduler, and common in test clusters that quietly became production - look like two nodes and fail like one. - **Volumes carved from one array.** Separate volume names, one set of controllers and one power feed. - **A hand-written assignment.** Someone placed the copies once, the cluster later grew, and nothing re-examined it. - **A replaced node.** A copy re-created after a failure lands wherever there was room, which may be beside its leader. The last one matters most in practice: a layout that was correct at creation is not self-maintaining, so the placement is worth re-checking after any node replacement. ## What the count is still worth, and what it costs Do not conclude that only placement matters. The two answer different questions: - the **count** fixes how many separate domains can fail before the stream is unreadable - two copies in two zones survive one zone, three in three survive two; - the **placement constraint** fixes which domains those are. And the count is not free. Each copy is a full stored replica, so one incoming byte becomes N stored bytes and roughly N-1 bytes crossing the cluster - **write amplification**. Raising the count to compensate for copies that are not separated spends that traffic and buys nothing, which is the single most common wrong reflex on this subject. ## What to check, and how to say it Given a stream and a claim, ask three questions in order: how many copies does this stream keep, which domains are they separated across, and what do the nodes holding them share underneath - machine, array, rack, power. If the answer to the second is "nothing in particular", the first number is decoration. Say that plainly in an interview: **the count is the depth of the guarantee, the placement is its axis, and neither is worth anything alone.**
- Two copies are on separate volumes of one storage array. Does that count as separation?Only against a single volume failing. The array, its controllers and its power are shared, so a failure there takes both copies at once. Separation has to be judged against the unit that actually fails together, which is why a placement constraint is expressed in failure domains - host, rack, zone - rather than in volume names.
- If placement is what matters, why does the count matter at all?The count fixes how many separate domains can fail before the stream is unreadable; placement fixes which domains those are. Two copies across two zones survive one zone; three across three survive two. Depth and axis are independent choices, and quoting one without the other says nothing.
- A stream's copies were correctly spread a year ago. Why re-check them now?Because a layout is not self-maintaining. Nodes get replaced, a cluster grows, a copy re-created after a failure lands where there was room. Unless the platform enforces the spread continuously from domain labels, the arrangement drifts back towards whatever is convenient, silently and without any change to the number.
Three spare keys are only spare if they are not all in one pocket. Lose the pocket and the count never mattered; what you were really buying was the second pocket.
saying these in an interview costs you the question
- Says three copies is three copies, wherever they sit.
- Assumes the cluster always spreads copies across hosts itself.
- Forgets two nodes can share one machine or one array.
- Thinks separate volumes on one host survive host loss.
- Raises the copy count to fix a placement problem.