skip to content

Records were acknowledged by three caught-up copies in one cluster, then a shared power failure lost them — why?

level: seniorimportance: should knowfreq 46%

answer

  1. three holders, zero on media
  2. one event, three memories
  3. copies cover separate failures
  4. forcing covers simultaneous ones
  5. more copies does not shrink this window

basics

~20 s

All three copies held the record in volatile memory rather than on persistent media, and one event took the three machines together, so it took the unforced span on each of them. Copies cover machines failing separately; forcing covers them failing together.

solid answer

~50 s

Three copies accepting a record means three machines have the bytes, not that any of them wrote them down. Each one put the record in its operating system's file cache and answered; the forced write was still pending on all three. When a single event removes those machines at the same moment, it removes each machine's unforced span at the same moment, and the acknowledged records inside it are gone. This is not a failure of the acknowledgement rule — it did exactly what it was configured to do. Copies and forcing answer two different questions: copies protect a record when one machine dies on its own, because the survivors still hold the bytes and will write them out; forcing protects it when the machines die together. If acknowledged records must survive the second case, the acknowledgement has to wait for the bytes to be on persistent media, which is a different setting with a different price.

go deeper

for a junior

Take away the headline: several machines holding a record in memory is not the same as the record being written down, and one event can empty all of that memory at once.

for a middle

Explain why the acknowledgement rule was not at fault, and state the split plainly: copies answer machines failing separately, forcing answers machines failing together.

for a senior

Diagnose without reflexes. Quote the window that was exposed, identify what made the failure simultaneous, and price the lever you propose in write latency and sustained throughput before recommending it.

for a principal

Turn it into a standard: a small number of durability tiers, each naming its copy count, acknowledgement rule and forcing policy, with a published loss window per tier and a stated default for streams nobody reviews again.

## What actually happened Walk the record through the cluster: 1. The writer sends it. The leader copy appends it, the operating system puts it in its **file cache**, and the append returns. 2. Two follower copies pull the record and do the same thing on their own machines: appended, cached, not forced. 3. The acknowledgement rule — whatever the operator set it to, up to and including every caught-up copy — is satisfied. The writer is told the record is stored and moves on. 4. A single event removes all three machines at once. 5. They come back. The record is on none of them. Every step behaved as designed. Three machines genuinely held the record. What none of them had done was put it on persistent media, and the event did not politely take them one at a time. ## Why the copies answered a different question This is the core of the subject, and it is worth stating as two separate sentences rather than one: - **Copies protect an unforced record against machines failing separately.** One machine dies; the other two are untouched, still hold the bytes in their caches, and will write them out in the normal course of things. The record survives. This is the overwhelmingly common failure, which is why the posture works most of the time. - **Forcing protects an unforced record against machines failing together.** Nothing about having more holders helps if they all lose their memory in the same instant. Adding a fourth or fifth copy does not change the second sentence at all, which is why "we lost data, raise the copy count" is the wrong reflex here. Spreading copies so that fewer of them share a power feed reduces how often the together-case arises — that is a placement decision, made elsewhere in the cluster's design — but it does not alter the fact that an unforced record lives only in memory. ## Events that take machines together An operator does not need a taxonomy here, just a short concrete list of things that actually remove several machines at the same moment: - a shared power feed, a rack, or a wider facility-level event; - a host or hypervisor fault affecting every guest on one physical machine; - an orchestrated action that stops or restarts many nodes in one go; - a change rolled out everywhere at once that makes the nodes stop in the same way. The common thread is a shared cause, not a shared location, and that is why "they were on three different machines" is not by itself an answer. ## The two levers and what each buys | Lever | Protects against | Cost | |---|---|---| | More copies of the stream within the cluster | One machine, or a few, failing separately | Stored bytes and catch-up traffic | | Waiting for the bytes to be forced before answering | Machines failing together, taking their memory | Write latency and sustained throughput | | A tighter forcing policy without waiting for it | Shrinks how much is exposed in the together-case | Throughput, proportional to how tight it is | The middle row is the only one that makes the scenario impossible rather than smaller, and it is the expensive one. Most estates choose the third row for most streams and the middle row for a few, which is a defensible posture as long as it was chosen rather than inherited. ## What varies across platforms - Some platforms let the wait for a forced write be expressed as part of what the writer waits for; others offer only a forcing policy running behind the acknowledgement. If the second is all you have, the together-case cannot be eliminated on that platform, only bounded. - On designs where durability comes from an **underlying shared replicated store**, the exposure is the store's, not the broker's, and it is its contract that has to be read. - On a **rented cluster**, what the provider forces and when is often not visible. The honest posture statement is what the provider publishes, not an assumption. - Hardware with a power-loss-protected device cache narrows the window's price sharply, which can make waiting for a forced write affordable on the few streams that need it. ## How to answer this in an interview Name the gap in one line — three copies had it in memory, none had it on media — then say which lever addresses the specific event, then state the price of that lever. Candidates who stop at "they should have had more copies" have not separated the two questions, and candidates who conclude "forcing should always be on" have not priced it.

  • Would raising the copy count have prevented this loss?
    No. Every additional copy is another machine holding the record in memory. If the event takes those machines together, it takes each of their unforced spans too. Raising the count improves survival when machines fail separately, which is a different and far more common event, but it leaves this scenario exactly as it was.
  • The team proposes waiting for a forced write before every acknowledgement. What do you ask before agreeing?
    Which streams, and at what price. It puts a device round trip in every writer's latency and caps sustained throughput at the hardware's forced-operation rate. That is usually right for a small number of streams whose records cannot be reconstructed, and wrong as an estate-wide default, so the answer is a tier, not a switch.
  • If the platform cannot make an acknowledgement wait for a forced write, what is left?
    Bound the exposure and declare it. Tighten the forcing policy as far as throughput allows, quote the resulting window in time and in bytes at peak, and tell the stream's owners that a simultaneous loss of the holding machines costs that much. An honest published number beats an implied guarantee the platform cannot make.

saying these in an interview costs you the question

  • Concludes the fix is simply more copies of the stream
  • Claims the acknowledgement rule must have been misconfigured
  • Thinks copies on different machines implies bytes on media
  • Treats simultaneous loss as impossible rather than priced
  • Proposes forcing every write with no throughput cost stated
  • Assumes every platform can make a write wait for forcing