skip to content

An operator floors caught-up copies at two on every stream, yet a leader failure still loses acknowledged writes; why did the floor never take effect?

level: seniorimportance: should knowfreq 46%

answer

  1. a condition, not a guarantee
  2. inert on the lower rungs
  3. two halves, two owners
  4. leader-only writes bypass it
  5. verify from the writing side

basics

~20 s

Because the floor is only consulted on a write that waits for every caught-up copy. Writers asking to be answered by the leader alone are answered by the leader alone, and the server-side floor never enters the decision.

solid answer

~50 s

The floor on caught-up copies is not an independent guarantee — it is a condition attached to the strongest acknowledgement rung. A writer that asks to be answered as soon as the leader holds the record gets exactly that, whatever the floor says, because the write was never going to involve a second copy in the first place. So the operator hardened one half of the pair and the writers kept asking for the weak half. When the leader is lost and a surviving copy is promoted, records that lived only on the lost machine are gone, and they were gone by policy, not by accident. The fix is to make the two halves agree: set the floor and also ensure the streams that need it are written with the rule that waits for every caught-up copy — and then verify from the writing side, not just the server side.

go deeper

for a junior

The takeaway to carry is simple: a server-side floor on copies only matters for writes that were going to wait for more than one copy anyway.

for a middle

Explain the pairing mechanically — which half the writer controls, which half the server controls, and why the floor is inert for the two lowest rungs.

for a senior

Demonstrate the diagnosis: read both halves per stream, use write latency as evidence of which rung is really in play, and prove it by removing a copy on a test stream.

for a principal

The organisational point is ownership: a durability posture split across a platform team and product teams drifts by default, so the standard has to be verifiable from one place.

## The shape of the failure This is the most common durability misconfiguration in broker operations, and it is invisible until the day a leader is lost. The operator has done real work: a floor on caught-up copies is set on every stream, copies are spread across failure domains, alerting exists. Then a machine dies, a surviving copy takes over, and a window of acknowledged records is missing. Nothing was broken. Everything did what it was configured to do. ## Why the floor did nothing The floor is **a condition on a rule, not a rule of its own**. It says: if a write is waiting for every caught-up copy, refuse it unless at least N copies are currently caught up. Consider a write whose acknowledgement rule is *the leader alone*. The server takes the record, the leader stores it, the writer is answered. At no point does the number of caught-up copies enter the decision — the write never asked for a second holder, so there is nothing for the floor to refuse. The floor sat there, correctly configured, never consulted. The halves of the pair live in different places, and that is exactly why they drift apart: - **The acknowledgement rule** is usually requested by the *writing application*, which is owned by a product team and deployed on its own schedule. - **The floor** is a *server-side* setting owned by the platform team. - The platform team can tighten the floor without touching a single writer, and it will show up as configured on every stream while changing nothing at all for writers on the low rungs. ## Then the leader is lost When a leader fails and a surviving copy takes over, that new leader can only serve what it holds. Records accepted by the old leader and never fetched by anyone else have exactly one holder, and that holder is gone. The writer was told they were accepted; nobody lied — "accepted" meant "one machine has it", which is precisely what was requested. The gap is bounded by how far behind the followers were, which is why the loss looks like a *window* of recent records rather than random gaps. Teams routinely misread this as a replication bug. ## Diagnosing it before it costs you 1. **Read the pair together, per stream.** The server-side floor tells you nothing on its own. The question is always "what does the floor say *and* what are writers actually asking for on this stream?" 2. **Check from the writing side.** The authoritative answer lives with the applications, not in the cluster's configured values. A stream can show a healthy floor while every writer on it asks for the cheapest rung. 3. **Watch write latency as evidence.** Writes that genuinely wait for every caught-up copy have a latency floor set by the slowest participating copy and by the distance between failure domains. Write latency that looks like a single local hop on a stream you believe is strongly acknowledged is a strong hint that the upper rung is not being requested. 4. **Test it deliberately.** Take one copy out of the caught-up set on a test stream. If writes keep flowing with no refusal, the floor is not being exercised — which is the whole finding. ## The asymmetry worth remembering | Situation | What the floor does | |---|---| | Write waits for every caught-up copy, enough copies current | Nothing; write proceeds | | Write waits for every caught-up copy, too few current | Refuses the write — the floor's whole purpose | | Write waits for the leader alone | Nothing, ever | | Write waits for nothing at all | Nothing, ever | Two of the four rows are the trap. A server-side setting that is inert for two of the four rungs is easy to mistake for a guarantee, because it is present, non-default and audited. ## How this differs across designs On designs that replicate by **majority write**, the majority is intrinsic to the write path, so there is no weak rung for a writer to select and this particular mismatch does not arise. Where a broker keeps **a single mirrored copy**, the equivalent question is whether the writer asked to be answered by the primary or by the mirrored pair — the same split, with two rungs instead of four. Where durability comes from **shared durable storage underneath the brokers**, the wait is on the underlying store and the operator has no per-stream floor to get wrong. ## The repair Make the two halves agree and own them together: decide per stream what the data requires, set the floor *and* the writer's rule to match, and re-verify from the writing side after any deployment. A durability posture that only one team can see is a posture that will be half-applied.

  • Why does the lost data appear as a recent window rather than as scattered gaps?
    Because the records that existed only on the failed leader are exactly those the followers had not yet fetched, which are the most recent ones. The size of the window tracks how far behind the followers were at the moment of failure, so it reads as a clean tail of missing records rather than as random holes in the stream.
  • Can a server force writers onto a stronger rung?
    Servers generally enforce floors rather than upgrade requests, so a write asking for a weak rung is usually answered at that rung rather than silently promoted. Treat the writer's request as the binding half and audit it directly; relying on the server to quietly strengthen what applications ask for is how this failure survives an audit.
  • What is the cheapest way to prove the floor is live on a stream?
    Exercise it on a test stream: remove one copy from the caught-up set and see whether writes are refused. A refusal proves the pair is wired together; writes continuing normally proves only that the writers are on a rung the floor never inspects, which is the finding you were looking for.

saying these in an interview costs you the question

  • Treats the floor as a guarantee independent of what the writer asks for
  • Audits server settings only and never checks the writing side
  • Blames replication for a window of records lost with the leader
  • Believes the server upgrades a weak request to a stronger rung
  • Assumes copies spread across failure domains fixes single-holder exposure
  • Cannot explain why the loss appears as a recent tail of records