skip to content

A stream's acknowledgement rule is every caught-up copy and its minimum-copies floor equals its copy count; what happens when one copy drops out?

level: middleimportance: must knowfreq 62%

answer

  1. two rules, not one
  2. the floor refuses, the rule waits
  3. zero tolerance at count
  4. writes stop, reads continue
  5. conventionally one below the count

basics

~20 s

Writes to that stream are refused until a third copy is caught up again. Setting the floor equal to the copy count turns the loss of any one copy into a write outage, which is why the floor is normally set one below the count.

solid answer

~50 s

Two separate rules interact here. The acknowledgement rule says what the writer waits for; the minimum-copies floor says the server must refuse the write outright unless at least that many copies are currently caught up. With three copies, a floor of three and a rule of every caught-up copy, the moment one copy drops out — a reboot, a slow disk, a lost network path — the floor is unmet and every write to that stream is rejected. Records already accepted are still readable; it is new writes that stop. The conventional posture is a floor one below the copy count: three copies with a floor of two tolerates losing one copy while still guaranteeing that an accepted record sits in at least two places. The floor is not a standalone setting — it only has meaning against the copy count you chose.

go deeper

for a junior

Remember that two settings are in play: one says what the writer waits for, the other says how many copies must be current before the server will accept a write at all.

for a middle

Work the interaction out loud: floor equal to count means zero tolerance, one copy out means writes refused, reads keep serving, and writes resume when a copy catches up.

for a senior

Show the diagnostic shape — write refusals on one stream with a healthy-looking cluster — and be honest about the emergency lever of lowering the floor and what it costs while lowered.

for a principal

Treat the pairing as policy, not tuning: the organisation decides which streams may stop writing to protect durability, and that decision belongs in a standard rather than in each team's head.

## Two rules, one outcome Operators routinely collapse these into a single idea, and this question exists to separate them. - **The acknowledgement rule** is what the *writer waits for*: no wait, the leader alone, a majority of copies, or every caught-up copy. - **The minimum-copies floor** is a *server-side refusal*: unless at least N copies of this stream are currently caught up, do not accept the write at all — answer the writer with an error instead. The floor exists because the top rung of the acknowledgement ladder is self-weakening. "Every caught-up copy" means every copy the leader *currently* counts as current. If copies keep dropping out until only the leader remains, "every caught-up copy" is satisfied by the leader alone, and the strongest-sounding rule has quietly become the weakest. The floor is the backstop that says: below this many current copies, refuse rather than accept a write that looks strong and is not. ## What happens at floor equals count With three copies, a floor of three, and writes waiting for every caught-up copy, the stream has **zero tolerance**. One copy dropping out for any reason at all — a rolling reboot, a slow disk that falls behind, a lost network path between failure domains — leaves two caught-up copies, the floor is unmet, and the server refuses every write to that stream. Specifically: 1. **New writes fail.** The writer receives an error, not a delay, and whether that is retried, buffered or dropped is now the writing application's problem. 2. **Already-accepted records are unaffected.** They remain readable up to the committed edge. This is a write outage, not a data outage. 3. **Other streams are unaffected**, unless they share the same copy placement and lost a copy on the same machine. 4. **Recovery is automatic but not instant.** Writes resume when a third copy has caught up, and how long that takes depends on how far behind it fell and how fast catch-up traffic moves. The operator-visible signature is distinctive: a write-error spike on one stream, with the cluster otherwise healthy and reads fine. Teams that have not seen it before usually go looking for a network fault. ## Why the conventional posture is one below the count | Copies | Floor | Tolerates | An accepted record lives in | Failure mode | |---|---|---|---|---| | 3 | 3 | Nothing | 3 copies | One copy out stops writes | | 3 | 2 | One copy out | At least 2 copies | Two copies out stops writes | | 3 | 1 | Two copies out | Possibly 1 copy | Silent single-copy acceptance | | 2 | 2 | Nothing | 2 copies | One copy out stops writes | The middle row is the usual answer: it survives a single-copy failure while still guaranteeing two holders for every accepted record. The bottom row shows why the floor is not free to lower during an incident — dropping it to one restores writes immediately and silently returns you to single-copy exposure, which is the emergency lever people pull at 3 a.m. and forget to put back. Notice also the two-copy row. A copy count of two cannot give both fault tolerance and a meaningful floor: a floor of two has no tolerance, and a floor of one is no floor. That is the real argument for a third copy, and it is an argument about the floor, not about surviving disk failures. ## The same trap in other shapes On designs that replicate by **majority write** rather than a leader with followers, the majority *is* the floor — it is built into the rule, and an operator cannot raise it into self-inflicted refusal the same way. Where a broker keeps **a single mirrored copy**, the equivalent question is whether a write is refused or quietly accepted by the primary alone when the mirror is unavailable; different products answer that differently, and it is worth knowing which your cluster does. Where durability comes from **shared durable storage underneath the brokers**, there is no per-stream copy count to floor at all. ## Putting it in one sentence The floor converts a durability risk into an availability risk, deliberately. Setting it equal to the copy count means you have chosen to stop writing rather than to write with less redundancy than you promised — and if you did not intend that choice, the first copy to reboot will tell you.

  • Why does the top acknowledgement rung need a floor underneath it at all?
    Because "every caught-up copy" is measured against whichever copies are current at that moment. If copies drop out until only the leader is left, the rule is satisfied by the leader alone and has become the weakest rung while still reading as the strongest. The floor refuses the write instead of accepting a promise the cluster can no longer keep.
  • During an incident someone lowers the floor to restore writes. What has been traded?
    Write availability has been bought with durability margin. Records accepted while the floor is lowered may sit on fewer copies than the standard promises, and nothing marks them afterwards. It is a legitimate emergency lever, but it needs an owner, a time box and an explicit restoration step, or the temporary posture silently becomes the permanent one.
  • What does an operator see when the floor is unmet?
    Write errors concentrated on the affected streams, with reads still serving normally and most cluster indicators healthy. The tell is that the error is a refusal rather than a timeout, and that it appeared at the moment a copy left the caught-up set — often a planned restart rather than a failure.

saying these in an interview costs you the question

  • Sets the floor equal to the copy count, so one lost copy stops writes
  • Thinks an unmet floor makes already-accepted records unreadable
  • Treats the floor and the acknowledgement rule as the same setting
  • Lowers the floor during an incident and never restores it
  • Believes two copies plus a floor of two is a fault-tolerant posture
  • Expects a timeout rather than an outright refusal when the floor is unmet