Before a broker answers a write, what can the writer be made to wait for, and what does each option cost in write latency?
answer
- a dial, not a guarantee
- four rungs of waiting
- each rung, one more holder
- top rung pays the slowest copy
- chosen per stream
basics
~20 sA write can be answered with no wait at all, once the leader holds it, once a majority of copies hold it, or once every caught-up copy holds it. Each rung up adds a network round trip of latency and removes one way to lose the record.
solid answer
~50 sThe acknowledgement rule is the server-side requirement met before a writer is told its record was accepted. The common ladder has four rungs: no wait (fire-and-forget), the leader alone, a majority of copies, and every caught-up copy. Fire-and-forget is the fastest and loses records to any failure or even a dropped connection, because nothing confirmed anything. The leader alone costs one hop and survives nothing worse than losing that leader. A majority or every caught-up copy costs the round trip to the slowest copy that has to answer, and in exchange the record survives losing the leader. The rule is a durability-versus-latency dial, and it is chosen per stream, not once for a whole cluster. Note that accepting a record is not the same as forcing its bytes to persistent media — that is a separate step.
go deeper
Be able to name the rungs in order and say, for each, how many places hold the record when the writer is told it was accepted. That alone answers the screening version of this question.
Explain where the latency of each rung comes from: a hop for the lower rungs, the slowest participating copy for the top one. Add that the rule is set per stream on top of a cluster default.
Show you have chosen this dial under pressure — which streams you put on which rung, and how you noticed a sick machine through producer latency rather than through an availability alarm.
Frame it as a cost envelope for the estate: each rung has a latency profile and an exposure profile, and the organisation needs a small number of named combinations rather than one argument per stream.
## What the answer to a write actually means A writer sends a record and the broker eventually answers. **The acknowledgement rule** is the server-side requirement that must be met before that answer goes back. It is the knob that decides how much of the cluster has seen a record at the moment the writer is told "accepted", and therefore how much has to fail before the record is gone. Two things it is *not*. It is not a statement that the bytes have been forced onto persistent media — acceptance into a broker process and a forced write to durable storage are separate steps with separate settings. And it is not about what a consumer later does with the record; a reader acknowledging work it has finished is a different mechanism entirely, on the other side of the system. ## The rungs of the ladder Four rungs is the common shape; some designs expose fewer. - **No wait (fire-and-forget).** The writer sends and carries on. There is no answer to wait for, so nothing can be retried intelligently: a record lost in the network, at a full buffer, or on a broker that was restarting simply never existed. Used for high-volume telemetry where a missing sample costs nothing. - **The leader alone.** On designs where one copy leads a stream and the others follow, the leader writes the record and answers immediately. One network hop of latency. The record exists in exactly one place; losing that one machine before a follower has fetched it loses the record. - **A majority of copies.** The write is answered once more than half the copies hold it. Latency is set by the *median* copy, so one slow machine does not stall writes. This is the natural rule on designs that replicate by majority write rather than by a leader-plus-follower arrangement. - **Every caught-up copy.** The write is answered once every copy the leader currently counts as current holds it. Latency is set by the **slowest** copy that has to answer, which is why this rung is the one that exposes a single sick disk as a producer-side latency spike. ## What each rung buys and costs | What the writer waits for | Holds the record when the answer arrives | Latency driver | Lost by | |---|---|---|---| | No wait | Unknown — possibly nothing | None | Any failure, silently | | The leader alone | One copy | One hop to the leader | Losing that leader before a follower copies it | | A majority of copies | More than half | The median copy | Losing more than half at once | | Every caught-up copy | Every current copy | The slowest current copy | Losing every current copy at once | The important shape of that table: latency rises roughly one round trip from the first rung to the second, and then as a function of *tail* behaviour, not average behaviour, from the third rung up. A cluster whose copies are spread across failure domains pays the inter-domain round trip on the upper rungs. ## The choice is per stream A cluster carries a default, and individual streams take a **per-stream override** on top of it. That matters because durability requirements are a property of the data, not of the machines: a payment event and a page-view event can live on the same cluster with different rules. An estate that sets one rule cluster-wide is either paying tail latency on data that does not need it, or running valuable data at the exposure level chosen for telemetry. ## Designs that do not offer all four rungs This is where a candidate who has only operated one platform gets caught. 1. Where a broker keeps **a single mirrored copy** rather than a configurable number, the choice collapses to two states: answered by the primary, or answered once the mirrored copy also holds it. There is no majority to wait for. 2. Where durability comes from **shared durable storage underneath the brokers** rather than from broker-held copies, the wait is for that underlying store to confirm, and copy count is not the operator's dial at all. 3. Where records are **deleted on acknowledgement** rather than retained in a log, the same ladder still applies to the write, even though there is no stored reading position anywhere in the system. ## What an interviewer is listening for That you state the rule as a trade, name what physically holds the record at each rung, and know where latency actually comes from — a hop for the low rungs, the slowest participating copy for the top one. Reciting rung names without the failure each one still permits is the shallow answer.
- Why does waiting for every caught-up copy make write latency sensitive to a single slow disk?Because the answer cannot go back until the last required copy has the record, the rule takes the maximum over the participating copies rather than the median. One machine with a degraded disk therefore sets the latency for every write on that stream, which is why this rung surfaces hardware trouble as a producer-side latency spike long before anything is unavailable.
- Does a write that has met the acknowledgement rule mean the record survives a power cut?Not by itself. Meeting the rule means the required copies accepted the record into their own storage path; whether those bytes have been forced out of volatile memory onto persistent media is a separate setting. The two questions are deliberately distinct, and an operator who treats them as one is describing a stronger guarantee than the cluster gives.
- Where is this rule chosen — by the writer or by the server?Usually both have a say. The writer asks for a level of waiting, and the server carries defaults plus its own floors that can make a request more demanding but not less. That split matters: an operator who tightens only the server side may find writers still asking for the weakest rung and getting it.
Posting a parcel: drop it in the box and walk away, get a receipt at the counter, get a receipt once it reaches the depot, or wait until every branch on the route has scanned it. Each extra scan takes longer and removes one way to lose it.
saying these in an interview costs you the question
- Thinks an acknowledged write is already on persistent media
- Believes fire-and-forget still retries a failed send
- Assumes the strongest rule is free because copies answer in parallel
- Treats the rule as one cluster-wide setting with no per-stream override
- Confuses this server-side wait with a consumer acknowledging finished work
- Cannot say what physically holds the record at each rung