skip to content

A volatile tier acknowledges a write before any replica holds a copy of it: what has the caller been promised?

level: juniorimportance: must knowfreq 66%

answer

  1. one node said yes
  2. the copy is behind, always
  3. acknowledged, but not yet propagated
  4. no other home for that state
  5. gone, not a replayable gap

basics

~10 s

Acknowledgment means one node has the write in memory, nothing more. Until it propagates, the write exists in exactly one place, and if that node is lost the write is simply gone.

solid answer

~50 s

Under acknowledge-then-propagate the primary applies the write, answers the caller, and only then sends it on to each replica. So the promise is that one node accepted it, not that a copy exists. The acknowledged writes no copy holds yet are the **un-propagated window**, and on this tier they are not a replayable gap in a durable record: memory is the only medium, so losing the primary vaporises them outright. The window is normally well under a millisecond on a healthy local network, but it is a moving quantity rather than a constant - `propagation lag` widens under a write burst, a saturated link or a busy copy. This ordering is the common case, not a law: some stores in this class acknowledge only after at least one copy holds the write, and some let the caller choose per write.

go deeper

for a junior

Recall the one sentence that matters: being told a write succeeded means one node has it in memory, not that a copy does. The replica is always slightly behind the primary, and that lag is where writes can disappear.

for a middle

Explain the ordering mechanically - apply, answer the caller, then send on - and name the un-propagated window as its consequence. Be able to say why the same ordering is far less serious on an engine that wrote the change to disk before answering.

for a senior

Show that you have watched the window move. Name what widens it under real conditions, and make the answer turn on what the affected entries actually hold rather than on the replication mechanism alone.

for a principal

Frame it as an exposure the organisation chooses rather than a defect. The question worth asking is how much recent state each workload can afford to vaporise, and whether that number has ever been written down and tested.

## The four moments of a write A volatile tier that keeps a full copy of its keyspace on a second machine has two nodes involved in every write. The **primary** accepts writes. The **replica** holds a copy of the whole keyspace and applies the stream of changes the primary sends it. Every write therefore passes through four distinct moments: 1. The primary receives the operation and applies it to its own memory. 2. The primary answers the caller - the client library returns and the application moves on. 3. The primary sends the change on to each replica. 4. Each replica applies the change to its own memory. The entire subject is the *order* of moment 2 against moment 4. Under **acknowledge-then-propagate**, moment 2 comes first: the caller is told yes while exactly one machine in the world holds the write. The interval between the caller being told yes and a copy actually holding the write is **the un-propagated window**, and its contents are acknowledged writes that no copy holds. ## Why this window is sharper here than on a durable engine A durable storage engine has the same ordering choice, and the same window - but there the un-propagated change is a **replayable gap**. The engine wrote the change into its own record on disk before answering. If the machine is repaired and brought back, the change is still there and can be shipped to the copy or replayed. The gap is a delay in a record that exists. On a tier whose medium is memory there is no second record. Nothing wrote the change anywhere but the primary's address space. When the primary is lost, the write was never anywhere else, and no operator action recovers it. It is not delayed, not queued, not pending - it is gone, and nothing in the system records that it ever happened. A caller that was told yes believes a thing that is no longer true anywhere. That asymmetry is the whole reason this question is asked. It is also why the answer depends on **what the state was**. If the entries were values reconstructable from a system of record, losing a few seconds of them costs some recomputation and a slow period. If they were claims, leases, deduplication records or quota counters with no other home, the same mechanism produces a double-charged customer, two holders of one lock, or a request processed twice. ## How large the window is, and why it is not a constant On a healthy link between machines in the same facility, propagation lag is typically well under a millisecond, so the window holds whatever arrived in that time. But the number people quote is a median in good conditions, and the number that matters is the one at the worst instant. The window in writes is roughly the lag at that instant multiplied by the arrival rate, and both terms move: - a **write burst** pushes arrivals above what propagation can carry, and the backlog grows until the burst ends; - a **saturated or shared network link** between the nodes stretches every hop; - a **busy copy** - under its own load, paused, or doing its own housekeeping - applies changes more slowly than they arrive; - a **long-running single operation on the primary** occupies it while writes queue behind it; - a copy that has briefly disconnected must catch up from a **bounded buffer the primary holds**, and if that buffer overflows the copy takes a whole fresh copy of the keyspace instead, during which it holds nothing recent at all. Adding more replicas does not shrink the window. Each copy has its own lag, and a write is un-propagated until the *fastest* copy has it; more copies add propagation load to the primary rather than removing it. ## What varies between stores in this class | Question | How stores differ | |---|---| | When is the caller told yes? | Most acknowledge before any copy holds the write. Some acknowledge only after at least one copy does. Some let the caller choose per write. | | Is the posture fixed? | Some fix it for the deployment; others expose it per call, so "we replicate asynchronously" does not settle what a given write actually did. | | How is lag reported? | Some report time behind the primary, some report outstanding un-propagated work, and some report from the primary's side rather than the copy's. | | Is there a write gate? | Some refuse writes when too few copies are current; others offer no such mechanism at all. | ## What this question is not about A promotion is the **event that exposes** the window, not the thing that creates it - how a node is declared gone and how a copy is promoted is a separate subject. Whether a read served from a copy is behind is likewise separate: that is about readers, this is about a writer who was told yes. And whether anything reached disk is a different posture measured against a different thing; a tier can be configured with neither disk persistence nor a waiting acknowledgment, and many are. ## The un-propagated window is the span between the caller being told yes and a copy holding the write. A primary lost inside it takes acknowledged writes with it; a primary lost after it does not ``` t0 caller -> primary : write accepted, caller told yes t0 primary : entry now held on one node t0+0.4ms primary -> replica: change sent t0+0.7ms replica : entry now held on two nodes primary lost at t0+0.2ms (inside the window) the write was acknowledged no copy holds it it is gone; nothing anywhere records that it happened primary lost at t0+0.9ms (after the window) a copy holds the write; it survives a promotion ```

  • Does running three replicas instead of one make the un-propagated window smaller?
    No. A write is un-propagated until the fastest copy holds it, so the window is governed by the best copy, not the count. Extra copies each carry their own lag and add propagation work to the primary, which under load can make the fastest copy slightly slower rather than faster.
  • Does an acknowledged write survive a restart of the primary?
    That is a different posture. Acknowledgment concerns whether a copy holds the write; surviving a restart concerns whether anything outside memory recorded it. The two are measured against different things, and a tier can be running with neither, which is a common and often deliberate configuration.
  • If the lost primary is repaired and rejoins, do its un-propagated writes come back?
    Not as accepted writes. The state was in memory and died with the process. If a copy was promoted meanwhile, what happens to a returning node that holds writes nobody else has is the split-brain question, not this one - but nothing re-injects them into the new primary.

A courier takes the parcel at your door and hands you a receipt on the spot. The receipt proves the courier has it; it says nothing about the depot. If the van is lost between your door and the depot, no record anywhere says the parcel existed - the receipt in your hand is now the only trace, and it refers to nothing.

saying these in an interview costs you the question

  • Says an acknowledged write is safe because a replica exists.
  • Treats propagation lag as a fixed constant of the deployment.
  • Assumes every store in this class acknowledges before propagating.
  • Calls the un-propagated writes a gap that gets replayed later.
  • Thinks adding more replicas shortens the un-propagated window.
  • Answers about stale reads instead of about a writer told yes.