skip to content

A tier stops acknowledging a write until at least one replica holds it: what does that buy, and what does it cost?

level: middleimportance: must knowfreq 58%

answer

  1. one hop per write, always
  2. in two memories, not one
  3. narrows the window, never closes it
  4. posture may be per write
  5. still nothing to do with disk

basics

~20 s

Waiting for a copy narrows the un-propagated window and adds a network hop to every write. It does not close the window: a write only one copy holds is still lost if that copy is not the one promoted.

solid answer

~40 s

Under **wait-for-a-copy acknowledgment** the primary applies the write, sends it on, and does not answer the caller until at least one copy confirms it holds it. What you buy is that an acknowledged write is, at the moment of acknowledgment, in two places rather than one - so a promotion does not automatically vaporise it. What you pay is a per-write latency bill: every write now carries a round trip to the acknowledging copy, which on a tier chosen for sub-millisecond response can be most of the response time. It narrows the window without closing it, and stores differ in whether the posture is fixed for the deployment or chosen per write - so "we replicate asynchronously" does not tell you what a given write actually did.

go deeper

for a junior

Know that there are two moments a store can pick between: answer the caller first, or wait until a copy has the write. Waiting is slower on every write and loses less when a node dies.

for a middle

Explain the round trip as the price and name it as a per-write cost, not a failure-path cost. Then state the limit yourself: one copy holding it is not the same as the window being closed.

for a senior

Show the choice being made per workload rather than per cluster, and know what your store does when no copy can confirm - block, refuse, or answer anyway - because that behaviour decides whether the posture is a latency change or an availability change.

for a principal

The decision is which writes justify the hop and who gets to declare that. Treat a uniform fleet-wide posture as a smell in both directions: one pays for safety nobody asked for, the other assumes nothing on the tier matters.

## The two postures, stated precisely **Acknowledge-then-propagate**: the primary applies the write, answers the caller, and sends the change on afterwards. The caller's latency is one round trip to the primary. The acknowledged writes no copy holds yet are the un-propagated window. **Wait-for-a-copy acknowledgment**: the primary applies the write, sends the change on, and holds the caller's answer until at least one copy confirms it holds the change. The caller's latency is now a round trip to the primary *plus* the primary's round trip to the fastest acknowledging copy. The second posture is what people usually mean when they reach for the word "synchronous", but on this tier it is better described as a timing than as a setting, because what is being promised is narrow: a copy holds it *in memory*, which on a volatile tier is all any copy ever has. ## What waiting actually costs The bill is paid on **every write**, not on the rare failure: - **Latency floor.** Write response time is now bounded below by the network distance to the acknowledging copy. A tier that answered in a few hundred microseconds may now answer in a millisecond or more. For a workload chosen precisely because this tier is fast, that can erase the reason the tier exists. - **Tail sensitivity.** The write path now inherits the copy's bad moments: its own load, a pause, a slow link. A copy that hiccups for 50ms makes every write in that window slow, not just one. - **In-flight work on the primary.** The primary is holding more outstanding writes at once, which costs it memory and connection capacity. - **Behaviour when no copy can answer.** This is where stores in this class genuinely diverge. Some block the write until a timeout expires and then answer anyway; some refuse the write; some degrade silently to answering without a copy. You cannot reason about the posture without knowing which of the three your store does, and it is worth confirming rather than assuming. ## What it does not fix This is the half candidates skip, and it is the half interviewers are listening for. 1. **A write held by exactly one copy is still lost if that copy is not the one promoted.** Waiting for *one* copy means one copy holds it; if a different copy is promoted, the write is not there. 2. **The window narrows, it does not become zero.** There is still an interval - the primary has applied the change and not yet heard back. A failure inside that interval leaves a write the caller was never told about, which is a smaller problem but not no problem. 3. **It says nothing about disk.** Whether anything survives a restart of the whole tier is a different posture, measured against a flush rather than against a copy. 4. **It does not make the tier a system of record.** Two copies of volatile state are still volatile state. ## Per deployment or per write Some stores fix the posture for the whole deployment. Others let the caller ask, on a given write, that the answer wait until some number of copies hold it. Where that per-write choice exists it is usually the right structure, because it lets you pay the hop exactly where the state justifies it: | What the affected entries hold | Sensible posture | |---|---| | Values reconstructable from a system of record | Acknowledge-then-propagate; losing a few costs recomputation | | Advisory counters and rate figures | Acknowledge-then-propagate; the exact value was never authoritative | | Claims, leases and deduplication records with no other home | Wait for a copy on those writes specifically | | Everything, uniformly | Rarely right - it pays the hop on writes that did not need it | The consequence for interview answers is that "we use asynchronous replication" is not a complete statement of the posture. The honest form is: this deployment acknowledges before propagating, *except* these writes, which wait for a copy - or, on a store with no per-write control, this deployment acknowledges before propagating and here is the window we accept for the state we keep here. ## The shape of a good answer A strong answer names the round trip as the price, names the promotion as the event the posture is defending against, and then volunteers the limit without being prompted: the write is in two memories rather than one, and two memories are still two memories. A weaker answer treats the waiting posture as a safety setting that makes writes durable, which is the misconception that sends teams to production believing a volatile tier has become a record of anything.

  • Why is waiting for a copy not the same as making the write durable?
    Durability is measured against something outside memory. Waiting for a copy puts the write in a second process's address space, which survives one node failing and does not survive the tier being restarted. They defend against different failures and neither substitutes for the other.
  • If waiting for one copy leaves a gap, why not wait for all of them?
    Because the write then inherits the slowest copy's worst moment, and a single unhealthy copy stalls every write. Waiting for more copies is a real dial, but it trades availability of the write path for a narrower window, and on a volatile tier that trade is rarely worth the far end.
  • What should you check before assuming this posture is available?
    Whether the store offers it at all, whether it is fixed for the deployment or selectable per write, and what it does when no copy can confirm - block until a timeout, refuse, or answer anyway. Stores in this class differ on all three.

saying these in an interview costs you the question

  • Says waiting for a copy makes the write durable.
  • Assumes the acknowledgment posture is always fixed for the whole deployment.
  • Ignores that the latency cost is paid on every write, not on failures.
  • Claims waiting for one copy closes the loss window entirely.
  • Cannot say what the store does when no copy can confirm.