skip to content

Your organisation runs volatile tiers on several products whose acknowledgment postures differ: what standing rule do you set, and what do you refuse to depend on?

level: principalimportance: nice to knowfreq 29%

answer

  1. standardise the number, not the mechanism
  2. the owning team owns the window
  3. a declared window must be tested
  4. refuse guarantees that vary by product
  5. change the state, not the setting

basics

~10 s

Standardise the number, not the mechanism: each tier declares the un-propagated window it accepts and proves it with a real promotion. Refuse to depend on guarantees the fleet's products implement differently.

solid answer

~50 s

You cannot standardise on a guarantee your products do not share, so a single fleet-wide acknowledgment posture is the wrong unit. What you can make uniform is the **question every tier must answer**: how many acknowledged writes would a promotion vaporise here, what did they hold, and when was that last tested. Then name the things you refuse to depend on because they vary - the existence of a per-write posture, the existence of a write gate, and what a lag number means and which side reports it. The cheapest structural move is usually not a replication setting at all: change what the state is, so that a lost window costs recomputation rather than correctness. And be explicit that no configuration of a volatile tier makes an acknowledged write recoverable - the two postures are a narrower and a wider window, not a safe end and an unsafe end.

go deeper

for a junior

Know that different in-memory products make different promises about when a write is acknowledged, so a practice copied from one deployment will not necessarily hold in another.

for a middle

Be able to explain why a single fleet-wide posture is hard: the control may not exist on every product, and the latency price falls unevenly across workloads with very different response-time budgets.

for a senior

Bring a tested number rather than a setting. Say what a promotion would vaporise on the tier you run, what the entries held, and when you last proved it by causing one under load.

for a principal

Own the framing: the decision is how much recent state each workload may vaporise, and the standard you set is that the number is declared, owned, tested and reviewed - not that a particular replication mode is switched on everywhere.

## Why a fleet-wide posture is the wrong unit The instinct is to write a rule: "all tiers wait for a copy before acknowledging". It fails on three counts. First, **the products differ**. Some expose a waiting acknowledgment, some do not; some allow it per write, some only per deployment; some hosted offerings fix the posture and expose no control at all. A rule the platform cannot enforce everywhere becomes a rule enforced in the places that needed it least. Second, **the cost is not uniform**. A tier answering in hundreds of microseconds and a tier answering in low milliseconds pay very different proportional prices for an extra hop. A blanket waiting posture quietly removes the reason some of these tiers were chosen. Third, **the risk is not uniform**. Most entries on most of these tiers are reconstructable from somewhere else. Paying a per-write hop to protect state that can be rebuilt is spending latency on nothing. ## What you can actually standardise 1. **A declared window per tier.** Every tier states the maximum un-propagated window it accepts, as a duration and as a count of writes at peak. The team that owns the state owns the number, not the platform team. 2. **A statement of what the window holds.** The number is meaningless without it. A tier that can honestly say "everything here is reconstructable" has finished; a tier holding claims, leases, deduplication records or quota counters has to justify the number against a concrete consequence. 3. **Evidence.** A declared window that has never been tested is a guess. Require that a promotion has been caused deliberately, under representative load, with the new primary's contents compared against what callers were told. 4. **A review trigger.** The declaration is re-examined when the write rate, the state on the tier, or the product underneath it changes - and a product migration is the most common way a tested number silently becomes false. ## What to refuse to depend on These vary across products in this class, so an architecture that assumes any of them is an architecture that cannot survive a change of store: - **That the posture can be chosen per write.** Some stores offer it; many do not. - **That a write gate exists.** The ability to refuse writes when copies fall behind is not universal, and where it exists the controls differ. - **That a lag number means the same thing.** Time behind, outstanding un-propagated work, reported by the primary or by the copy - four different numbers that all appear under similar labels. - **That a promoted copy holds everything an acknowledged write touched.** Even under a waiting posture, the copy that confirmed and the copy that is promoted need not be the same one. - **That anything about this reaches disk.** Disk behaviour is an independent posture, measured against a flush rather than against a copy, and it deserves its own declaration. ## The cheaper move: change the state, not the replication The highest-leverage decision available is usually upstream of any replication setting. If the state on the tier can be reconstructed - from a durable record, from the request that produced it, from a recomputation - then the whole subject collapses from a correctness question into a performance question. A promotion costs a slow period instead of an incident, and no write pays a hop. Where that is not possible, the second-best move is to shrink the set that cannot be rebuilt rather than to harden everything. A tier where only the claim records are irreplaceable is a tier where a per-write waiting posture, if the product offers one, is cheap because it applies to a minority of writes. ## What to measure, organisationally Replacing "is replication healthy" with a better question is most of the win. The better question is: *if the primary had failed at the worst instant of the last seven days, how many acknowledged writes would have ceased to exist, and what were they?* That framing does three useful things - it forces the tail rather than the average, it makes the answer a number someone owns, and it keeps the loss framed as what it is rather than as a delay. ## The admission to make out loud The two postures are a wider window and a narrower one. There is no configuration of a volatile tier that makes an acknowledged write recoverable, because recoverability requires a second record and this tier's medium does not keep one. Every distribution decision here is therefore a decision about **how much recent state the workload can afford to vaporise**, and an organisation that has never named that quantity for a given tier has not decided - it has defaulted.

  • A team proposes a fleet-wide rule that every tier waits for a copy. What is your objection?
    It cannot be enforced, because not every product in the fleet offers the posture, and where it can be enforced it charges a per-write hop to protect state that is mostly reconstructable. The unit of decision is the state on a tier, not the fleet.
  • What single question would you put on every tier's review?
    If the primary had failed at the worst instant of the last week, how many acknowledged writes would no longer exist anywhere, and what did they hold? It forces the tail rather than the average, makes the number someone's property, and keeps the loss framed as loss.
  • When does a previously tested window declaration silently become false?
    When the write rate grows, when new kinds of state are put on the tier without review, or when the product underneath changes - a migration between stores in this class can change the default posture, the lag metric's meaning and the availability of a write gate all at once.

saying these in an interview costs you the question

  • Sets one acknowledgment posture for the whole fleet regardless of product.
  • Assumes every store offers a per-write acknowledgment choice.
  • Treats a declared loss window that has never been tested as a fact.
  • Compares lag numbers across products as if they measured the same thing.
  • Believes some configuration makes an acknowledged write recoverable.
  • Hardens replication instead of making the state reconstructable.