skip to content

How do you decide whether a system genuinely needs zero-data-loss replication given its latency and availability costs, and what would you configure differently for a payments ledger versus a clickstream ingestion pipeline?

level: principalimportance: should knowfreq 32%

answer

  1. cost of loss, then reconstructability
  2. replayable upstream = no real RPO problem
  3. sync in-region, async cross-region
  4. quorum headroom or you chose an outage
  5. RPO is not RTO; idempotent writes for timeouts

basics

~20 s

Decide from the cost of losing the last second of writes and whether they can be reconstructed upstream. A ledger cannot reconstruct them, so it pays a same-region synchronous quorum plus asynchronous cross-region copies. Clickstream data is replayable from the producer, so asynchronous replication and relaxed local commit are the right trade.

solid answer

~50 s

Start with two questions per data class: what does losing the last second of committed writes cost, and can those writes be reconstructed from somewhere else (a producer's log, an upstream queue, the client)? Zero loss is worth paying for only when the answer is expensive and no. **Payments ledger**: losses are unrecoverable and auditable. Configure synchronous commit at the flush level with a quorum of any one of several standbys inside one region, so a zone loss is survivable without stalling writes; add asynchronous cross-region replicas for disaster recovery, accepting a small RPO for that rarer event; make write APIs idempotent so timeouts are safe to retry; alert loudly on any degrade to asynchronous. **Clickstream ingest**: the producer or broker can replay, volumes are huge, and per-event value is tiny. Asynchronous replication, batched commits, and relaxed local flush are correct; the money goes to throughput and to retention in the upstream log instead. The general rule: buy durability per data class, not per cluster.

go deeper

for a junior

Recognise that important data may need synchronous replication and that it costs latency; a concrete example of each side is enough.

for a middle

Compare the two workloads on data loss cost and latency, and pick asynchronous or synchronous with a stated reason.

for a senior

Give the full configuration for each shape, including quorum size, acknowledgement level, cross-region strategy, and idempotency for ambiguous timeouts.

for a principal

Drive the decision from business cost and reconstructability, own the availability commitment implied by any zero-loss promise, and set policy per data class with alerting on degraded mode.

## Frame the decision, do not default it The weak answer is important data gets synchronous replication. The strong answer prices the guarantee. Three inputs decide it. **Cost of loss.** Multiply the value of the writes in your worst-case exposure window by the probability of a sudden, unplanned primary loss. With asynchronous replication the exposure is roughly the current lag, often well under a second. Losing 300 ms of ledger entries means unreconcilable money and an audit finding. Losing 300 ms of page-view events means a rounding error in a dashboard. **Reconstructability.** This is the question most candidates skip. If the writes exist upstream, in a message broker with retention, in an idempotent producer that can replay, in a mobile client's outbox, then a small database RPO is not a real business RPO: you replay. Zero-loss replication is for data that is born in the database and exists nowhere else. **Latency budget.** Synchronous commit adds a round trip plus a remote flush to every commit. Within a region, across availability zones, that is typically well under two milliseconds and usually acceptable. Across regions it is tens of milliseconds and irreducible by tuning, because it is distance. If your write path already has a tight budget, a cross-region synchronous commit will consume most of it. ## The availability consequence you must own A zero-loss promise is a promise not to acknowledge when the second copy is unavailable. With a single standby that turns standby loss into a write outage. The engineering answer is headroom: require any one acknowledgement out of three standbys placed in different availability zones, so any single zone or node failure is invisible to writes, while every acknowledged transaction still lives on at least two machines. If you cannot afford that many nodes, you are implicitly choosing between an outage and a possible loss, and that choice belongs to the business, documented, not to a default in a configuration file. ## Payments ledger: a concrete shape - Synchronous acknowledgement at the flush level (the standby has the change on its own durable storage), because a correlated restart must not lose an acknowledged entry. - Quorum of any 1 of 3 standbys across zones within one region: zero loss for the common failure, no write stall for single-node or single-zone loss. - Asynchronous replica in a second region for disaster recovery. Paying cross-region latency on every commit is almost never justified; a regional disaster with a few seconds of RPO is a documented, accepted risk. - Idempotent write APIs keyed by a client-supplied request identifier, because a synchronous timeout leaves the client unable to tell whether the transaction happened. - Failover automation that compares replication positions before promoting, and fences the old primary. - Alerting on degraded mode, and a tracked metric for time spent running without the guarantee. - Backups and point-in-time recovery regardless, since replication does not protect against logical errors. ## Clickstream ingest: a different shape - Asynchronous replication, because the broker retains events and the pipeline can replay a lost window. - Relaxed local commit durability (acknowledging before the primary's own flush), trading a fraction of a second of crash exposure for a large throughput gain, since events are batched and individually low-value. - Larger batch sizes and fewer, bigger transactions; the per-commit round trip you avoided is the whole point. - Investment goes to upstream retention and replay tooling, which is the real durability mechanism here, rather than to database acknowledgement levels. ## Mixing within one database The levels are usually settable per transaction, so a single cluster can run ledger writes strictly and telemetry writes loosely. This is often better engineering than splitting into two clusters purely for durability, though it demands discipline: the strict setting must be applied in the code path, not assumed from a global default, and it must be covered by a test. ## What to say about numbers Be explicit that recovery point objective and recovery time objective are separate. Synchronous replication addresses RPO, the amount of data you may lose. Fast automated failover addresses RTO, how long you are down. A team can meet RPO of zero and still be down for twenty minutes, or fail over in fifteen seconds while losing two seconds of writes. Both come from stated business requirements, and they buy different things. ## How to present it Give the decision procedure first (cost of loss, reconstructability, latency budget, availability headroom), then two concrete configurations that differ in every relevant dimension, and end with the honest statement that a zero-loss promise is also an availability commitment which must be sized and alerted on.

  • Why is asynchronous replication usually the right choice for the cross-region copy even in a zero-loss system?
    Because commit latency would inherit the inter-region round trip, tens of milliseconds that no tuning removes, on every write for the entire life of the system, in exchange for protection against a rare regional disaster. The normal decision is a synchronous quorum within the region for the common failure modes plus an asynchronous cross-region replica with a small documented RPO for the disaster case.
  • A team says they have RPO zero because they replicate synchronously. What else would you check?
    Whether the acknowledgement level is flush rather than merely received; whether the configuration can silently degrade to asynchronous and whether that degrade is alerted; whether failover promotes the most advanced standby and fences the old primary; and whether write APIs are idempotent, since a synchronous timeout leaves the client unsure. Also whether backups exist, because replication does not cover logical errors.

saying these in an interview costs you the question

  • Applying zero-loss configuration everywhere because it sounds safer, without pricing latency and availability
  • Ignoring that upstream replayability can make database RPO irrelevant
  • Proposing synchronous commit across regions for an interactive write path
  • Treating RPO and RTO as the same requirement
  • Promising zero loss without idempotent write APIs, so ambiguous timeouts cause duplicate or lost work

context