skip to content

How do the replication commit mode (asynchronous versus synchronous) and the frequency with which write-ahead log segments are shipped off the primary determine the recovery point objective you can honestly promise? What do you pay for a tighter one?

level: middleimportance: must knowfreq 55%

answer

  1. recovery point = freshness of the surviving copy
  2. sync = ~0 loss, pay a round trip per commit
  3. async = lag at failure, unbounded tail
  4. archive timeout bounds the archiving path
  5. promise the alarmed ceiling, not the median

basics

~20 s

Your recovery point equals the newest data that survived the primary. Synchronous commit to another node gives near-zero loss; asynchronous gives loss equal to replication lag; log archiving gives loss up to the archive interval. Tighter costs write latency and couples your availability to the standby.

solid answer

~1 min

The recovery point you can promise is exactly **the freshness of the newest copy that survives the failure**, so it is determined by whichever off-primary durability mechanism is fastest. - **Synchronous replication** - the primary does not acknowledge a commit until a second node has the record durably. Recovery point is effectively zero for failures the standby survives. The price is that every commit pays a network round trip, and if the standby is unreachable you must choose between blocking writes (availability loss) and silently degrading to asynchronous (a recovery-point promise you no longer keep). - **Asynchronous replication** - the recovery point equals replication lag at the failure instant: usually milliseconds, but it becomes minutes under a write burst, a long transaction, a slow standby disk, or a network stall. So the honest promise is not the typical lag, it is the lag ceiling you alert on. - **Log archiving** - archives complete segments plus a timeout-driven flush, so the recovery point is roughly the archive interval plus transfer time. Shortening the timeout tightens it at the cost of many small, mostly-empty segments. Honest practice: promise the p99 of the measured lag, not the median, and alarm on the ceiling.

code

text · 5 lines
text
sync standby (same region)   -> ~0 loss on host/zone failure   | +1 RTT per commit, blocks if standby down
sync standby (cross region)  -> ~0 loss on region failure      | +30-80 ms per commit
async standby                -> loss = byte lag at failure     | unbounded tail; alarm on the ceiling
WAL archiving, timeout 60s   -> loss <= 60s + transfer         | many small segments, longer replay
nightly backup only          -> loss <= 24h                    | cheapest

go deeper

for a junior

Know the mapping: synchronous means essentially no data loss, asynchronous means you lose whatever had not shipped yet, and shipping logs periodically means you lose up to that period.

for a middle

Explain the mechanisms and their prices - commit round trip, availability coupling, unbounded lag tails - and identify the archive timeout as the knob that bounds the archiving path.

for a senior

Insist on stating the recovery point per failure scenario and as a monitored ceiling, cover degraded-mode alerting, byte versus time lag measurement, and archiver backlog as a silent regression.

for a principal

Weigh the per-commit latency tax on all traffic against a rare event, decide which tiers deserve synchronous cross-region commit, and design quorum topologies that keep the guarantee without coupling availability to a single standby.

## The rule that generates every answer After you lose the primary, the best you can recover is **the most recent state that exists somewhere else and is durable**. Everything about recovery-point engineering is therefore a question of how fast, and how reliably, writes leave the primary. ## Synchronous commit In synchronous replication the primary withholds the commit acknowledgement until at least one standby confirms it has the write-ahead log record durably (or applied, depending on the level configured). If the primary is destroyed the instant after acknowledgement, the standby already holds that transaction. Recovery point for that failure class is zero. Costs and caveats: - **Latency.** Every commit pays at least one network round trip. Within an availability zone that is sub-millisecond; across regions it can be 30-80 ms, which for a chatty write path is often unacceptable. This is a per-commit tax on all traffic, forever, to protect against a rare event. - **Availability coupling.** If the standby is down or unreachable, a strict synchronous primary must block commits. Many teams configure a fallback to asynchronous so writes keep flowing - which is a reasonable availability choice but silently voids the zero-recovery-point promise for the duration. If you do that, you must alert on it and be honest that the guarantee is conditional. - **Quorum forms.** Requiring acknowledgement from any one of several standbys, or a majority, restores availability without dropping to asynchronous, at the cost of more replicas to run. - **Scope.** Synchronous to a standby in the same rack or zone protects against host loss, not against site or region loss. The guarantee is only as broad as the failure domain the standby sits outside of. ## Asynchronous replication The primary commits locally and streams to the standby in the background. Normal-case lag is small - often single-digit milliseconds - but it is **not bounded**. It grows with write bursts that exceed standby apply throughput, network congestion or a cross-region link hiccup, standby I/O saturation or maintenance such as a vacuum or index build, and long-running transactions that hold back applying. The crucial interview point: the recovery point you can promise is not the median lag, it is the **worst lag you tolerate before you consider the system out of compliance**. If you promise 5 seconds, you need monitoring on lag with an alarm at 5 seconds and a defined action when it fires. Otherwise you promised the happy path. Measure lag in **bytes and in time**: byte lag (how much log the standby has not received) is the direct predictor of data loss; time lag is easier for people to reason about but is misleading on an idle system, where the reported lag can grow simply because nothing new has been written. ## Log archiving cadence Continuous archiving copies completed write-ahead log segments to durable off-host storage. Two knobs govern the recovery point: - **Segment fill rate.** A segment is archived when it fills. On a busy system that is frequent; on a quiet one, a segment may take hours to fill, so recent writes sit only on the primary. - **Archive timeout.** A timeout forces the current segment to be closed and archived after N seconds regardless of fill. This bounds the recovery point at roughly N plus transfer time, and it is the setting that actually delivers the promise. The cost of a short timeout is many partially-used segments - more objects, more storage and request cost, more archive operations, and more segments to replay at restore time. Values in the tens of seconds to a few minutes are typical; single-digit seconds usually means you should have used streaming replication instead. Also remember that archiving must be **monitored for backlog**. If the archive command starts failing, the primary retains log locally and the recovery point silently degrades from minutes to "since the archiver broke", while every backup job still reports success. Failed archiving also fills the primary's disk, turning a recovery-point problem into an outage. ## Combining mechanisms Real designs layer them: a synchronous standby in a second zone for zero loss on host and zone failure, an asynchronous standby in another region for site loss with a small recovery-point exposure, and continuous archiving to object storage as the mechanism that also supports recovering from logical damage that replication would faithfully copy. Your promised recovery point is the best of the mechanisms that survives the failure you are describing - which is why a serious answer states the recovery point **per failure scenario**, not as a single number. ## What to say Name the mechanism-to-number mapping, insist that asynchronous promises must be stated as a monitored ceiling rather than a typical value, mention the archive timeout as the knob that bounds the archiving path, and be explicit about the price: write latency and availability coupling for synchronous, unbounded tail risk for asynchronous, storage and replay overhead for aggressive archiving.

  • A synchronous standby becomes unreachable. What are the options and what does each do to the recovery point promise?
    Either block commits until a standby acknowledges, preserving the zero-loss guarantee at the cost of a write outage, or fall back to asynchronous so writes continue while the guarantee lapses. A middle path is quorum acknowledgement across several standbys, so losing one does not force the choice. Whichever you pick, the degraded state must raise an alert, because the system is running with a recovery point it no longer promises.
  • Why can measured replication lag look large on a mostly idle database, and how should you measure it instead?
    Time-based lag is often derived from the timestamp of the last replayed transaction, so on an idle primary it simply grows with wall-clock time even though the standby is fully caught up. Measure byte lag - the difference between the write-ahead log position generated and the position received or flushed by the standby - which reflects genuine outstanding data, and use time lag only as a human-readable secondary signal.
  • How does shortening the write-ahead log archive timeout affect the system beyond the recovery point?
    It forces partially filled segments to be archived, so the number of archived objects rises sharply on low-traffic systems, increasing storage, request costs, and the number of segments to fetch and replay during a restore, which lengthens recovery time slightly. It also increases archive command invocations, so a marginal archiver becomes a bottleneck sooner. Values of tens of seconds to a few minutes are the usual balance.

saying these in an interview costs you the question

  • Quoting the typical replication lag as the recovery point objective while nothing alarms on the tail.
  • Claiming synchronous replication guarantees zero data loss without mentioning that a fallback-to-asynchronous setting silently voids it.
  • Believing a same-zone synchronous standby protects against regional loss.
  • Forgetting that a failed or backlogged log archiver degrades the recovery point invisibly while backup jobs still report success.
  • Assuming a shorter archive timeout is free, ignoring the segment, storage, and replay overhead it creates.

context