Your order database has a synchronous standby in a second zone and an asynchronous copy in a second region — what differs when you promote each?
answer
- acknowledgement timing, not copy speed
- distance decides which is affordable
- lag equals the data you lose
- promotion breaks every open connection
- fence the old primary before promoting
basics
~20 sPromoting the in-zone synchronous standby loses no acknowledged writes, because every commit already waited for it. Promoting the cross-region asynchronous copy loses whatever the lag held — seconds or minutes of confirmed orders that exist nowhere else.
solid answer
~50 sSynchronous replication means a write is not acknowledged until the second copy has it. Zones in one region are close enough for that round trip to cost a small amount of latency on every commit, so a promotion there is clean: the standby is already current and no confirmed order is lost. A second region is too far for that — the round trip would be added to every write — so the cross-region copy is fed asynchronously and sits behind by a lag that grows under write bursts. Promoting it accepts that lag as data loss: orders the customer was told had succeeded simply are not there. Both promotions also break every open connection, so clients reconnect and in-flight work fails, and the old primary must be stopped from accepting writes before the new one starts.
go deeper
Remember the difference in one line: with synchronous replication the caller waits until both copies have the write, with asynchronous it does not. That single sentence explains why only one of them can lose confirmed data.
Explain why distance forces the choice, and state the consequence of each promotion: a latency tax on every write against a data-loss window equal to the lag at the moment you promoted.
Cover what promotion does beyond the data — dropped connections, repointing clients, fencing the old primary, retries that must be safe to repeat — and say how you measure lag rather than assuming it.
Turn it into a signed number: how much confirmed data the business will accept losing in a regional event, and what you would pay to make that number smaller. Say who owns that figure.
## Two copies making two different promises A replica is defined by **when the write is acknowledged relative to when the copy has it**. - **Synchronous**: the primary does not confirm the write to the caller until the second copy has durably accepted it. The cost is paid on every single write, forever. The benefit is paid once, at promotion: there is nothing to lose. - **Asynchronous**: the primary confirms immediately and ships the change afterwards. The cost is paid once, at promotion — everything not yet shipped is gone. The benefit is paid on every write: no waiting. Distance is what forces the choice. Zones inside one region are separated for failure isolation but sited close enough that a round trip is small, which is precisely what makes synchronous replication practical there. A second region is far away by design — that is the point of it — so the round trip would be added to every commit, and virtually nobody accepts that. Hence the shape almost every stack ends up with: **synchronous inside the region, asynchronous across regions**. | | Synchronous standby, second zone | Asynchronous copy, second region | |---|---|---| | Cost per write | a small latency addition on every commit | none | | Data lost at promotion | none of the acknowledged writes | whatever the lag held | | Protects against | losing one zone | losing the whole region | | Typical promotion time | short, often automated | longer, usually declared by a person | | Failure while healthy | a sick standby can slow or stall writes | lag grows quietly and nobody notices | ## The lag is not a constant The most useful thing to say about asynchronous replication is that the lag is a **variable you must measure, not a property you can assume**. It grows when the write rate rises, when a bulk job runs, when the link between regions is congested, and when the receiving side is busy applying what it already has. That means the lag is usually largest exactly when you are most likely to need the copy — during peak trading, or during the incident that is about to become a regional one. A plan that quotes a comfortable normal lag and never alarms on it is a plan built on the best case. ## What promotion does to everything else Promotion is not only a data event. Four things happen around it: 1. **Every open connection breaks.** Connection pools reconnect, and whatever was in flight fails. Callers need retries, and the operations behind them need to be safe to repeat, because some of those failed requests did commit before the connection dropped. 2. **Clients have to find the new writable copy.** Whether that is a name that repoints, a proxy that moves, or configuration that changes, that step has its own delay and its own cached copies sitting in callers. 3. **The old primary must be fenced.** If it comes back and still believes it is writable while the promoted copy is also taking writes, you have two divergent truths and a painful reconciliation. Stopping the old one accepting writes is part of promoting, not a follow-up task. 4. **Failing back is a second migration.** The promoted copy has accepted writes the old one never saw, so returning is a planned cutover with its own lag and its own window, not an undo. ## A replica is not a backup Both of these copies replicate faithfully — which means they replicate the mistake too. A wrong bulk update or a delete of the wrong rows is applied to the standby and to the far copy as diligently as any other change, usually within seconds. Replication protects against **losing a failure domain**; a backup with a retention window protects against **something being wrong in the data**. They are different controls for different risks, and a design that has only one of them is missing half its coverage. ## How to answer this out loud State the definition in terms of acknowledgement, not in terms of speed — 'the caller waits' against 'the caller does not wait'. Then give each promotion its consequence: no data loss and a latency tax, against no latency tax and a data-loss window equal to the lag. Then add the parts most candidates skip: broken connections, repointing clients, fencing the old primary, and the fact that neither copy is a backup. If you are asked which you would choose, the answer is usually both, for different failures, and the interesting question is what lag you are willing to sign for.
- Why not replicate synchronously to the second region as well?Because every commit would then wait for a round trip across that distance, which is an order of magnitude longer than a hop between zones in one region. Write latency would rise on every request for a failure that happens very rarely. Some systems accept it for a small, high-value subset of writes, but applying it to the whole order path is usually indefensible.
- What can go wrong with a synchronous standby while everything is healthy?It becomes a dependency of your write path. If the standby slows or stops acknowledging, commits slow or stall with it, so a partial failure in the second zone can degrade the primary. Systems usually offer a way to drop out of synchronous mode automatically — which restores write availability and quietly turns your zero-loss promise into a lagging one until it is restored.
- How do you know what the lag was at the moment you promoted?By measuring it continuously and alarming on it, not by asking afterwards. Record the last position the far copy acknowledged and compare it with the primary's, and keep the series so the value at the moment of the incident is recoverable. Without that, you cannot say what was lost and you cannot tell customers what to re-do.
A courier who waits for the signature before telling you the parcel arrived, against one who tells you immediately and posts the confirmation later. The second is faster for you, and leaves a window in which you believe something the far end has no record of.
saying these in an interview costs you the question
- Describing asynchronous replication as merely slower rather than lossy
- Assuming replication lag is a stable number you can quote
- Treating a replica as a backup against a mistaken delete
- Forgetting that promotion drops every open connection
- Promoting without stopping the old primary from accepting writes
- Believing a failback is simply the failover run in reverse