skip to content

Your database primary's commits suddenly hang for every client, and the only change in the environment is that one standby became unreachable. Explain the mechanism, and how you would configure the system so that losing a standby cannot stall writes.

level: seniorimportance: must knowfreq 44%

answer

  1. writes hang, reads fine, CPU idle
  2. waiting on remote ack, already durable locally
  3. cancel = client uncertain about the outcome
  4. ANY 1 of 3 for headroom
  5. MySQL semi-sync timeout degrades; always alert

basics

~20 s

Synchronous commit makes the primary wait for a standby's acknowledgement before returning; with the only such standby gone, every commit waits forever. Fixes: use a quorum of any k of n with n greater than k, configure an automatic timeout that degrades to asynchronous, and monitor and alert on degraded mode.

solid answer

~50 s

The transactions are not stuck in the engine; they are already written locally and are waiting for a remote acknowledgement that will never come. That is by design: a synchronous configuration promises that acknowledged means replicated, so with no standby to replicate to, the honest behaviour is to not acknowledge. The subtlety is that the transaction is already durable locally. If you cancel the wait or the server restarts, that transaction becomes visible anyway, so clients can end up uncertain about the outcome. Configuration options. List more standbys than the quorum requires, for example any 1 of 3 across zones, so any single loss is invisible. Use a timeout that falls back to asynchronous, which MySQL semi-synchronous does natively; PostgreSQL has no automatic degrade, so a controller such as Patroni must rewrite the synchronous configuration. And decide deliberately: automatic degrade preserves availability but silently drops your zero-loss promise, so it must be alerted on.

go deeper

for a junior

Say that the primary is waiting for the standby's acknowledgement and that with the standby gone the wait never completes, so writes hang.

for a middle

Add how to spot it (writes hang, reads fine, idle server, sessions in a replication wait) and that more standbys than the quorum requires prevents it.

for a senior

Cover the locally-durable-but-unacknowledged subtlety, timeout-based degrade versus controller-driven reconfiguration, and alerting on degraded mode.

for a principal

State the policy question outright: whether an outage or a possible lost transaction is worse for this data, and encode that choice in the quorum size, the degrade rule, and the alerting.

## The mechanism In a synchronous configuration the commit path is: write the change records to the primary's log, flush them, send them to the standby, and wait for the acknowledgement before returning success to the client. If the standby is unreachable, the last step never completes. Every committing session parks in that wait, connections accumulate, the pool exhausts, and the application looks fully down even though the database is healthy and read queries still work. This is not a bug. You asked for the guarantee that an acknowledged transaction exists on more than one machine; with only one machine reachable, the system cannot honour it, so it declines to acknowledge. The failure mode is the direct price of the promise. ## The uncomfortable detail The waiting transaction is already committed locally: its records are durable in the primary's log and the locks are released at the appropriate point. If an operator cancels the backend, or the primary restarts, that transaction is visible to everyone afterwards even though the client never received a success. So a stalled synchronous cluster produces transactions in an indeterminate state from the client's point of view. This is exactly why write APIs should be idempotent and why unique request keys matter: on timeout the client cannot know whether the write happened. ## Diagnosis in the moment The signature is distinctive: writes hang while reads are fine, the primary's CPU and I/O are idle, and sessions sit in a wait state naming synchronous replication rather than a lock. Check which standbys are connected and what the synchronous configuration requires; if the required count exceeds the connected count, you have found it. ## Making it not happen **Configure a quorum with headroom.** The root mistake is requiring exactly the standbys you have. Requiring any one acknowledgement out of three listed standbys means two can vanish before writes block. This is the primary defence and it costs only extra replicas. **Use an automatic timeout to degrade.** MySQL semi-synchronous replication has a source timeout: if no replica acknowledges within it, the source reverts to asynchronous and continues, resuming semi-synchronous when a replica returns. This trades the zero-loss promise for availability automatically. PostgreSQL deliberately has no such setting: a synchronous standby is a promise it keeps until a human or a controller changes the configuration. Cluster managers implement the degrade by rewriting the synchronous standby list, and typically offer a strict mode that refuses to degrade for deployments where losing a transaction is worse than an outage. **Decide the policy explicitly, and alert on it.** The silent danger of automatic degrade is that you keep the availability and lose the guarantee without noticing; a ledger that has quietly been running asynchronous for three weeks fails its audit at the worst moment. Every degrade should page or at least raise a visible alert, and time spent degraded should be a tracked metric. **Do not cancel the waits as a routine fix.** Cancelling a synchronous wait releases the client with an error while the transaction is locally durable, which manufactures the ambiguity described above. It is an emergency lever, not an operating procedure. ## The underlying trade-off With a single standby you cannot have both a zero-loss promise and availability under standby loss: the moment the second copy is gone, either you stop acknowledging (lose availability) or you acknowledge with one copy (lose the promise). Adding standbys does not remove the trade-off, it makes the bad case rarer by requiring more simultaneous failures to reach it. That is the whole design conversation, and stating it plainly is what distinguishes a senior answer. ## How to present it Name the wait immediately, note the locally-durable-but-unacknowledged subtlety, then give the three configuration levers (quorum headroom, timeout-based degrade, explicit policy with alerting) and close with the trade-off statement: with one standby, zero loss and availability under standby loss cannot both hold.

  • If an operator cancels a query that is waiting on the synchronous acknowledgement, what is the state of that transaction?
    It is already committed locally and will be visible to other sessions, but the client received an error or nothing at all. The client therefore cannot tell whether its write took effect, which is why such APIs need idempotency keys and a way to query the outcome. It is an emergency action, not a routine remedy.
  • Why does PostgreSQL not offer an automatic timeout that falls back to asynchronous?
    Because the setting is a durability promise, and silently downgrading it would mean acknowledged transactions might exist on one node only without anyone knowing. PostgreSQL leaves the decision to an external controller so the degrade is an explicit, observable operational act, and many controllers offer a strict mode that refuses to degrade at all.

saying these in an interview costs you the question

  • Diagnosing it as a lock or deadlock problem when no locks are involved
  • Killing waiting sessions as the standard fix, ignoring the resulting client ambiguity
  • Assuming any synchronous setup automatically degrades to asynchronous when a standby dies
  • Turning off synchronous replication during an incident and never turning it back on, with no alert to catch it
  • Claiming you can have both guaranteed zero loss and full write availability with a single standby

context