After a failover has already promoted a replica, the old primary comes back online. Why can you usually not just restart it as a replica, and how does a tool such as pg_rewind get it back into the cluster without a full rebuild?
answer
- shared prefix then fork = divergence
- replication never un-applies
- rewind copies only blocks changed after fork
- needs clean shutdown + checksums/wal_log_hints
- MySQL: errant GTIDs, clone instead
basics
~20 sThe old primary may have committed transactions the new one never received, so its history diverged at the failover point and it cannot follow the new stream. A rewind finds the divergence point and undoes only the blocks changed after it, copying those from the new primary, which is far cheaper than a full rebuild.
solid answer
~1 minWhen the old primary died it may have had committed changes that were never shipped. Those changes are on an abandoned branch of the log (in PostgreSQL terms, an old timeline). If you simply start it as a replica of the new primary, its data files contain rows the new history does not, and the log positions do not line up, so replication either refuses to start or the node is silently inconsistent. The safe options are rebuild or rewind. A **rebuild** takes a fresh base backup from the new primary; always correct, but it moves the whole database, which is painful at terabyte scale. A **rewind** (pg_rewind) is a targeted undo. It compares the two nodes' timeline histories to find the divergence point, scans the old primary's WAL from that point to list every block it modified, copies just those blocks plus the relevant configuration and log files from the new primary, and leaves the node in a state where normal recovery can replay forward from the new primary's stream. Cost scales with what changed after divergence, not with database size. Requirements: the old primary must be cleanly shut down first, and the cluster needs data checksums or wal_log_hints enabled so modified blocks are detectable. MySQL's GTID equivalent is to detect errant transactions and, if they exist, clone or rebuild the node.
code
text · 9 lines# 1. make sure it is down cleanly (never let it accept writes first)
pg_ctl -D /var/lib/pgsql/data stop -m fast
# 2. rewind against the current primary
pg_rewind --target-pgdata=/var/lib/pgsql/data \
--source-server="host=db-node-2 user=rewind dbname=postgres" \
--progress
# 3. configure it to follow db-node-2, then start; it replays forward from the fork pointgo deeper
Know that the returned old primary may have extra data, so it cannot just follow the new primary and must be rewound or rebuilt.
Explain divergence and the two remedies with their cost profiles, and that replication cannot remove already-applied changes.
Name the preconditions and pitfalls: clean shutdown, checksums or wal_log_hints, WAL retention, discarded transactions, and keeping the node out of the routing layer while you decide.
Make rejoin part of the automated failover contract and set the standard build so rewind is always possible, with rebuild capacity and time budgeted for the cases it is not.
## Why the returning node is dangerous At the moment of an unplanned failover, the old primary might hold transactions that were committed locally but never streamed. It also very likely wrote WAL after the last record anyone else saw, for instance during its own crash recovery on restart. The new primary meanwhile started its own branch of history from the last position it had. So the two nodes share a common prefix and then diverge. That is exactly the shape that plain replication cannot repair: replication moves forward from a position, it never removes changes a node has already applied. Starting the old primary against the new one produces either an immediate error (mismatched timeline or GTID set) or, in the worst configurations, a node whose data contains extra rows forever. The same reasoning applies to any surviving replica that was ahead of the promoted node. ## Option A: rebuild Take a new base backup from the current primary (pg_basebackup, a snapshot restore, MySQL CLONE or a physical backup) and start the node as a fresh replica. Simple and always correct. The problems are time and load: copying several terabytes can take hours and steals IO and network from a cluster that has just had an incident. It is still the right answer when the divergence is large, the node's storage is suspect, or you cannot meet the rewind preconditions. ## Option B: rewind (PostgreSQL) pg_rewind makes the diverged node a copy of the new primary at the fork point, cheaply. Conceptually: 1. Connect to the source (the new primary) or read its files, and read both nodes' timeline history to compute the **divergence LSN**: the last point at which the two histories agreed. 2. Scan the target's WAL from that LSN forward and build the set of data blocks it touched afterwards. These are the only blocks that can differ in a way that matters. 3. Copy those blocks from the source, plus files that must be taken wholesale (configuration-ish files, relation forks it cannot reason about, and the source's WAL needed for the next step). 4. Leave a control file that makes the node start in recovery from the divergence point, so it replays the new primary's stream forward and converges. Work is proportional to changes since divergence, so a node that was only seconds ahead rewinds in seconds. **Preconditions that trip people up:** - The target must be **cleanly shut down**. If it crashed, start it once to complete crash recovery and shut it down cleanly, which itself can extend the divergence. - The cluster needs **data checksums** or **wal_log_hints** enabled, because otherwise hint-bit changes are not WAL-logged and modified blocks cannot be identified reliably. This must be set before the incident; retrofitting requires a rebuild or an offline checksum-enabling pass. - The source must be sufficiently ahead and must still retain the WAL covering the divergence, so slots and retention settings matter. - Rewinding discards the old primary's extra transactions. If they matter, extract them from its WAL or backups before rewinding, because the operation is not reversible. ## MySQL equivalent With GTIDs the concept is explicit: the returning node's executed GTID set may contain **errant transactions**, that is, GTIDs the new primary does not have. Tooling such as Orchestrator detects this. There is no block-level rewind, so the standard remedies are to clone or re-provision the node, or, when the errant transactions are provably harmless, to inject empty transactions so the sets align, which is a manual, risky action reserved for well-understood cases. ## Fitting it into failover automation Good automation treats rejoin as part of failover, not as an afterthought: on the old primary's return, it must be prevented from accepting writes, checked for divergence, then rewound or rebuilt automatically. Patroni, for example, will run pg_rewind when configured and the preconditions are met, otherwise reinitialise the node. If your runbook is manual, the first step must still be ensuring the node cannot serve or accept application traffic while you decide. ## Interview framing Explain divergence in one sentence, say replication cannot undo applied changes, give rebuild versus rewind with their cost profiles, and then show operational depth by naming the preconditions (clean shutdown, checksums or wal_log_hints, WAL retention) and the fact that the extra transactions are discarded.
- What must be configured in advance for pg_rewind to be usable at incident time?Either data checksums or wal_log_hints must be enabled on the cluster, because pg_rewind identifies changed blocks from WAL records and hint-bit updates are not logged otherwise. You also need the source primary to retain the WAL covering the divergence point and a role with sufficient privileges for the rewind connection. None of these can be turned on usefully after the divergence has happened, which is why they belong in the cluster's standard build.
- How long does a rewind take compared to a rebuild?A rewind's cost is proportional to the volume of blocks the diverged node modified after the fork point, so a node that was ahead by seconds typically rewinds in seconds to minutes regardless of database size. A rebuild copies the entire dataset, so it scales with total size, hours for terabytes, and consumes network and IO on the freshly promoted primary. That asymmetry is why rewind is the default path in automated managers, with rebuild as the fallback.
saying these in an interview costs you the question
- Assuming the old primary can always be restarted as a replica of the new one
- Starting the returned node with application traffic still routed to it
- Believing a rewind preserves the old primary's extra transactions
- Expecting to enable wal_log_hints or checksums after the divergence to rescue the situation
- Thinking MySQL has a block-level equivalent rather than errant-GTID detection plus clone or rebuild