skip to content

When a replica is promoted to primary, for example using the PostgreSQL pg_promote() function or MySQL's CHANGE REPLICATION SOURCE command, what actually changes on that node, and what must then happen to the other replicas?

level: middleimportance: must knowfreq 46%

answer

  1. promotion = end of recovery, not data movement
  2. new timeline ID + .history file
  3. repoint replicas; ahead ones must rewind
  4. slots, archiving, backups follow the writer
  5. routing is a separate step

basics

~20 s

Promotion ends recovery on that node: it finishes applying what it already received, opens for writes, and starts its own branch of the change stream (a new PostgreSQL timeline or a new binlog source). Every other replica must be repointed at it, and any replica ahead of it must be rewound or rebuilt.

solid answer

~60 s

On the promoted node, promotion means leaving replica mode. In PostgreSQL, pg_promote() (or the promote signal) ends recovery: the node replays what it already has, performs its end-of-recovery work, increments the **timeline ID**, writes a timeline history file, and opens read-write. In MySQL the equivalent is stopping the replication threads, ensuring the relay log is applied, and clearing the replication source so the server accepts writes; other servers are then pointed at it with CHANGE REPLICATION SOURCE TO. The important part is that a new branch of history begins at the promotion point. Everything downstream must be attached to that branch: - Remaining replicas must follow the new primary. With GTIDs or timeline-aware streaming they can often just be repointed; with file/offset positions the coordinates must be translated. - A replica that received changes the new primary never got is ahead on an abandoned branch and cannot follow; it needs a rewind or a rebuild. - Replication slots, delayed replicas, and backup/archiving jobs must be recreated or repointed. And nothing about promotion moves client traffic; the routing layer must be updated separately.

code

sql · 10 lines
sql
-- PostgreSQL: end recovery on the standby and open it read-write
SELECT pg_promote(wait => true, wait_seconds => 60);
SELECT pg_is_in_recovery();   -- now false

-- MySQL: point a surviving replica at the newly promoted server
STOP REPLICA;
CHANGE REPLICATION SOURCE TO
  SOURCE_HOST = 'db-node-2',
  SOURCE_AUTO_POSITION = 1;
START REPLICA;

go deeper

for a junior

Know that promotion takes the node out of recovery and makes it writable, and that other replicas must then follow the new primary.

for a middle

Name the concrete mechanics: timeline increment or new binlog stream, applying pending changes first, and repointing the remaining replicas.

for a senior

Cover the ahead-of-primary replica case, slot and archive re-creation, and the fact that routing is a separate and equally necessary step.

for a principal

Talk about making this repeatable: automation that owns the whole fleet transition, GTID or timeline-aware replication to avoid manual coordinate translation, and rehearsals that prove it.

## What a replica is doing before promotion A physical replica is permanently in recovery: it receives the primary's change log (PostgreSQL WAL, MySQL binlog copied into relay logs), applies it, and refuses writes. Its state is defined by a position in that stream: an LSN in PostgreSQL, a file/offset or GTID set in MySQL. ## What promotion does on that node Promotion terminates recovery. Concretely: - The node applies whatever it has already received (a manager may first wait for pending relay log entries to be applied so nothing received is discarded). - It performs end-of-recovery bookkeeping: PostgreSQL writes an end-of-recovery record and checkpoint, and MySQL stops the IO and applier threads. - It begins its own branch of history. In PostgreSQL this is explicit: the **timeline ID** increments (say 1 to 2) and a .history file records the LSN at which timeline 2 diverged from timeline 1. In MySQL the equivalent is that the new primary begins writing its own binlog, and under GTIDs new transactions carry its server UUID. - It opens for writes, and read-only enforcement is lifted. PostgreSQL exposes this as pg_promote() (a function callable from SQL) or a promote signal via pg_ctl. In MySQL, promotion is not a single command: you stop replication on the target and let writes in, then use CHANGE REPLICATION SOURCE TO on the other servers to point them at it. Promotion is fast because there is no data movement, only bookkeeping, but it is not instantaneous: applying a large backlog of received-but-unapplied changes can dominate, which is one reason apply lag is monitored separately from receive lag. ## Why the branch matters The new timeline or binlog stream means downstream nodes must be attached to the correct branch. Three cases: 1. **Behind the new primary, on the same branch.** The straightforward case: repoint and it streams forward. PostgreSQL streaming replication can follow a timeline switch if the replica can read the history file, and GTID-based MySQL replication can work out what it still needs. 2. **Exactly current.** Repointing is trivial. 3. **Ahead of the new primary.** It applied transactions from the old primary that the new one never received. Those changes are on an abandoned branch, so following the new primary would mean its data contains rows the new history does not; replication will error or, in bad configurations, diverge silently. Such a node must be rewound to the divergence point or rebuilt from a new base backup. Case 3 is common precisely in unplanned failovers, and it applies to the old primary too when it returns. ## Everything else that has to move Promotion changes the writer, not the ecosystem around it: - **Client routing.** Connection pools, proxies, service discovery, or a virtual IP must now point at the new node. Until then the promotion has achieved nothing user-visible, and stale routes are also a split-brain risk if the old node is still up. - **Replication slots and retention.** PostgreSQL physical slots live on the node serving the stream, so slots for surviving replicas must exist on the new primary or WAL retention will not protect them. MySQL binlog retention likewise now matters on the new node. - **Backups and archiving.** The archive/backup job followed the old primary; it must follow the new one, and the backup catalogue must understand the new timeline. - **Monitoring and scheduled jobs** that assumed a fixed hostname for the writer. - **Read-only settings.** A demoted or promoted node's read-only flags, and any application-level guards, must match its new role. ## Practical failure modes - Promoting a replica that still had unapplied relay log entries and discarding them, losing more than necessary. - Repointing an ahead-of-primary replica and having replication break hours later on a duplicate key. - Forgetting slot or retention setup so a lagging replica falls off and requires a full rebuild. - Promoting successfully but never updating routing, so the application still writes to the old node. ## Interview framing Say promotion ends recovery and starts a new branch of the log, name the PostgreSQL timeline increment as the concrete artifact, then immediately cover the fleet: repoint replicas, rewind or rebuild those ahead, recreate slots and backups, and change client routing.

  • A surviving replica had applied changes the newly promoted primary never received. What are your options?
    It is on an abandoned branch, so it cannot simply follow the new primary. Either rewind it to the divergence point using a tool such as pg_rewind (or re-provision it under GTIDs when the extra transactions can be identified), or rebuild it from a fresh base backup of the new primary. Rewinding is far faster on large databases, but either way those extra transactions are discarded.
  • Why does a PostgreSQL promotion increment the timeline ID rather than just continuing the WAL sequence?
    Because after a promotion two different histories can exist from the same LSN onward: the old primary might have written more WAL that nobody else has. The timeline ID plus its history file makes that fork explicit, so archives and replicas can tell which branch a WAL segment belongs to and refuse to mix them. Without it, WAL files from divergent branches could be applied together and silently corrupt the database.

saying these in an interview costs you the question

  • Thinking promotion copies or rebuilds data rather than ending recovery
  • Believing every surviving replica can always just be repointed
  • Forgetting to update client routing, so the application still writes to the old node
  • Ignoring replication slots and WAL/binlog retention on the new primary
  • Promoting before the received relay log or WAL is applied, discarding data unnecessarily

context