skip to content

A read replica repeatedly falls minutes behind its primary during peak hours, even though the replication network link is nowhere near saturated. What are the usual causes of apply-side replication lag, and how would you narrow it down?

level: seniorimportance: must knowfreq 52%

answer

  1. Prove apply lag first: received current, replayed behind
  2. Many writers on primary vs one applier — throughput deficit
  3. Long/bulk transactions land as one slab at commit
  4. No PK/index → full scan per row event → catastrophic
  5. Standby long queries → recovery conflict pauses replay

basics

~20 s

Usual causes: apply is far less parallel than the primary's concurrent writers; long or huge transactions arriving as one burst; row changes on the replica lacking a usable index so each change scans; replica hardware or I/O weaker; and replay blocked by conflicting long-running queries on the replica. Narrow it down by finding which stage and which transaction stalls.

solid answer

~1 min

Once the receive side is current and the apply side is not, the causes cluster: - **Apply concurrency mismatch.** The primary commits with hundreds of concurrent sessions; the replica replays with one recovery process (or a limited pool of applier workers). A write rate that the primary absorbs easily can exceed what serialized replay can sustain. - **Long or bulk transactions.** A change stream that ships a transaction at commit delivers a huge unit of work all at once; a 20-minute batch job lands as a spike the applier must chew through. - **Missing usable index on the replica for row-level changes.** If each replicated UPDATE/DELETE must locate its row without a primary key or matching index, the applier scans per row — throughput collapses by orders of magnitude. - **Weaker or busier replica hardware**, backups running on it, or I/O contention from serving reads. - **Replay conflicts**: on PostgreSQL, long-running standby queries collide with replay of cleanup, pausing recovery (or cancelling queries), and `hot_standby_feedback` shifts the cost onto the primary. Narrowing it: confirm receive is current, look at what the applier is executing right now, check per-transaction size, check indexes on the replicated tables, check whether replay is waiting rather than working.

go deeper

for a junior

Name two plausible causes — a big batch job on the primary and a replica that is smaller or busier — and know that applying changes is separate from receiving them.

for a middle

Explain the concurrency asymmetry and transaction shape, and know that row-level apply needs a primary key or index to avoid per-row scans.

for a senior

Lead with the diagnostic split, distinguish applier-working from applier-blocked, and discuss recovery conflicts and the hot_standby_feedback trade-off explicitly.

for a principal

Separate transient spikes from a sustained throughput deficit, and turn the answer into policy: lag budgets, batch-chunking rules, index parity between primary and replicas, and when the real answer is sharding or fewer writes.

## Step zero: prove it is apply lag Before theorising, split the pipeline. If the replica has received and flushed nearly everything the primary has produced but replayed far less, the backlog is on the apply side and network/bandwidth theories are dead. In PostgreSQL that is a large `flush_lsn → replay_lsn` gap; in MySQL it is a large gap between the received position and the executed position. Anything else — a gap before the received position — is a sender, network, or replica-disk story instead. ## Cause 1: the concurrency asymmetry This is the structural reason apply lag exists at all. A primary executes writes with as much parallelism as its clients supply: 200 sessions committing small transactions simultaneously. The replica's job is to reproduce the *result* in a way that preserves ordering guarantees, and historically it does so with a single replay process. One CPU replaying what many CPUs produced only works because replay is cheaper per change than the original execution (no parsing, no planning, no client round trips). When the primary's write rate rises past that advantage, lag grows without bound — it is a **throughput deficit, not a spike**, and it will not recover until the write rate falls. Remedies are architectural: enable multi-threaded/parallel apply where the engine supports it (MySQL's multi-threaded applier with logical-clock/writeset-based parallelism; PostgreSQL parallelizes only certain aspects of recovery, so physical standbys remain largely serial), reduce write volume, or shard. ## Cause 2: transaction shape How work is packaged matters as much as its volume. - **Long-running write transactions** may be transmitted only when they commit (typical for row-based logical streams and for MySQL's binlog, which writes on commit). A transaction that ran for 20 minutes on the primary arrives as one indivisible slab; the replica cannot start it earlier and cannot split it, so lag jumps by roughly the transaction's duration the moment it lands. - **Bulk operations** — mass UPDATE/DELETE, index builds, table rewrites, large imports — generate change volume out of all proportion to their statement count. In row-level replication a single `DELETE FROM t WHERE created_at < ...` touching 10 million rows becomes 10 million row events. - **DDL** is typically serialized and can block the applier entirely while it runs. The operational fix is batching: chunk large writes into bounded transactions with pauses, and schedule them against a lag budget. ## Cause 3: no usable index on the apply path This one produces the most spectacular numbers. In row-level replication the applier receives "the row that had these identifying values changed to this" and must locate that row locally. If the table has a primary key (or a unique index designated as the row identity) and the replica has it too, this is one index probe per row: fast. If the table has no primary key — or the replica's copy of the table is missing the relevant index, which happens when someone tailors indexes on a reporting replica — the applier does a **full scan per row event**. A batch of 100,000 row changes on a 50-million-row table becomes 100,000 sequential scans. Lag goes from milliseconds to hours. Diagnosis: look at what the applier thread is executing and its state; a scanning applier shows as CPU/IO-bound on one table with no progress on the position. The cure is a primary key or a suitable unique index on both sides. ## Cause 4: the replica is simply weaker or busier Replicas are frequently provisioned smaller, live on cheaper storage, take the nightly backup, feed an ETL extract, or serve heavy analytical reads. Replay competes with all of that for I/O and buffer cache. A replica whose working set does not fit in memory pays a random read for every page the applier must modify, while the primary had those pages hot. This is why a replica can lag under exactly the workload the primary handles comfortably. ## Cause 5: replay is blocked, not slow On PostgreSQL hot standbys, replay of WAL that removes rows still visible to a long-running standby query creates a **recovery conflict**. Depending on configuration, replay pauses up to `max_standby_streaming_delay` and then cancels the offending query. So a single analyst running a 40-minute report can pin replay and manufacture 40 minutes of lag for every reader. Turning on `hot_standby_feedback` prevents the conflict by making the primary retain the needed row versions — trading replica lag for table bloat and delayed cleanup on the primary. That trade is a real decision, not a free fix. Similar blocking happens on any engine when the applier waits on a lock held by a local session (for example, a session on the replica holding a lock on a table the applier must modify in a writable/logical setup). ## A narrowing procedure 1. Split receive vs apply; stop here if the gap is before receive. 2. Ask whether the applier is **working or waiting** — CPU/IO busy versus blocked on a conflict or lock. 3. If working: what is it executing? Identify the table and the transaction; check its size and whether it is a bulk/DDL event. 4. Check the row-identity/index situation for that table on both sides. 5. Correlate lag onset with the primary's write pattern — batch jobs, cron, deploys, backfills — and with replica-side activity such as backups and long reports. 6. Distinguish a **drainable spike** (lag rises then falls; capacity is adequate) from a **sustained deficit** (lag rises monotonically; apply throughput is below write throughput and only a structural change fixes it).

  • How does hot_standby_feedback change the picture, and what does it cost?
    It makes the standby report its oldest running query's snapshot to the primary, so the primary refrains from cleaning up row versions the standby still needs. Recovery conflicts disappear, so replay stops pausing and lag from long standby reads goes away. The cost lands on the primary: dead tuples accumulate, vacuum falls behind, tables and indexes bloat, and transaction-ID/xmin horizon advance is held back — so it trades replica lag for primary maintenance debt.
  • You find lag rises steadily every night at the same time and drains by morning. Is that a capacity problem?
    Not necessarily. A spike that reliably drains means apply throughput exceeds the average write rate and the replica only falls behind during a burst — usually a batch job or backfill. It becomes a problem only if peak lag breaches the staleness or recovery-point budget, or if the backlog risks exceeding log retention. The usual fix is chunking the batch job into bounded transactions with pauses, not resizing the replica.

saying these in an interview costs you the question

  • Jumping to 'the network is slow' without splitting receive from apply
  • Assuming adding CPUs to the replica fixes serialized replay
  • Not knowing that a missing primary key makes row-based apply scan per row
  • Treating a steadily rising lag trend as a transient spike
  • Recommending hot_standby_feedback as a free fix with no mention of primary bloat
  • Believing lag on the replica cannot be caused by activity on the replica itself

context