skip to content

questions

5

What is replication lag between a primary database and its replica, and in what two units is it normally measured?

level: juniorimportance: must knowfreq 66%

answer

  1. Distance between primary's newest change and replica's applied change
  2. Two units: bytes (log position) and seconds (staleness)
  3. Received ≠ applied — quote lag at the applied position
  4. Idle primary makes seconds-lag lie; bytes stay honest
  5. Heartbeat row gives truthful time lag

basics

~20 s

Replication lag is how far a replica trails its primary. It is measured in bytes — how much of the primary's change log the replica has not consumed yet — and in seconds — how old the newest change the replica has applied is.

solid answer

~60 s

A replica is kept current by streaming the primary's change log (WAL in PostgreSQL, binlog in MySQL) and replaying it. **Replication lag is the distance between the primary's newest change and the last change the replica has durably applied.** Two units are used, and they answer different questions: - **Bytes (or log positions)** — the difference between the primary's current log position and the replica's applied position. This is a direct volume measure: how much work the replica still owes. Good for spotting growth trends and for retention safety (falling too far behind can push the replica off the primary's retained log). - **Seconds (time)** — how stale the data on the replica is: the wall-clock age of the newest change it has applied. This is what product requirements are written in ("reports may be up to 5 seconds stale") and what an RPO estimate uses. They are not interchangeable. A megabyte behind is seconds on a fast replica and minutes on an overloaded one; a system with no writes has zero seconds of lag no matter what. Production monitoring tracks both.

go deeper

for a junior

Define it clearly — the replica trails the primary because changes travel through a log and must be replayed — and name both units, bytes and seconds.

for a middle

Add that lag exists per pipeline stage (sent, received, flushed, applied) and that the applied position is the one that matters to a reader.

for a senior

Emphasise that the two units fail in different ways, that a heartbeat row is how you get truthful time lag, and that trend matters more than a spike.

for a principal

Frame lag as two separate budgets — a staleness budget for read traffic and a recovery-point budget for failover — and note that log retention turns extreme lag into a rebuild rather than a catch-up.

## The mechanism that creates lag A relational engine writes every change into an ordered, append-only log before (or as) it changes data pages: the **write-ahead log (WAL)** in PostgreSQL, the **binary log (binlog)** in MySQL. Replication reuses that log. The primary streams log records to the replica; the replica receives them, writes them to its own disk, and then **applies** (replays) them so its copy of the data catches up. Every one of those steps costs time, so a replica is always at least slightly behind. **Replication lag is that distance.** It only reaches zero momentarily, when the replica has applied everything the primary has produced so far. Note the ordering: *received* is not *applied*. A replica can hold ten seconds of change records on its disk and still be serving reads from data that is ten seconds old, because the apply step has not caught up. Lag is normally quoted against the **applied** position, since that is what a query on the replica can actually see. ## Unit 1: bytes / log position Log positions are monotonically increasing offsets — PostgreSQL calls them LSNs (log sequence numbers), MySQL uses file plus offset. Subtracting the replica's position from the primary's current position gives a **byte distance**: how much change data the replica still owes. Why this unit matters: - **It is honest when the system is idle.** If nobody writes for an hour, byte lag is zero because there is genuinely nothing outstanding. - **It measures backlog volume**, which is what predicts recovery time and what threatens log retention. If the primary keeps only, say, the last few gigabytes of log (or a replication slot is retaining log on the replica's behalf), byte lag tells you how close you are to either losing the replica entirely (it must be rebuilt from a fresh copy) or filling the primary's disk. - **It is comparable across replicas** fed from the same primary. What it does not tell you: how stale the data *feels*. Ten megabytes of small OLTP transactions and ten megabytes from one bulk load are very different amounts of staleness. ## Unit 2: seconds / time behind Time lag answers "how old is the newest data visible on this replica?" It is usually derived from the timestamp embedded in the change record the replica most recently applied, compared against now. Why this unit matters: - **Product and SLA language is in seconds.** "The dashboard may lag by up to 5 seconds" is a staleness budget; "we can lose at most 1 second of writes on failover" is a recovery point objective. Both are time statements. - **It is directly actionable** — an operator can decide to drain read traffic away from a replica whose data is a minute old. Its weaknesses are the mirror image of the byte metric. **On an idle primary, time lag collapses to zero or becomes undefined even if the replication link is broken**, because the replica has applied everything it has ever been sent and there is no newer timestamp to compare against. It is also sensitive to clock differences between machines and to how the engine computes it (in a chained setup, timestamps originate from the top-level primary, not the intermediate node). ## Why both, always Mature monitoring records both units side by side because their failure modes do not overlap: | Situation | Byte lag | Time lag | |---|---|---| | Healthy, busy system | small, stable | small, stable | | Replica apply is too slow | grows steadily | grows steadily | | Write burst, replica keeps up | spikes then drains | barely moves | | Primary idle, link broken | grows only when writes resume | reads ~0 — misleading | | Bulk load / index build | large | may stay small | A common practical addition is a **heartbeat**: a tiny row on the primary updated with the current timestamp every second. The replica reads that row and subtracts it from its own clock. Because a write is always flowing, the heartbeat gives a truthful time-lag number even on an otherwise idle system, and it detects a dead replication link instead of reporting zero. ## What normal looks like On a healthy asynchronous replica on the same network, lag is typically single-digit to low-hundreds of milliseconds, with brief spikes during checkpoints, bulk writes, or schema changes. The signal that matters is not a spike but a **trend**: lag that keeps climbing means the replica's apply throughput is lower than the primary's write throughput, and it will not recover on its own until the write rate drops or the replica gets faster.

  • Why can byte lag and time lag disagree sharply during a bulk data load?
    A bulk load produces an enormous volume of change-log records in a very short window, so byte lag spikes into gigabytes. But those records all carry near-identical, very recent timestamps, so as long as the replica is chewing through them steadily the newest applied record is only a second or two old and time lag stays small. The reverse also happens: a single long, slow statement can produce few bytes but hold the applied timestamp far back.
  • If a replica reports zero seconds behind, can you conclude replication is healthy?
    No. Time-based lag on most engines is computed from the change record currently being applied, so when there is nothing to apply — because the primary is idle, or because the receiving side is disconnected and no new records are arriving — the metric reads zero or null. You need the byte distance from the primary's own view of the connection, plus a heartbeat and a connection-liveness check, before calling replication healthy.

Think of a scribe copying a stack of dictated notes. Byte lag is the height of the pile still waiting on the desk; time lag is how old the newest note the scribe has finished actually is. If the speaker stops talking, the pile empties and the scribe looks caught up — even if the courier bringing new notes died an hour ago.

saying these in an interview costs you the question

  • Saying lag is one number, without distinguishing received/written from applied
  • Treating seconds-behind as trustworthy on a low-write or idle system
  • Assuming lag zero means the replica is a synchronous copy — asynchronous replicas are simply momentarily caught up
  • Claiming lag is purely a network problem, ignoring apply-side cost
  • Believing lag self-corrects; sustained lag means apply throughput is below write throughput

context

open as a page

Replication from a primary to a replica happens in several stages. Which stages can lag independently, and what do PostgreSQL's pg_stat_replication view and MySQL's Seconds_Behind_Source field each tell you about them?

level: middleimportance: must knowfreq 56%

basics

~20 s

Stages: the primary generates log records, sends them, the replica writes them, flushes them to disk, then replays them. pg_stat_replication exposes sent/write/flush/replay positions and the matching lag intervals per stage. MySQL's Seconds_Behind_Source covers only the apply stage — receive lag is separate.

open as a page

A read replica repeatedly falls minutes behind its primary during peak hours, even though the replication network link is nowhere near saturated. What are the usual causes of apply-side replication lag, and how would you narrow it down?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Usual causes: apply is far less parallel than the primary's concurrent writers; long or huge transactions arriving as one burst; row changes on the replica lacking a usable index so each change scans; replica hardware or I/O weaker; and replay blocked by conflicting long-running queries on the replica. Narrow it down by finding which stage and which transaction stalls.

open as a page

Time-based replication-lag metrics can report zero or null while a replica is genuinely far behind. Why does that happen, and how would you measure lag so the number is trustworthy?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Time-based lag is derived from the newest change record the replica applied. With no new records — an idle primary, or a dead receiver whose backlog has drained — there is nothing to compare against, so it reads zero or null. Trustworthy measurement combines byte distance from the primary, thread/connection state, and a heartbeat write.

open as a page

How would you choose replication-lag alerting thresholds for a fleet of read replicas, and what should happen automatically when a replica exceeds them?

level: principalimportance: should knowfreq 33%

basics

~20 s

Derive thresholds from what each replica is for: a staleness budget in seconds for read traffic, a recovery-point budget for failover candidates, and a hard byte threshold from log retention. Alert on sustained growth, not single spikes. Above the staleness budget, drain the replica from the read pool automatically, with hysteresis before re-adding.

open as a page