skip to content

What is replication lag between a primary database and its replica, and in what two units is it normally measured?

level: juniorimportance: must knowfreq 66%

answer

  1. Distance between primary's newest change and replica's applied change
  2. Two units: bytes (log position) and seconds (staleness)
  3. Received ≠ applied — quote lag at the applied position
  4. Idle primary makes seconds-lag lie; bytes stay honest
  5. Heartbeat row gives truthful time lag

basics

~20 s

Replication lag is how far a replica trails its primary. It is measured in bytes — how much of the primary's change log the replica has not consumed yet — and in seconds — how old the newest change the replica has applied is.

solid answer

~60 s

A replica is kept current by streaming the primary's change log (WAL in PostgreSQL, binlog in MySQL) and replaying it. **Replication lag is the distance between the primary's newest change and the last change the replica has durably applied.** Two units are used, and they answer different questions: - **Bytes (or log positions)** — the difference between the primary's current log position and the replica's applied position. This is a direct volume measure: how much work the replica still owes. Good for spotting growth trends and for retention safety (falling too far behind can push the replica off the primary's retained log). - **Seconds (time)** — how stale the data on the replica is: the wall-clock age of the newest change it has applied. This is what product requirements are written in ("reports may be up to 5 seconds stale") and what an RPO estimate uses. They are not interchangeable. A megabyte behind is seconds on a fast replica and minutes on an overloaded one; a system with no writes has zero seconds of lag no matter what. Production monitoring tracks both.

go deeper

for a junior

Define it clearly — the replica trails the primary because changes travel through a log and must be replayed — and name both units, bytes and seconds.

for a middle

Add that lag exists per pipeline stage (sent, received, flushed, applied) and that the applied position is the one that matters to a reader.

for a senior

Emphasise that the two units fail in different ways, that a heartbeat row is how you get truthful time lag, and that trend matters more than a spike.

for a principal

Frame lag as two separate budgets — a staleness budget for read traffic and a recovery-point budget for failover — and note that log retention turns extreme lag into a rebuild rather than a catch-up.

## The mechanism that creates lag A relational engine writes every change into an ordered, append-only log before (or as) it changes data pages: the **write-ahead log (WAL)** in PostgreSQL, the **binary log (binlog)** in MySQL. Replication reuses that log. The primary streams log records to the replica; the replica receives them, writes them to its own disk, and then **applies** (replays) them so its copy of the data catches up. Every one of those steps costs time, so a replica is always at least slightly behind. **Replication lag is that distance.** It only reaches zero momentarily, when the replica has applied everything the primary has produced so far. Note the ordering: *received* is not *applied*. A replica can hold ten seconds of change records on its disk and still be serving reads from data that is ten seconds old, because the apply step has not caught up. Lag is normally quoted against the **applied** position, since that is what a query on the replica can actually see. ## Unit 1: bytes / log position Log positions are monotonically increasing offsets — PostgreSQL calls them LSNs (log sequence numbers), MySQL uses file plus offset. Subtracting the replica's position from the primary's current position gives a **byte distance**: how much change data the replica still owes. Why this unit matters: - **It is honest when the system is idle.** If nobody writes for an hour, byte lag is zero because there is genuinely nothing outstanding. - **It measures backlog volume**, which is what predicts recovery time and what threatens log retention. If the primary keeps only, say, the last few gigabytes of log (or a replication slot is retaining log on the replica's behalf), byte lag tells you how close you are to either losing the replica entirely (it must be rebuilt from a fresh copy) or filling the primary's disk. - **It is comparable across replicas** fed from the same primary. What it does not tell you: how stale the data *feels*. Ten megabytes of small OLTP transactions and ten megabytes from one bulk load are very different amounts of staleness. ## Unit 2: seconds / time behind Time lag answers "how old is the newest data visible on this replica?" It is usually derived from the timestamp embedded in the change record the replica most recently applied, compared against now. Why this unit matters: - **Product and SLA language is in seconds.** "The dashboard may lag by up to 5 seconds" is a staleness budget; "we can lose at most 1 second of writes on failover" is a recovery point objective. Both are time statements. - **It is directly actionable** — an operator can decide to drain read traffic away from a replica whose data is a minute old. Its weaknesses are the mirror image of the byte metric. **On an idle primary, time lag collapses to zero or becomes undefined even if the replication link is broken**, because the replica has applied everything it has ever been sent and there is no newer timestamp to compare against. It is also sensitive to clock differences between machines and to how the engine computes it (in a chained setup, timestamps originate from the top-level primary, not the intermediate node). ## Why both, always Mature monitoring records both units side by side because their failure modes do not overlap: | Situation | Byte lag | Time lag | |---|---|---| | Healthy, busy system | small, stable | small, stable | | Replica apply is too slow | grows steadily | grows steadily | | Write burst, replica keeps up | spikes then drains | barely moves | | Primary idle, link broken | grows only when writes resume | reads ~0 — misleading | | Bulk load / index build | large | may stay small | A common practical addition is a **heartbeat**: a tiny row on the primary updated with the current timestamp every second. The replica reads that row and subtracts it from its own clock. Because a write is always flowing, the heartbeat gives a truthful time-lag number even on an otherwise idle system, and it detects a dead replication link instead of reporting zero. ## What normal looks like On a healthy asynchronous replica on the same network, lag is typically single-digit to low-hundreds of milliseconds, with brief spikes during checkpoints, bulk writes, or schema changes. The signal that matters is not a spike but a **trend**: lag that keeps climbing means the replica's apply throughput is lower than the primary's write throughput, and it will not recover on its own until the write rate drops or the replica gets faster.

  • Why can byte lag and time lag disagree sharply during a bulk data load?
    A bulk load produces an enormous volume of change-log records in a very short window, so byte lag spikes into gigabytes. But those records all carry near-identical, very recent timestamps, so as long as the replica is chewing through them steadily the newest applied record is only a second or two old and time lag stays small. The reverse also happens: a single long, slow statement can produce few bytes but hold the applied timestamp far back.
  • If a replica reports zero seconds behind, can you conclude replication is healthy?
    No. Time-based lag on most engines is computed from the change record currently being applied, so when there is nothing to apply — because the primary is idle, or because the receiving side is disconnected and no new records are arriving — the metric reads zero or null. You need the byte distance from the primary's own view of the connection, plus a heartbeat and a connection-liveness check, before calling replication healthy.

Think of a scribe copying a stack of dictated notes. Byte lag is the height of the pile still waiting on the desk; time lag is how old the newest note the scribe has finished actually is. If the speaker stops talking, the pile empties and the scribe looks caught up — even if the courier bringing new notes died an hour ago.

saying these in an interview costs you the question

  • Saying lag is one number, without distinguishing received/written from applied
  • Treating seconds-behind as trustworthy on a low-write or idle system
  • Assuming lag zero means the replica is a synchronous copy — asynchronous replicas are simply momentarily caught up
  • Claiming lag is purely a network problem, ignoring apply-side cost
  • Believing lag self-corrects; sustained lag means apply throughput is below write throughput

context