skip to content

A Redis replica reconnects after a 20-second network blip and performs a full resynchronization instead of a partial one. Which mechanism should have prevented that, and how do you tune it?

level: seniorimportance: should knowfreq 42%

answer

  1. PSYNC replid offset+1 -> +CONTINUE or +FULLRESYNC
  2. Backlog = circular buffer, repl-backlog-size default 1mb
  3. Size = peak write bytes/s x tolerated outage
  4. repl-backlog-ttl 3600 frees it when no replicas
  5. PSYNC2: replid2 lets promoted replica accept partials

basics

~20 s

Partial resync uses the primary's replication backlog, a fixed-size circular buffer of the recent write stream (repl-backlog-size, default 1mb). If the replica's offset has fallen out of it, or the replication ID no longer matches, the primary must send everything again. Size the backlog as peak write bytes per second times the outage you want to survive.

solid answer

~60 s

On reconnect the replica sends `PSYNC <replid> <offset+1>`. The primary can answer `+CONTINUE` and ship only the missing bytes if two conditions hold: the replication ID matches its current history, and the requested offset is still inside the **replication backlog**, a circular buffer of the recent replication stream sized by `repl-backlog-size` (default 1 MB). Twenty seconds of a busy write stream easily exceeds 1 MB, so the offset falls off the end and the primary is forced into a `+FULLRESYNC` with a fork and a whole-dataset transfer. The sizing rule is straightforward: backlog bytes should be at least peak replication throughput in bytes per second multiplied by the longest disconnection you want to ride out. It is a flat memory cost on the primary, so tens or hundreds of megabytes are usually cheap insurance. Other causes of an unexpected full sync: the primary restarted and lost its replication ID, the replica was pointed at a different primary, or `repl-backlog-ttl` (default 3600s) freed the backlog after no replicas were attached. Watch `sync_full`, `sync_partial_ok` and `sync_partial_err` in `INFO stats`.

code

text · 15 lines
text
# sample twice, 10s apart, at peak traffic
> INFO replication
master_repl_offset:840512300
# ... 10 seconds later ...
master_repl_offset:860912300
# (860912300-840512300)/10 = ~2.04 MB/s

# survive a 60s outage: 2.04MB/s * 60 ~= 123MB -> round up
> CONFIG SET repl-backlog-size 268435456
> CONFIG REWRITE

> INFO stats
sync_full:12
sync_partial_ok:3
sync_partial_err:9   # replicas asked to resume and could not

go deeper

for a junior

Know that Redis keeps a small buffer of recent writes so a briefly disconnected replica can catch up without copying everything again.

for a middle

Name repl-backlog-size, explain PSYNC with replication ID and offset leading to +CONTINUE or +FULLRESYNC, and why the 1 MB default is often too small.

for a senior

Do the sizing arithmetic from measured master_repl_offset deltas, cover repl-backlog-ttl, replication-ID mismatch after restart, and monitor sync_full versus sync_partial_ok.

for a principal

Treat it as trading a fixed slice of primary memory against fork and bandwidth cost during network instability, and account for PSYNC2 behaviour so promotions do not stampede every replica into a full sync at once.

## The backlog is the whole trick Redis maintains, on every primary that has (or recently had) replicas, a fixed-size circular buffer called the replication backlog holding the most recent bytes of the replication stream. It exists solely so a briefly disconnected replica can resume rather than restart. Independently, the primary tracks a monotonically increasing offset: the total number of replication-stream bytes ever produced in the current history, and a replication ID naming that history. ## What happens on reconnect The replica remembers the last offset it applied and the replication ID it was following. It reconnects and sends `PSYNC <replid> <offset+1>`. The primary checks two things. First, does the requested replication ID match its own current ID (or, thanks to PSYNC2, its recorded previous ID)? Second, is `offset+1` still within the range currently held in the backlog? If both are satisfied, it answers `+CONTINUE` and writes only the missing byte range, and the replica is caught up in milliseconds. Otherwise it answers `+FULLRESYNC` and pays for a fork, a snapshot and a full transfer. ## Why 20 seconds was enough to lose it The backlog defaults to 1 MB, which is tiny for anything but a low-traffic instance. If the primary produces, say, 2 MB/s of replication stream, the entire buffer is overwritten every half second, so any disconnection longer than that is unrecoverable by partial resync. Compute the requirement directly: `repl-backlog-size >= peak_write_bytes_per_second * max_tolerated_disconnect_seconds`. To measure the numerator, sample `master_repl_offset` from `INFO replication` a few seconds apart and divide the delta by the interval, at peak traffic rather than at the daily average. Then add headroom, because the cost is a fixed allocation on the primary and not proportional to the dataset. ## The other reasons a full sync happens - **Replication ID mismatch.** If the primary process restarted, it comes up with a fresh history and no memory of the old one (unless it reloaded an RDB that stores the replid), so every replica must fully resync. Repointing a replica at a different primary has the same effect. - **Backlog released.** `repl-backlog-ttl` (default 3600 seconds) frees the buffer once no replica has been attached for that long, so a replica that has been away longer cannot resume. - **Promotion without PSYNC2 support.** Before Redis 4.0, promoting a replica to primary changed the history and forced all sibling replicas into full syncs. PSYNC2 fixed this: a promoted replica keeps the old replication ID as a secondary ID (`master_replid2` with `second_repl_offset`), so siblings that were following the old primary can partially resync against it. Chained sub-replicas also preserve offsets across the change. ## Operating it `INFO stats` gives you `sync_full`, `sync_partial_ok` and `sync_partial_err`. A healthy deployment with occasional network noise should show partial successes climbing and full syncs staying near constant; a rising `sync_full` means real load on the primary each time. `INFO replication` exposes `repl_backlog_active`, `repl_backlog_size`, `repl_backlog_first_byte_offset` and `repl_backlog_histlen`, from which you can see how much history you actually hold. Note also that the backlog is not the same thing as the per-replica output buffer that fills during a full sync; both can force a full resync, but the backlog covers reconnects while the output buffer covers the initial transfer window, and they are tuned separately. The judgement call is simply memory: a few hundred megabytes of backlog on a primary is usually far cheaper than repeated forks, transfer bandwidth and degraded replicas whenever the network hiccups.

  • What is the downside of setting repl-backlog-size very large?
    It is a flat memory allocation on the primary, held for as long as replicas are attached, so it eats into the same RAM budget as your dataset and can push you toward maxmemory and eviction. It does not otherwise slow anything down, so the practical approach is to size it from measured peak throughput times the outage you want to survive, with headroom, rather than guessing large.
  • After a replica is promoted to primary, why do its sibling replicas usually not need a full resynchronization?
    Because of PSYNC2: the promoted node keeps the old replication ID as a secondary ID together with the offset at which the switch happened. Siblings that were following the old primary present that ID and their offset, and the new primary accepts a partial resynchronization if the offset is still inside its backlog. Without that mechanism, every promotion would trigger a fork and full transfer on every sibling at once, exactly when the system is already under stress.

saying these in an interview costs you the question

  • Confusing the replication backlog with the AOF or with the per-replica output buffer
  • Believing the default 1 MB backlog is adequate for a busy primary
  • Thinking a replica can resume from any offset regardless of age
  • Assuming a primary restart preserves the replication history for partial resync
  • Sizing the backlog from average rather than peak write throughput

context