skip to content

Async Replication & Replica Reads

You will learn how replicas sync — full RDB transfer, then a command stream with a backlog for partial resync — and why reads from replicas can be stale or even lost on failover. Interviewers ask about WAIT and replication lag to test whether you know Redis replication is asynchronous by default.

part ofRedisoverview, primer and where to startread it →
on this pageshow

questions

5

What does the Redis REPLICAOF command do to the instance it is run on, what can that instance still serve, and how do you detach it again?

level: juniorimportance: must knowfreq 50%

answer

  1. REPLICAOF host port, was SLAVEOF pre-5.0
  2. Replica discards its own data on attach
  3. replica-read-only yes = -READONLY error on writes
  4. REPLICAOF NO ONE = promote, keeps data
  5. Runtime only; config file for restart survival

basics

~20 s

REPLICAOF host port makes the instance a replica: it discards its own dataset, copies the primary's, and then applies the primary's write stream. It serves reads but rejects writes (replica-read-only is on by default). REPLICAOF NO ONE detaches it into an independent primary, keeping the data it has.

solid answer

~50 s

`REPLICAOF <host> <port>` (called SLAVEOF before Redis 5.0) turns the instance you run it on into a replica of the given primary. It connects, performs a synchronization that replaces its own dataset with the primary's, and from then on continuously applies the primary's write stream. Because `replica-read-only yes` is the default, write commands sent to it are refused with `-READONLY You can't write against a read only replica`, while reads are served normally. `REPLICAOF NO ONE` reverses it: the instance stops replicating and becomes a standalone primary, keeping whatever data it has already received. It does not roll back or re-fetch anything, and it starts a new replication history. The command changes runtime state only, so put `replicaof <host> <port>` in the config file if the role must survive a restart. `INFO replication` shows `role`, `master_link_status` and the connected replicas.

code

text · 18 lines
text
# on the instance that should become a replica
> REPLICAOF 10.0.0.4 6379
OK
> INFO replication
role:slave
master_host:10.0.0.4
master_link_status:up
slave_read_only:1

> SET k v
(error) READONLY You can't write against a read only replica.

# detach and become an independent primary, keeping current data
> REPLICAOF NO ONE
OK
> INFO replication
role:master
connected_slaves:0

go deeper

for a junior

Core recall: REPLICAOF makes this node a read-only copy of a primary, REPLICAOF NO ONE makes it standalone again, and INFO replication shows the role.

for a middle

Add that the replica's own data is discarded on attach, why read-only is the default, and that the change is runtime-only unless written to config.

for a senior

Discuss chained replicas, the risks of writable replicas, and the operational traps around restarts with stale or missing replicaof lines and dual-primary divergence after manual promotion.

for a principal

Position REPLICAOF as the primitive that automated failover tooling drives, and note that safe promotion is really about fencing the old primary and repointing clients, not about the command itself.

## Setting up a replica Redis replication is single-primary: one instance accepts writes and any number of replicas copy it. You create the relationship from the replica side with `REPLICAOF <host> <port>`; nothing is configured on the primary, which simply accepts replication connections. The command returns OK immediately and the actual synchronization proceeds in the background, so `INFO replication` on the replica moves through `master_link_status:down` to `up` once the initial sync completes. Critically, becoming a replica is destructive to the replica's own data: the instance abandons its dataset and adopts the primary's. Pointing a populated instance at a primary by mistake is a data-loss event, not a merge. ## What a replica can serve With the default `replica-read-only yes`, the replica accepts read commands and rejects writes with a `-READONLY` error. This exists to protect the replication invariant: a replica's dataset must be a copy of the primary's, and a locally-applied write has no way to reach the primary or other replicas. You can set `replica-read-only no` to make a *writable replica*, and it is occasionally used for ephemeral per-node scratch data, but locally written keys diverge from the primary and are wiped by the next full synchronization, so it is a sharp tool. Replicas can themselves have replicas (chained or cascading replication): a sub-replica syncs from another replica rather than from the primary, which spreads the fan-out cost of shipping the stream at the price of extra lag on the tail of the chain. ## Detaching `REPLICAOF NO ONE` promotes the instance to an independent primary. It keeps its current dataset, which is the primary's data as of the last thing it applied, and begins accepting writes. It also starts a fresh replication history, generating a new replication ID while remembering the previous one so that other replicas of the old primary can, in some cases, attach to it without a full resynchronization. This manual promotion is the raw mechanism underneath automated failover systems, which additionally handle election, fencing the old primary, and repointing clients; on your own you must do that repointing yourself, and you must ensure the old primary cannot keep taking writes, or you get two divergent primaries. ## Runtime versus configuration Both `REPLICAOF` forms change the running state and are persisted only if you also write the config, for example via `CONFIG REWRITE`. A replica restarted without `replicaof` in its configuration file comes back as a standalone primary holding an old copy of the data, which is a classic way to accidentally serve stale data or accept writes that nothing else will ever see. Conversely, leaving a stale `replicaof` line in a promoted node's config makes it demote itself, and lose its data, on the next restart. ## What to check `INFO replication` is the one-stop view. On a primary it lists `role:master`, `connected_slaves` and one line per replica with its address, state and offset. On a replica it shows `role:slave`, `master_host`, `master_link_status`, `master_last_io_seconds_ago`, and whether an initial sync is in progress. Those fields are how you verify that a newly attached replica really caught up rather than merely accepting the command.

  • An engineer runs REPLICAOF on an instance that already holds production data. What happens to that data?
    It is discarded. The instance synchronizes from the primary and replaces its dataset with the primary's copy, so anything unique to it is gone unless a snapshot exists on disk from before. This is why the direction of the command matters so much: you run REPLICAOF on the node that should be overwritten, never on the source of truth.
  • Why is replica-read-only enabled by default?
    Because a write applied only on a replica cannot propagate anywhere: replication flows one way, so the replica would diverge from the primary and from its siblings. Worse, the divergence is silent until the next full synchronization erases it. Read-only replicas keep the copy semantics honest, and the setting can be relaxed deliberately for throwaway local data.

saying these in an interview costs you the question

  • Running REPLICAOF on the primary instead of the replica and wiping the good dataset
  • Expecting a replica to merge its existing keys with the primary's
  • Thinking replicas accept writes by default
  • Assuming REPLICAOF survives a restart without a config change
  • Believing REPLICAOF NO ONE discards data or rewinds the replica

context

open as a page

Walk through what happens on both the primary and the replica during a Redis full synchronization, when a replica connects for the first time.

level: middleimportance: must knowfreq 55%

basics

~20 s

The replica sends PSYNC with no known history; the primary replies FULLRESYNC with its replication ID and offset, produces an RDB snapshot from a forked child (to disk or streamed directly), and buffers all writes made meanwhile in that replica's output buffer. The replica flushes its data, loads the RDB, then applies the buffered and ongoing stream.

open as a page

Redis replicates asynchronously. When a client receives +OK for a SET command, what has actually been guaranteed, and what does the WAIT command add on top?

level: seniorimportance: must knowfreq 48%

basics

~20 s

+OK means only that the write was applied in the primary's memory. Replicas are sent the write after the reply, so an acknowledged write can vanish if the primary dies before propagating and a replica is promoted. WAIT numreplicas timeout blocks until that many replicas acknowledge the offset and returns how many did; it never rolls anything back.

open as a page

Your service sends read traffic to Redis replicas. What consistency properties can a read from a replica be relied on for, and which ones can it not? Explain what the `master_repl_offset`, `slave_repl_offset` and `master_link_status` fields of the Redis `INFO replication` section let you observe, and what changes when the replica is configured with `replica-serve-stale-data no` and its link to the primary drops.

level: middleimportance: should knowfreq 45%

basics

~20 s

A replica read reflects the primary at an earlier offset — staleness tracks replication lag but is never bounded by a promise. No read-your-writes after a primary write, and values can appear to move backwards between replicas. INFO replication shows the offsets and master_link_status; replica-serve-stale-data no makes a disconnected replica error instead of answering.

open as a page

A Redis replica reconnects after a 20-second network blip and performs a full resynchronization instead of a partial one. Which mechanism should have prevented that, and how do you tune it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Partial resync uses the primary's replication backlog, a fixed-size circular buffer of the recent write stream (repl-backlog-size, default 1mb). If the replica's offset has fallen out of it, or the replication ID no longer matches, the primary must send everything again. Size the backlog as peak write bytes per second times the outage you want to survive.

open as a page