Your service sends read traffic to Redis replicas. What consistency properties can a read from a replica be relied on for, and which ones can it not? Explain what the `master_repl_offset`, `slave_repl_offset` and `master_link_status` fields of the Redis `INFO replication` section let you observe, and what changes when the replica is configured with `replica-serve-stale-data no` and its link to the primary drops.
answer
- Replica read = consistent snapshot of an *older* offset
- No read-your-writes after writing to the primary
- Two replicas at different offsets → value moves backwards
- INFO replication: master_repl_offset − slave_repl_offset = lag
- master_link_status:down + replica-serve-stale-data no → -MASTERDOWN
basics
~20 sA replica read reflects the primary at an earlier offset — staleness tracks replication lag but is never bounded by a promise. No read-your-writes after a primary write, and values can appear to move backwards between replicas. INFO replication shows the offsets and master_link_status; replica-serve-stale-data no makes a disconnected replica error instead of answering.
solid answer
~50 sA replica applies the primary's write stream asynchronously, so a read there is a consistent snapshot of the primary as of *some* past offset — never a guarantee of currency. Three consequences: (1) no read-your-writes — a client that wrote to the primary may read the old value from a replica right after; (2) non-monotonic reads — two replicas sit at different offsets, so a client bouncing between them can see a value go backwards; (3) staleness is unbounded — it tracks lag, and lag has no ceiling under load or link loss. `INFO replication` is the observable: `master_repl_offset` on both sides, `slave_repl_offset` on the replica, and `master_link_status:up|down` plus `master_last_io_seconds_ago`. Offset delta is your lag proxy. With the default `replica-serve-stale-data yes`, a replica whose link is down keeps answering with arbitrarily old data. Set to `no`, it replies `-MASTERDOWN` for most commands — trading a correctness risk for an availability one. (In Redis Cluster, a connection additionally needs `READONLY` to be served by a replica at all.)
code
text · 15 lines# on the primary
127.0.0.1:6379> INFO replication
role:master
connected_slaves:2
slave0:ip=10.0.0.5,port=6379,state=online,offset=97431102,lag=0
slave1:ip=10.0.0.6,port=6379,state=online,offset=97402118,lag=1
master_repl_offset:97431190
# on replica slave1 — 29,072 bytes of stream behind slave0
127.0.0.1:6379> INFO replication
role:slave
master_link_status:up
master_last_io_seconds_ago:0
slave_repl_offset:97402118
slave_read_only:1go deeper
Know that a replica lags behind the primary, so a read there can be out of date, and that you check INFO replication to see how far behind it is.
Explain read-your-writes and non-monotonic reads concretely, name master_repl_offset / slave_repl_offset / master_link_status as the observables, and state what replica-serve-stale-data no does when the link drops.
Frame it as unbounded staleness with no protocol-level bound, show how you alert on offset delta and link status, and describe concrete mitigations — primary-pinned reads after a write, session stickiness to one replica, per-role stale-data settings.
Treat it as a per-read consistency budget rather than a cluster-wide setting: classify reads by staleness tolerance, decide where a correctness risk should be converted into an availability event, and be explicit that read scaling on replicas is buying throughput with a weaker consistency model the application has to be designed around.
## What a replica read actually is A Redis replica receives the primary's stream of write commands and applies them in order, asynchronously. There is no coordination on the read path: the replica answers from whatever it has applied so far. So a read from a replica is a **consistent point-in-time view of the primary's history at some earlier offset** — it is not garbage, it is not a mix of old and new, it is simply the past. Every ordering the primary produced is preserved; only currency is lost. That one sentence generates every anomaly below. Redis offers **no bound** on how far in the past that offset is. In a healthy cluster the gap is sub-millisecond; under write bursts, during a fork for an RDB save, on a saturated link, or while the connection is broken entirely, it can be seconds or minutes. Nothing in the protocol caps it, and nothing tells the reading client how stale its answer was. ## The three anomalies **No read-your-writes.** A client writes to the primary, gets `+OK`, then reads the same key from a replica and sees the previous value — or nothing, if the key is new. Any flow that writes and then redirects to a read path (form submit → detail page, write API → read API) is exposed. This is not a bug or a misconfiguration; it is the definition of asynchronous replication. **Non-monotonic reads.** Replicas advance independently, so replica A can be at offset 1000 while replica B is at 940. A client load-balanced across both reads the new value from A, then the old value from B — the observed state moves *backwards in time*. This surprises engineers far more than plain staleness, because a single client sees causality violated. The mitigation is session stickiness: pin a client to one replica for the duration of a session so at least its own view advances monotonically. **Unbounded staleness during link loss.** If the replica loses its connection to the primary, it does not stop serving. With the default `replica-serve-stale-data yes`, it keeps answering from a dataset frozen at the moment the link broke, for as long as the outage lasts. The data is not merely a few milliseconds behind — it is arbitrarily old, and the client cannot tell. ## The observable: INFO replication `INFO replication` is where you see all of this from outside: - On the primary, `master_repl_offset` is the total number of bytes of replication stream it has produced. Each connected replica is also listed with its own acknowledged offset. - On the replica, `slave_repl_offset` is how much of that stream it has processed. **The delta between the primary's `master_repl_offset` and a replica's offset is your lag metric** — measured in bytes of stream, not seconds, so translate it against your write throughput. - `master_link_status` is `up` or `down`. `down` means the replica is currently disconnected and whatever it serves is frozen. - `master_last_io_seconds_ago` shows how long since the replica heard from the primary — a useful early warning while the link is technically still `up`. Alert on both: offset delta above a threshold, and `master_link_status:down` for longer than a few seconds. Comparing offsets *across* replicas is also what tells you whether non-monotonic reads are a live risk in your fleet. ## What `replica-serve-stale-data no` changes This directive only takes effect when the link is down (or the replica is still performing its initial synchronization and has no usable dataset yet). Default `yes`: serve whatever is held. Set to `no`: the replica replies `-MASTERDOWN Link with MASTER is down and replica-serve-stale-data is set to no` to essentially every data command. A small allowlist still works so you can operate the node — `INFO`, `REPLICAOF`, `PING`, `SUBSCRIBE`, `SHUTDOWN`, `AUTH` and similar. So the setting converts a **correctness risk into an availability event**. That is a per-workload decision, not a global best practice. Serving a cached rendered page a few minutes late is harmless; serving a stale entitlement, balance, feature flag or rate-limit counter can be worse than returning an error the caller can retry against the primary. A common shape is: `no` on replicas backing correctness-sensitive reads, `yes` on replicas backing pure caches. ## Practical stance Decide staleness tolerance **per read**, not per system. Route reads that must reflect a just-completed write to the primary — for a short window after that client's write, or permanently for that endpoint. Send bulk, cacheable, tolerant reads to replicas. Pin sessions to one replica if a client must never see time run backwards. And monitor offset lag, because every anomaly here is a function of it. TTL-expiry behaviour on replicas is a separate mechanism covered by the eviction/expiry material; and in Redis Cluster a connection must send `READONLY` before a replica will serve it at all, which is covered with cluster redirection.
- A user updates their profile and the very next page load shows the old value, because reads go to replicas. How do you fix it without giving up read scaling?Route reads to the primary only on the narrow path that notices: pin that user's session to the primary for a few seconds after a write, or until you have observed a replica offset at or beyond the offset at write time. Even simpler, return the newly written value in the write response so the page never needs to re-read. Both keep the bulk of read traffic on replicas while removing the read-your-writes anomaly where it actually hurts.
- Would you set `replica-serve-stale-data no` everywhere as a default?No — it converts a correctness risk into an availability event, and which one you prefer depends on the data. For a replica fronting a rendered-page or object cache, serving minutes-old data during a link blip is harmless and erroring would cause a needless outage. For entitlements, balances, or anything a decision is made on, an explicit `-MASTERDOWN` the caller can retry against the primary is far better than silently authoritative stale data. Set it per replica role.
- How do you actually measure replication lag, and what's the catch with the number you get?Compare `master_repl_offset` on the primary against `slave_repl_offset` on the replica, or read the per-replica `offset` lines in the primary's `INFO replication`. The catch is that the delta is in **bytes of replication stream**, not seconds — the same byte gap means milliseconds under heavy writes and minutes under light ones. Pair it with `master_last_io_seconds_ago` and `master_link_status`, and alert on all three rather than on bytes alone.
A replica is a newspaper edition, not a live feed. Every story in it is internally consistent and really happened — it just went to press at some point you don't control. Two newsstands can hold different editions, so walking between them can make the news appear to go backwards.
saying these in an interview costs you the question
- Assuming replica reads are always current, or that lag stays in microseconds regardless of load
- Claiming Redis bounds staleness or that a client can request a maximum-staleness read
- Thinking a replica returns a torn or inconsistent mix of old and new state, rather than a consistent older snapshot
- Not realizing two replicas at different offsets let a single client observe a value going backwards
- Believing `replica-serve-stale-data no` affects normal operation — it only applies when the link is down or the initial sync hasn't finished
- Reading the offset delta as a time measurement instead of bytes of replication stream